discover_mcp_tools binds the owner secret scope only around the config load (#113746), so a
routed profile's reconciliation ran with no ambient scope. The adopter's stdio identity then
resolved unscoped, and with a source-tagged secret name get_secret raised UnscopedSecretError:
the stack's None sentinel kept that safe (the share was refused) but two profiles holding the
same value never shared the owner's child.
The omitted-name config load and the per-name identity resolution now run under this profile's
own secret scope (_owner_secret_scope), outside the registry lock.
Salvage resolution: the out-of-lock, once-per-name resolution, the None refuse sentinel and the
per-server refusal were already on the stack (_adopter_identity_digest / resolved_ids), so this
keeps that one implementation and takes the contributor's scope binding. The contributor's
unscoped-routed test is folded as an assertion into the kept multi-credential test (test budget);
their 'one unresolvable identity refuses only that share' test duplicates the kept 'boom' case and
is dropped, as are the test tweaks written against their _resolved_identity signature.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
(cherry picked from commit a8627071e7367cd544af77f921b4403a0b6c3e36)
_adopter_identity_digest returned "" for a config with no url/command and
for an unavailable live endpoint, but None for a resolver failure. Only
None is refused unconditionally by _same_server_route; "" is a comparable
string, so two unconnectable sides (an owner record of "" and an adopter's
"") compared equal and adopted. Return None on every can't-resolve branch
so there is a single refuse sentinel.
The non-owner reload test used an empty config and relied on that
""=="" match; give it a connectable URL config so the adopter resolves a
real identity.
The adopter resolves each foreign-held name's connection identity before
taking the registry lock, and a resolver failure is caught per server so
it refuses adoption of that server alone. Nothing pinned that: dropping
the per-server except would let one broken secret backend abort discovery
for the whole scope and stop every healthy shared connection from being
adopted. Extend the existing credentials test so a raising resolver for
one server leaves a healthy same-identity connection adopted.
Every multiplex discovery pass resolved an identity (secret-scope reads,
PATH lookup, live-endpoint probe) for each judged name, although the
digest is only read by a cross-profile `_same_server_route` comparison,
and that needs another profile's connection for the name. A profile's
first pass, and names it owns itself, paid for nothing. The first
registry-lock snapshot now also collects the names held under a foreign
scope and only those are resolved, still outside the lock; a foreign key
that appears after the snapshot has no digest and is refused, as before.
The same up-front loop let any resolver error other than
LiveEndpointUnavailable escape and abort discovery for the whole scope,
even for servers that share nothing. Such an error now refuses adoption
for that one server (None digest, fail-closed) and is logged once at
warning.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The module function `_resolved_identity(name, config)` (adopter
recomputation) shared its name with the attribute
`server._resolved_identity` (owner's published digest) and the
`resolved_identity` kwarg; a gateway test monkeypatched the function
while its fake set the attribute, which read as one thing. Renaming the
function to `_adopter_identity_digest` makes the owner/adopter split
visible at every call site.
Cross-profile adoption only works while the owner (transport) and the
adopter (`_resolved_identity`) hash byte-identical inputs, but each side
assembled the list itself: the stdio owner unpacked `_stdio_launch` and
re-listed `[command, safe_env, stdio_cwd]`, and both HTTP sides ran
`_http_endpoint` -> `_apply_identity_header` as separate copies. A new
launch field or header overlay added on one side would silently stop
every share (fail-closed, but no error).
`_connect_inputs(name, config)` now returns the list both sides digest
(stdio `[command, env, cwd]`, HTTP `[url, headers]`) plus the configured
header names the strict-redirect boundary needs, so `_run_stdio`,
`_run_http` and the adopter hash the same object by construction.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
The stale-overlay loop re-derived each name's config by hand
(`servers` first, else the profile config) right next to `resolved_ids`,
which is keyed off `judged`. Two copies of one precedence rule means an
edit to either lets the static config and the resolved digest compared
in `_same_server_route` come from different sources. Read both from
`judged`.
The stack's shape gate allows two invariant tests. Keep the real
two-profile gateway tests (secret-source env on stdio, identity_header
value_from: profile on HTTP) and drop the unit-level npx/runtime-file/
published-digest and ssl_verify/strict_redirect_headers parametrized
tests; the fixture updates that make existing multiplex tests publish an
owner identity stay.
_same_server_route recomputed the adopter's resolved identity per
candidate while _register_connected_into_current_scope held _core._lock:
a PATH lookup for stdio commands, secret-scope reads and the live-endpoint
AppResolver probe all ran under the global MCP registry lock, repeatedly
per live key. Resolve once per judged name before taking the lock and
pass it in via a resolved_identity kwarg (None refuses the share).
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Both are consumed by the HTTP transport but were in neither config_fingerprint nor
_connection_identity, so a profile whose config differs only in TLS verification or
redirect-header policy adopted another profile's live connection under that profile's policy.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
(cherry picked from commit 8b7175a2cf0e1935ea77805db9b99e4cba9fdcea)
The connecting loop hashed the resolved inputs, then the transport resolved them again; a runtime-file rotation between the two reads published endpoint A's digest alongside a session opened with B. The transport now resolves once per attempt (_http_endpoint / _stdio_launch), connects with that value and publishes its digest; the adopter recomputes through the same resolvers. The loop clears the digest per attempt.
(cherry picked from commit f543a4d3aadcdf7ec10093ebb99e14a740d43494)
The cross-profile identity digest covered config headers, identity_header,
secret-source env and cwd, but not two more per-profile resolutions the
transport performs: _run_http() swaps in a server_json live endpoint (URL and
bearer token) read from the active profile, and _run_stdio() resolves a bare
npx/npm/node through _resolve_stdio_command(), which can land under the
profile's own HERMES_HOME. Two profiles with those differences hashed equal
and could share one live connection.
_resolved_identity now resolves both exactly as the transport does. The new
test module imports the repository YAML module and passes explicit encodings
so it collects under scripts/run_tests.sh and clears the Windows-footgun gate.
(cherry picked from commit 637872d55bc81609657ed82139546709f6b99ccd)
_same_server_route judged "same credentials" from the static config alone, but a
connection is also opened with per-profile values the config never shows: a stdio
child's env carries every external secret-source value (Bitwarden, 1Password,
secrets.command) from the owner's secret scope, identity_header value_from: profile
resolves to the owner's profile name, and a stdio child's default cwd is the owner
session's runtime cwd. A multiplexed profile with a byte-identical config therefore
adopted the owner's live connection and its MCP calls ran as the owner.
The connecting task now records a SHA-256 of those resolved inputs on every connect
attempt, in its own scope; cross-profile adoption and the stale-overlay check
recompute it under the adopter's scope and refuse to share on a mismatch. Only the
hash is kept. Same-profile checks are unchanged.
(cherry picked from commit e59d0dfb9918db03482ccd4bf5a7599efc3cd72e)
Gate mutation showed deleting the clear of _session_row_replay_pending after
a successful write survived the suite. A flag that never clears re-appends
any unmarked history dicts on every flush (e.g. history rehydrated per turn
on a cached agent). Assert it is None after the turn-3 flush.
When the heal recreated the session row but its single retry write then
failed (database locked, turn lease, disk), the replay was carried only by
the per-call _replay_history argument. The next flush found a live row, hit
no FK error, ran no heal, and stamped the history prefix durable without
writing it: the durable transcript silently lost everything but the tail
(scratch probe: rows ['a2'] instead of ['one','a1','two','a2']).
Set _session_row_replay_pending in the heal branch right after the markers
are stripped, so any exit after a heal (recreate failed, retry failed)
leaves the replay pending until a write succeeds. That makes the separate
recreate-failed assignment and the _replay_history parameter redundant;
the retry is a plain _adoption_budget=0 call. The fails-closed test now
covers FK -> heal -> locked retry -> full replay on the next flush.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
When the session row is gone and the heal cannot recreate it, the flush
fails closed with the flush markers already stripped. The next flush then
recreates the row up front via _ensure_db_session, so no FK error fires,
no heal runs, and _db_flush_collect stamps the history prefix as durable
without writing it. The recreated row ends up holding only the new tail
(probe: rows ['a2'] instead of ['one', 'a1', 'two', 'a2']).
Remember the failed heal per session id and have the next flush for that
session replay the history prefix; clear it once a flush succeeds.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
The two heal tests stayed green with agent/session_persistence.py reverted
to before the muted-turn and confirmed-missing fixes, so neither fix was
guarded. Extend the existing tests (no new test functions):
- recreate test: a heal during a muted notification turn must hide only
that turn's new rows, not the replayed history (reverted: history rows
and their in-memory dicts turn 'hidden').
- fails-closed test (renamed from ..._fails_open_... to match its
docstring): an FK failure while the row still exists must not heal or
duplicate the transcript (reverted: `one, a1, one, two`), and a raising
parent-row lookup for a delegate child must fail the flush closed
instead of leaking OperationalError.
The getattr(..., 787) fallback guarded a constant every supported
Python (>=3.11) always defines; reference it directly and drop the
module-level alias.
The FK classifier behind session_row_missing also matches the sessions
table's own parent_session_id / system_prompt_hash FKs, and session
create is an upsert. Healing on the classification alone could replay
the whole transcript into a still-live row and duplicate it. Heal only
when get_session(session_id) confirms the row is gone.
The delegate-child branch's get_session(parent_id) was unguarded: a
store error there escaped the flush's error handler. Route both lookups
through one guarded helper (as the compression-tip helper does); a failed
lookup is not proof of deletion, so it falls through to the existing
"will retry next flush" return.
While here: _db_flush_failed now returns which retry to take
("adopted" / "healed" / None) instead of the caller string-comparing
_last_persistence_error_cause; its docstring covers the heal branch;
messages is required (the only caller always passes it); the marker
reset reuses _strip_persistence_markers. Tests drop the change-detector
assertion on the lingering cause after a successful heal and the Scenario
B docstring now says fail-closed, which is what it asserts.
The session-row heal retried with conversation_history=None so the
history prefix would be replayed onto the recreated row. On a muted
notification-reply turn that also made _db_flush_collect treat every
history row as new: each replayed row was written display_kind=hidden
and the flag was stamped onto the live in-memory history dicts, hiding
the user's whole past transcript from pollers from then on.
Keep the history set on the heal retry and pass an explicit
replay_history flag instead: history rows are written again (not stamped
durable) but keep their original visibility, and only this turn's new
rows get the mute treatment.
A session deleted while its chat is still running is recreated under the
same id with the full in-memory transcript on the next save (the gateway
session-key mapping expects the id to be stable). Say so next to
`hermes sessions delete`, in English and the zh-Hans mirror.
Fixes#123583
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
Rewrite the heal test to the real turn shape (messages = history + tail
with conversation_history=history), which is the shape that exposed the
partial-replay bug, and fold the 'deleted twice' case in as extra delete
rounds. Keep the fail-open Scenario B test. Stack budget is two tests.
Deleting a parent session cascades to its delegate children. A live child
then can never be recreated: create_session(parent_session_id=...) hits
the parent FK on every flush for the agent's lifetime. When the recreate
fails and the parent row is gone, create once without the parent for that
call only; agent._parent_session_id is restored because the relay and
hooks still key on it.
Without entries in _STORAGE_FAILURES and _PERSISTENCE_CAUSE_EXPLANATIONS a
failed heal fell back to the generic disk/lock advice, which sends users
chasing the wrong problem.
The heal predicate (errno 787 or a 'foreign key constraint' substring)
and the classifier phrase were two separate definitions of the same
failure. Move the SQLITE_CONSTRAINT_FOREIGNKEY code check into
classify_persistence_error, document the bucket, and branch the heal on
agent._last_persistence_error_cause == 'session_row_missing' so the two
can't drift.
The heal retry passed the caller's conversation_history through, so
_db_flush_collect's id()-based history shortcut treated the prefix as
already durable and skipped it. The row was recreated with only the new
tail and the flush returned True: a silent partial restore. On the
session_row_missing retry, pass conversation_history=None so the full
in-memory transcript lands on the recreated row.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
`_ensure_db_session` trusts the cached `_session_db_created` flag as proof
the row exists, and the flush path only retries row creation while that
flag is False. Any store-side removal of the row under a live agent —
`hermes sessions delete`, the Desktop/web delete, bulk prune, a
profile-repair move, an in-place store rebuild — leaves the flag stale, so
every later turn's append fails the FK and is dropped with one WARNING per
turn. The agent keeps answering; the durable transcript silently stops
growing, and because the failed transaction leaves no rows behind there is
no post-hoc trace in the store.
Classification (`hermes_state_errors.py`): add the `session_row_missing`
cause — matched by `SQLITE_CONSTRAINT_FOREIGNKEY` (787) when the code
survives, else the RPC-wrapped phrase — so the turn-end explanation names
the real failure instead of "unknown".
Heal (`agent/session_persistence.py::_db_flush_failed`): on that cause,
drop the stale flag, reset the flush markers, call `_ensure_db_session()`,
and replay once within the flush's existing `_adoption_budget`. Because the
deletion erased the session's message rows too, the replay clears the
per-message persisted markers so the FULL in-memory transcript lands on the
recreated row, not just the current tail (mirrors
`_db_flush_adopt_compression_tip`). If row creation also fails, return
False without appending into a guaranteed rollback — fail-open, batch stays
unmarked for the next flush. No new fail-closed path.
(cherry picked from commit ad97047da0f9801a4fdca4478bfd68085f02c3a8)
assert_export_safe grew an include_inactive flag whose only production caller
(the console export guard) always passes True, leaving the live-only branch
and its default unused. Drop the parameter and count every row, which is what
the transfer export materializes; the console call and docstring follow. The
existing guard tests seed live rows only and are unchanged.
export_all re-derived the live-row clause by hand; use the existing
_active_clause(include_inactive, False) so "live row" stays defined in one
place. Behaviour is unchanged.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
A transfer row's active/compacted flags were read by truthiness, so a
hand-edited or foreign JSONL row with "active": "0" (truthy string) imported
live - putting compaction-archived turns back into model context - and
"active": null archived a live row. Coerce both flags with the same
_coerce_or(..., int, default) used for the int session columns: a missing,
null, or unparsable active flag means live (older exports carry none), and
compacted defaults to 0.
The same pass partitions rows into live and archived up front, replacing the
id()-set complement, the dead "_row_id" guard (_insert_message_rows always
sets it), and the subtract-and-reparse counter fixup: message_count and
tool_call_count now come straight from the live rows, with identical values
for well-formed exports. The kept round-trip test re-imports its payload with
stringified/null flags and asserts every row keeps its state.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
The console export guard tells users whose session exceeds the in-memory cap
to use the Sessions page's streaming Export instead, but that endpoint read
get_messages() with the default live-only projection. Its JSON therefore
dropped every compaction-archived turn, and re-importing it collapsed a
9-turn compacted session to its 3 live rows - the exact loss this stack fixes
for the console path.
Page with include_inactive=True (keyset after_id is only incompatible with
include_compacted). Each row now carries its active/compacted flags, which
import_sessions re-archives. Probe: 9-turn compacted session -> web export
(11 rows) -> web import -> display 9, live 3, message_count 3 (was 3/3).
_import_session_row now calls _insert_message_rows(..., prune_checkpoints=False),
so the positional-only failing_insert stub raised TypeError before its own
"interrupted write" and the rollback test failed. Forward **kwargs so the stub
still performs the insert and then fails, keeping the rollback assertion honest.
test_adoption_keeps_every_row_checkpoint_sidecar already asserts every
row (content, active/compacted flags, checkpoint sidecar) survives
adoption and that the donor retires, so the end-to-end resume test that
checked the same compacted-history carry-over is redundant. Keeps the
stack at two invariant tests.
The console `sessions export` backup went through export_session /
export_all, whose batched read filters `AND active = 1`, so an
export-then-import restore silently dropped every compaction-archived
turn (display 9 -> 3 in the C13 probe). export_all now takes
include_inactive (storage-order rows with their active/compacted flags,
which import_sessions restores as archived), the console export passes
it for both single-session and all-session exports, and
assert_export_safe sizes the same projection it guards.
Ported from #123267's transfer/backup hunks; its import-normalisation
hunks are superseded by #122680's checkpoint-correct import.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
import_sessions inserted every row live and pruned shadowed checkpoints before restoring the archived flags. A newer archived carrier (a rewound turn) then stripped the newest live checkpoint, and every archived row lost its own. Rows written before #102374 pruning each carry one, so adopting such a donor fed the model a context without its checkpoint while the donor was retired as unrecoverable.
The import now prunes once the archived rows are archived again, keyed on the live rows only; every other writer still prunes on insert. The blank line at the end of the adoption test file is dropped.
(cherry picked from commit f9ad62b37517cbbdc27f75e9e58900048258bd7b)
session.resume with a `profile` for an id that profile's store lacks
adopts the lineage from the launch store and retires the donor with
end_reason=adopted_by_profile, which is not recoverable.
Adoption copied live rows only. In-place compaction (the default)
archives earlier turns under the same id (active=0, compacted=1), so
those turns reached neither the profile copy nor any live row: a
15-turn chat with one compaction showed 30 messages before and 22
after, with the rest behind the unrecoverable archive. The divergence
guards compared live counts, so they could not see the loss.
Adoption now exports every row with its active/compacted flags
(export_session*/_lineage gain include_inactive), and import_sessions
restores a row exported archived as archived, so the copy shows the
same display history and feeds the model the same live context. Its
counters count live rows only. Both divergence guards compare all rows,
so a donor is never retired while any of its rows is missing here.
fe8b643db6 kept adoption on live rows because import_sessions inserted
every row as live context. That still holds for the display projection
(include_compacted), which the JSON export, /save json and the
dashboard keep excluding; adoption uses the new storage-order export.
(cherry picked from commit c1ae532f973277a9a10da79ada3534e827a492dc)
The remote guard contract only pinned refusal of an existing target-side
binary. Nothing pinned the other half: a path the backend proves absent must
stay writable, or a guard that over-reports "exists" (e.g. treating a
"missing" probe as present) would block every new .pdf/.sqlite on Docker/SSH
without a red test. Extend the kept T1 test instead of adding a function
(test budget); verified it fails 6/6 under a missing->exists mutation.
The substitution pairs were built by looping over a one-element tuple (left
over from a multi-pair version), and _bash_safe() only existed to do a
function-local import from a module already imported at the top. Import
_bash_safe_path directly and build the three pairs as a literal list.
The helper docstring restated the inline why-comments (exact probe string,
backend $HOME tilde fallback, which statuses prove absence). Keep only the
contract callers need: three return values and that anything not proven
absent is "unavailable" and must fail closed. Inline comments stay.
Keep the two contracts that pin #122662: (1) write_file / patch replace /
V4A update against a binary that exists only on an unhinted non-local
backend (VercelSandboxEnvironment, scoped config claiming 'local') is
refused with bytes untouched, for .sqlite and .pdf; (2) a probe transport
failure (raise or error return) fails closed. Creation-allowed, locality
and tilde-fallback cases are covered by the first invariant's setup or by
tests/tools/test_binary_document_write_guard.py, and the salvage shape
caps a stack at two invariant tests.
_probe_regular_file returns "bad_size" only after `[ -f ]` succeeded and
`wc -c` output was unparseable, so a regular file IS present on the
execution target. Treat it as "exists" (the precise existing-binary
refusal) instead of the generic "unavailable" retry message; both refuse,
but the retry hint is wrong for a file that is known to exist. Mirrors
how the other _probe_regular_file callers treat bad_size.
Idea and the original report/Docker reproduction come from #122663, the
first submitted fix for this bug.
Fixes#122662
Co-authored-by: liuzikaii <2319582736@qq.com>
_check_binary_document_write stat'ed the controller host only, so a binary
(.sqlite/.pdf/...) that existed solely in the task's execution target
(Docker/SSH/... namespace) was treated as a new file and destroyed by a
plain-text write_file/patch that even reported verified:true (#122662).
The existence decision now goes through one tri-state helper
(_target_regular_file_state: exists/absent/unavailable) that asks the LIVE
file-ops layer - the same backend the write executes on. Host-backed envs
keep today's Path.is_file semantics (OSError -> proceed); other backends are
probed via ShellFileOperations._probe_regular_file, whose missing/not_regular
answers prove absence while every other status fails closed with a retry
message. Locality comes from the live environment object
(_file_ops_uses_host_paths), never env_type strings or class-name hint tables
(VercelSandboxEnvironment is unclassified there). The PDF and generic binary
refusal messages are unchanged verbatim; opaque-document and SQLite-sidecar
unconditional refusals are untouched.
_stale_overwrite_blocker keeps its host-only probe on purpose: remote reads
never record a full_write_baseline (version stability requires host metadata),
so converting it would refuse every remote overwrite of a file the task fully
read. Tracked as follow-up.
(cherry picked from commit 8b5eb26d57d7ef7d1d975de0ef5e1e8abd9a513d)
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
Follow-up to the #123592 salvage. The canonical endpoint now comes from the
provider profile, which also covers OpenRouter (absent from PROVIDER_REGISTRY),
so pinning https://openrouter.ai/api/v1 keeps the native catalog too. Both
sides go through normalize_route_base_url, the helper the rest of the route
comparisons use, so scheme/host case and default ports no longer read as a
relay. The provider is normalized once.
Follow-up to the #123857 salvage. request_overrides is static config, so the
drop warning fired on every turn, retry and subagent call. It now fires once
per process. The Astra sanitizer docstring no longer suggests it handles the
key. The test is parametrized per route, and it pins the extra_body escape
hatch that the warning points to.
request_overrides are merged into the top-level Responses.create() kwargs,
where prompt_cache_options has no SDK parameter: the call fails with
TypeError before any request is sent, on every route. Drop the key after
the merge with a warning, matching the wire-boundary sanitizer precedent.
Wire-only fields still reach capable proxies via extra_body.
(cherry picked from commit c8e8227fdde75236246233565f9cb8ed74c6bd9c)