Every multiplex discovery pass resolved an identity (secret-scope reads,
PATH lookup, live-endpoint probe) for each judged name, although the
digest is only read by a cross-profile `_same_server_route` comparison,
and that needs another profile's connection for the name. A profile's
first pass, and names it owns itself, paid for nothing. The first
registry-lock snapshot now also collects the names held under a foreign
scope and only those are resolved, still outside the lock; a foreign key
that appears after the snapshot has no digest and is refused, as before.
The same up-front loop let any resolver error other than
LiveEndpointUnavailable escape and abort discovery for the whole scope,
even for servers that share nothing. Such an error now refuses adoption
for that one server (None digest, fail-closed) and is logged once at
warning.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The module function `_resolved_identity(name, config)` (adopter
recomputation) shared its name with the attribute
`server._resolved_identity` (owner's published digest) and the
`resolved_identity` kwarg; a gateway test monkeypatched the function
while its fake set the attribute, which read as one thing. Renaming the
function to `_adopter_identity_digest` makes the owner/adopter split
visible at every call site.
Cross-profile adoption only works while the owner (transport) and the
adopter (`_resolved_identity`) hash byte-identical inputs, but each side
assembled the list itself: the stdio owner unpacked `_stdio_launch` and
re-listed `[command, safe_env, stdio_cwd]`, and both HTTP sides ran
`_http_endpoint` -> `_apply_identity_header` as separate copies. A new
launch field or header overlay added on one side would silently stop
every share (fail-closed, but no error).
`_connect_inputs(name, config)` now returns the list both sides digest
(stdio `[command, env, cwd]`, HTTP `[url, headers]`) plus the configured
header names the strict-redirect boundary needs, so `_run_stdio`,
`_run_http` and the adopter hash the same object by construction.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
The stale-overlay loop re-derived each name's config by hand
(`servers` first, else the profile config) right next to `resolved_ids`,
which is keyed off `judged`. Two copies of one precedence rule means an
edit to either lets the static config and the resolved digest compared
in `_same_server_route` come from different sources. Read both from
`judged`.
_same_server_route recomputed the adopter's resolved identity per
candidate while _register_connected_into_current_scope held _core._lock:
a PATH lookup for stdio commands, secret-scope reads and the live-endpoint
AppResolver probe all ran under the global MCP registry lock, repeatedly
per live key. Resolve once per judged name before taking the lock and
pass it in via a resolved_identity kwarg (None refuses the share).
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Both are consumed by the HTTP transport but were in neither config_fingerprint nor
_connection_identity, so a profile whose config differs only in TLS verification or
redirect-header policy adopted another profile's live connection under that profile's policy.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
(cherry picked from commit 8b7175a2cf0e1935ea77805db9b99e4cba9fdcea)
The connecting loop hashed the resolved inputs, then the transport resolved them again; a runtime-file rotation between the two reads published endpoint A's digest alongside a session opened with B. The transport now resolves once per attempt (_http_endpoint / _stdio_launch), connects with that value and publishes its digest; the adopter recomputes through the same resolvers. The loop clears the digest per attempt.
(cherry picked from commit f543a4d3aadcdf7ec10093ebb99e14a740d43494)
The cross-profile identity digest covered config headers, identity_header,
secret-source env and cwd, but not two more per-profile resolutions the
transport performs: _run_http() swaps in a server_json live endpoint (URL and
bearer token) read from the active profile, and _run_stdio() resolves a bare
npx/npm/node through _resolve_stdio_command(), which can land under the
profile's own HERMES_HOME. Two profiles with those differences hashed equal
and could share one live connection.
_resolved_identity now resolves both exactly as the transport does. The new
test module imports the repository YAML module and passes explicit encodings
so it collects under scripts/run_tests.sh and clears the Windows-footgun gate.
(cherry picked from commit 637872d55bc81609657ed82139546709f6b99ccd)
_same_server_route judged "same credentials" from the static config alone, but a
connection is also opened with per-profile values the config never shows: a stdio
child's env carries every external secret-source value (Bitwarden, 1Password,
secrets.command) from the owner's secret scope, identity_header value_from: profile
resolves to the owner's profile name, and a stdio child's default cwd is the owner
session's runtime cwd. A multiplexed profile with a byte-identical config therefore
adopted the owner's live connection and its MCP calls ran as the owner.
The connecting task now records a SHA-256 of those resolved inputs on every connect
attempt, in its own scope; cross-profile adoption and the stale-overlay check
recompute it under the adopter's scope and refuse to share on a mismatch. Only the
hash is kept. Same-profile checks are unchanged.
(cherry picked from commit e59d0dfb9918db03482ccd4bf5a7599efc3cd72e)
The helper docstring restated the inline why-comments (exact probe string,
backend $HOME tilde fallback, which statuses prove absence). Keep only the
contract callers need: three return values and that anything not proven
absent is "unavailable" and must fail closed. Inline comments stay.
_probe_regular_file returns "bad_size" only after `[ -f ]` succeeded and
`wc -c` output was unparseable, so a regular file IS present on the
execution target. Treat it as "exists" (the precise existing-binary
refusal) instead of the generic "unavailable" retry message; both refuse,
but the retry hint is wrong for a file that is known to exist. Mirrors
how the other _probe_regular_file callers treat bad_size.
Idea and the original report/Docker reproduction come from #122663, the
first submitted fix for this bug.
Fixes#122662
Co-authored-by: liuzikaii <2319582736@qq.com>
_check_binary_document_write stat'ed the controller host only, so a binary
(.sqlite/.pdf/...) that existed solely in the task's execution target
(Docker/SSH/... namespace) was treated as a new file and destroyed by a
plain-text write_file/patch that even reported verified:true (#122662).
The existence decision now goes through one tri-state helper
(_target_regular_file_state: exists/absent/unavailable) that asks the LIVE
file-ops layer - the same backend the write executes on. Host-backed envs
keep today's Path.is_file semantics (OSError -> proceed); other backends are
probed via ShellFileOperations._probe_regular_file, whose missing/not_regular
answers prove absence while every other status fails closed with a retry
message. Locality comes from the live environment object
(_file_ops_uses_host_paths), never env_type strings or class-name hint tables
(VercelSandboxEnvironment is unclassified there). The PDF and generic binary
refusal messages are unchanged verbatim; opaque-document and SQLite-sidecar
unconditional refusals are untouched.
_stale_overwrite_blocker keeps its host-only probe on purpose: remote reads
never record a full_write_baseline (version stability requires host metadata),
so converting it would refuse every remote overwrite of a file the task fully
read. Tracked as follow-up.
(cherry picked from commit 8b5eb26d57d7ef7d1d975de0ef5e1e8abd9a513d)
Follow-up to the ssh path fix salvaged from #121686.
- Read the ssh anchor raw (session record, override, TERMINAL_CWD): the
shared workspace-root helper expands ~ on the Hermes host, so
TERMINAL_CWD='~/proj' still resolved into the container home.
- Resolve ~ to the remote home the SSH environment detects at connect,
bringing the environment up through the file tools' own creator
(_get_file_ops, same cwd and cache) when none is live. A failed bring-up
is remembered per container for 30s so one call's several resolutions
don't each retry; a live environment is always used first. SSHEnvironment
now records whether the home was detected, and a guessed /home/<user>
(echo $HOME failed) is not used. Results are
absolute and stable from the first call (read tracking and staleness
checks key on them), and '..' normalizes to the real target: relative
traversal like ../../../etc/x from ~ was refused on main and slipped past
the sensitive-path guard on the PR head.
- If the remote home cannot be detected, an ssh ~-path that climbs above ~
cannot be classified; the write guard refuses it.
- ~user passes through for the remote shell instead of becoming ~/~user.
- coerce_ssh_remote_cwd maps paths under the host subprocess home onto ~/,
except when that home is the OS user's real home.
- The outside-workspace warning compares in the remote namespace (it fired
on every correct relative write when the anchor was ~).
- The backend type is looked up once per resolution again (the PR head did three per local path).
The session cwd can already be the Hermes host's /opt/data/home. That
path is not on the SSH target, so the implicit cd exits 126 and the
command never runs.
(cherry picked from commit 4e99397bb4e18416b6a3e0b850d51604c0ad0d1d)
Relative writes and a bare ~ were expanded against the container
subprocess home and then sent to the SSH target, which does not have
that directory.
(cherry picked from commit f4a2549d8276762e53ed2d773dbd4de8c7e74b3b)
Drop the unreachable consumer-side _capped() wrapper (the pump already
stops before enqueueing past the cap), bound the pump queue and make
every put poll the stop event so an abandoned consumer can't wedge the
pump thread. Timeout message derives from _RECV_TIMEOUT_S; sample_rate
reuses DEFAULT_XAI_SAMPLE_RATE.
GeminiStreamer and the synchronous _generate_gemini_tts both sent the key
as a ?key= query param, so requests' HTTPError text (full prepared URL)
carried it into TTS logs on any 4xx/5xx — and text_to_speech_tool logs
unexpected exception text. Send it as the x-goog-api-key header in both
paths instead.
The unterminated-think flush half of the original PR is dropped: current
main already normalizes each clause before provider use (4aac89b429).
(cherry picked from commit 02387442c9a3830e1b4b7a727722472f2f0f76c1)
The old shape used extra_headers and a guessed message framing, and the
tests mocked the seam. Rewritten against the verified wire protocol:
voice/language/codec/sample_rate ride in the URL query string, bearer
auth header, text.delta + text.done out, base64 audio.delta envelopes in
until audio.done. Covered by a loopback WS test suite that pins the wire
format, early chunk delivery, and error-envelope raises.
The pump enforces the 16 MiB per-sentence byte budget before enqueueing
frames and closes the socket when it is exceeded: the consumer-side
_capped() runs only after q.get(), so without a pump-side check a fast
upstream piles decoded PCM into the queue ahead of the consumer. A
consumer that stops early flips a stop event so the pump closes instead
of filling a queue nobody drains. A lowered-cap loopback test verifies
the producer stops consuming upstream data.
(cherry picked from commit e6b5e29930fc026a759a1c205eb0fa97dd4a5688)
The #37589 uv/uvx fallback spelled its directory table (~/.local/bin,
/opt/homebrew/bin, /usr/local/bin) inline in _launcher_fallback, which the
managed-runtime ratchet flags as an unreviewed known_path_table outside
hermes_platform/. known_dirs is the module whose contract is 'every table in
Hermes lives here', so add uv_tool_dirs() there and compose it at the call
site. Behavior is unchanged: same four directories probed in the same order
(managed <home>/bin first, then uv's install order), bare uv/uvx still
resolves under a GUI-style PATH that lacks them.
macOS Desktop/launchd processes inherit the bare /usr/bin:/bin:/usr/sbin:/sbin
PATH, which carries none of uv's install locations, so an MCP server configured
as `command: uvx` fails with ENOENT at execvp from Desktop even though it works
from an interactive terminal. The stdio resolver already falls back to
well-known install dirs for bare npx/npm/node; extend the same treatment to
uv/uvx, probing managed <home>/bin, ~/.local/bin (uv's installer default),
/opt/homebrew/bin and /usr/local/bin (Homebrew AS/Intel).
Design salvaged from #37665, #67125 and #67178.
Co-authored-by: Morad37 <mohamed.origami@gmail.com>
Co-authored-by: Ignacio Rodriguez <ignacio@agenticolabs.io>
Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
After the locate click, type now refuses unless the located editable is
document.activeElement, so characters are not delivered as page input when
focus stayed on body. The keystroke loop checks an abort signal between
characters; preview.act cancel (timeout or interrupt) and a local Stop both
set it, so queued keystrokes stop. A printable press on body/html is refused
unless allow_shortcut is set.
A session asked to clean up older Pythons removed the uv-managed base
interpreter its own venv depended on; the next boot died with 'uv
trampoline failed to spawn Python child process' and no agent tool could
repair it, because the agent itself no longer started (#58748). Prior
uninstall detection (85ce25687e) only flagged package-manager commands.
Add agent/runtime_self_protection.py and wire it into both layers:
- The approval floor (_floor_block) now blocks shell commands that
delete the running interpreter, its own venv, the pyvenv.cfg base, or
the uv-managed install directory — rm/rmdir/rd/del/erase/Remove-Item
with any flags, find <root> -delete, and uv python uninstall of the
running version (including --all). The floor runs before yolo /
approvals.mode=off / cron approve mode, so no session setting can
bypass it.
- The file-safety write classifier denies write/patch/move/delete to the
same paths, so the file tools cannot overwrite the interpreter either.
Only the runtime the process itself boots from is protected; every other
venv and interpreter on the machine stays manageable.
Fixes#58748
[ -f ] and [ -e ] follow symlinks, so a dangling link probed as missing and
read_file_raw reported not_found: V4A Add followed the link and created its
target, Move replaced the link, both reporting success. The shell size
probe, the compound read probe and the native stat now classify any
symlink entry as not a regular file.
The fence drops backend output around the read, not output printed while
the transport runs: a BASH_ENV DEBUG hook firing for base64 alone puts
text inside the payload that still decodes (TERM is b"LDL"), and the edit
paths wrote it back. Emit wc -c in its own fenced segment and hand bytes to
a writer only when their length matches; the od fallback is checked the
same way.
Consolidates #120559 by @JoaoMarcos44 into this PR. Both PRs landed on the same
shape for #120514 — read_file_raw must be a byte-preserving mutation boundary,
framed so the backend's merged stdout is never decoded as payload — but the
sibling one step in FRONT of it was still unfenced here.
_sample_file_bytes whitespace-joined the whole reply before base64-decoding it,
so a backend that announces something on connect had that noise decoded into the
sample: "TERM" is four base64 characters and prefixes the sample with b"LDL".
That sample is the binary-admission gate read_file_raw consults, so noise there
decides whether a file is editable at all and what a refusal reports about its
bytes. Same class as the read this PR already fixed, one call earlier.
_fenced_read() is that framing extracted once — sentinel, payload segment, and
the body's own exit status in its own trailing segment — and the three byte
transports (sample, base64, od fallback) now share it instead of repeating the
split/parse/decode triplet, with _failed_read() for the diagnostic they all
surfaced by hand.
The mock double for the sample transport moves to the fenced reply shape the
other doubles in that file already build, so it stays honest about the fence.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Validation treated every read_file_raw error as "the path is free" for an
Add target and a Move destination; only _apply_add checked not_found. With no
byte transport (native reads off, no base64 or od) an `Add new-src` then
`Move new-src -> dst` validated, and the mv replaced the existing dst while
the patch reported success.
_occupied() frees a path only on not_found (overlay-aware), and _apply_move
re-checks the destination right before mv, since validation's answer is stale
once earlier ops of the patch have applied.
Three defects in the byte-exact read this PR introduced.
base64 is not on every backend (busybox, distroless). The sample path
already degrades when it is missing; this one returned the read as a
failure, and read_file_raw is also _apply_add's existence check, which
treated any error as "the path is free". A backend with cat but no base64
turned `*** Add File` over an existing file into a silent overwrite that
reported success:
main: refused, "file already exists — use Update File"
PR head: b'KEEP ME -- months of work\n' -> b'clobbered by Add'
So: fall back to od (POSIX, in busybox) when base64 exits 127, and when
neither exists report a transport error. ReadResult grows not_found, set
only where the path is genuinely absent, and _apply_add refuses unless it
sees that flag — a read that FAILED can no longer pass as an absent path.
getattr keeps a producer without the field failing closed rather than
raising. The doubles in test_patch_parser that meant "absent" now say so.
The native fast path stat'd the path and then opened it, two lookups on a
name. A swap to a FIFO in between blocks the thread, and nothing times out
that. One O_RDONLY|O_NONBLOCK open, fstat on THAT descriptor, then read, so
a non-regular file is handed to the shell path and its timeout instead.
Tests: the od fallback round-trips byte-exactly and patches; an Add over a
file the backend cannot read is refused with the file intact; the native
read hands a FIFO to the shell rather than opening it. The first two go red
if their fix is removed. The third pins the property, not the race — with
one descriptor there is no window left to swap into, so reverting to
stat-then-open does not flip it.
Not addressed here: inserted text is still encoded UTF-8 regardless of the
file's declared encoding, so an edit adding non-ASCII to a latin-1 file
writes mixed bytes. That is the write half, it predates this PR, and it
needs source-encoding detection.
The byte-exact read whitespace-joined the whole reply before decoding, so
anything a backend prepends to its merged stdout became part of the base64.
A remote shell announcing TERM is four base64 characters, which decodes to
b"LDL" and lands at the head of the file the edit paths then write back:
patching HEADER\nVERSION=1\n left LDLHEADER\nVERSION=2\n on disk.
Fence the payload between a per-call sentinel, the shape the compound read
probe already uses and for the same reason. _new_sentinel's underscores are
outside the base64 alphabet, so noise that lands inside the fence fails
validation instead of decoding into bytes, and noise outside it is dropped.
The read's exit status rides in its own trailing segment, so a failed read
is still told apart from an empty file without && chaining or exit.
stderr stays merged rather than silenced: on a failed read that segment
carries the backend's own diagnostic ("No such file or directory"), which
the callers surface, and a reply with no fence at all hands the backend's
text straight back so a wrapper cd failure still reports itself.
The two doubles this PR added keyed off the old command string; they now
build a fenced reply from the sentinel in the command they were handed, so
they stay honest about the fence the transport depends on.
The new test drives the noisy backend through the real transport: the
existing tests use LocalEnvironment, which takes the native read path on
POSIX and never exercises base64 at all.
Review follow-up (ehz0ah, teknium1). Keying the cleanup on the
__HERMES_FENCE_ marker still rewrote a real line that contains that
text, and nothing has emitted the wrapper since the spawn-per-call layer
(d684d7ee7e), so the cleanup can only ever eat the file's own bytes.
The same class also hit mode="replace", the default: patch_replace (and
V4A past the 1000-byte binary sample) read the file through the text
transport, which decodes with errors="replace", so any byte UTF-8 cannot
decode came back as U+FFFD and was persisted on lines the edit never
touched, while the diff and the post-write check (same lossy read) showed
nothing.
_read_exact_bytes reads natively on the local POSIX host (regular files
only) and over base64 elsewhere; read_file_raw, patch_replace and
_verify_patch_persisted decode it with surrogateescape, which write_file's
encode already inverts (#79178), so every untouched byte is written back
exactly and no readable file becomes a read error (V4A's Add/Move/Delete
existence checks are unchanged). A garbled transport reply refuses
instead of writing stray output into the file. The Python linter parses
bytes so a declared legacy coding still lints clean. Local edits spawn
two fewer shells (replace 5 -> 3, V4A 9 -> 7).
read_file_raw feeds the V4A write-back, and it ran the whole file through the
terminal fence-leak cleanup, which deletes every OSC sequence and BEL byte. A
prompt script's title escape or a beep line was silently rewritten on apply, and
the diff and post-write hash both started from the stripped text, so nothing
showed it. A leaked wrapper always carries the fence marker: clean only those
lines when the content is going to be written back.
The four CI reds from #41225:
- persist_on_release is a background-only terminal modifier: join the
sandbox blocked sets (precedent: heartbeat, 9acd0d33b6) in
_TERMINAL_BLOCKED_PARAMS and the stub-drift tests' mirrors.
- The gateway shutdown sweep passes source="gateway_shutdown" so
persisted jobs are still killed on host exit; the two shutdown tests'
kill_all fakes now accept and assert that kwarg instead of raising
TypeError that _quiet_step silently swallowed.
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.
Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
(_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
explicit operator stops (process_manage kill, /stop slash + RPC mirror,
CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
the host is going away and survivors would become PPID=1 orphans
Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
A heredoc body always ends in a newline, so on the heredoc-stdin backends
(Modal, Daytona, Vercel) every stdin payload arrived with one extra byte.
write_file and patch verify an exact sha256 of the written file, so even
with the command grouped every write failed verification and left
content + "\n" on disk.
Feed the heredoc through a process substitution that re-emits the body
minus that last character. The heredoc stays outside <( ) because bash
3.2 mis-parses heredoc bodies inside it, and the reader tolerates read's
EOF status under an inherited set -e.
The tool-result spill no longer needs its +1 byte allowance for heredoc
mode, so the size check is exact on every backend again.
read_preview followed the global right-rail tab. follow() copied the
interacted zone into that id, so one group overwrote an explicit open
in another. Resolve from the hovered or focused zone, and skip that
copy when the explicit open lives in a different group. When more
than one preview is mounted, include active_tab_id and the open tabs.
A profile-scoped parent donated its environ to the host multiplexer, so a
named launcher was treated as the primary adapter owner and its platform
token became the primary claim. Spawn the host with served_profile_child_env
for the default root, mark multiplex active before that primary load, and
name the env-derived side in a duplicate-credential refusal.
Seeding the disk-watch baseline on first observation also swallowed the
case where the process started before login: file absent, then an
external `hermes mcp login` writes it, and the provider never reloaded.
Only skip the reload when the provider already holds tokens in memory.
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.
The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.
Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
The first CLI launch after a PM install starts tirith's download in the
background, so ensure_installed() returns None and the CLI printed
"tirith security scanner enabled but not available". The scanner was on
its way, not missing; the same warning also fired when lazy installs are
disabled by the operator's own policy.
missing_is_expected() tells those by-design states apart. The CLI logs
them and keeps the visible warning for a missing explicit tirith_path or
a finished download that failed.
The heartbeat that reports long-running command progress swallowed any
activity-callback exception with a bare pass, so a broken callback left
no trace. Log it at debug with exc_info like the sibling best-effort
handlers in the same module; the command still never fails because of
a progress hiccup.
utf-8-sig exists to tolerate BOMs that Windows tooling adds to files
users edit. /proc and /sys files are generated by the Linux kernel, never
BOM'd and absent on Windows, so -sig there only muddies the read/write
policy. Switch every literal /proc/ and /sys/ read to utf-8 and teach
the footgun read rule that string literals starting with /proc/ or
/sys/ are exempt (user-edited files keep utf-8-sig).
The pre-existed branch of _restore_snapshot unlinked only SKILL.md, so a batch
[create gamma (adopting an empty leftover), write_file gamma references/a.md,
failing op] left references/a.md behind while reporting "all touched skills
rolled back" — and the next create was refused as occupied, wedging the user
until they hand-deleted the dir. The batch already records each applied op's
name/file_path; hand those to the rollback so it unlinks exactly what the batch
wrote and rmdir()s the emptied dirs up to the skill dir. rmdir() fails on
anything else, so a file that landed out-of-band still survives.
One ownership rule for an adopted empty dir on both paths: the single-op create
also always rmdir()s after a blocked scan instead of tracking a created_dir flag
(rmdir removes only an empty dir, so the flag added nothing but a second rule).
create_targets loses its never-used default; _create_skill reuses
mkdir_under_hermes_home for the parent instead of inlining its assert + mkdir.