When a compute-host turn has not settled by session close, the session's
real active-session lease is parked in _deferred_active_session_leases and
the turn's completion callback is supposed to release it. If that callback
is lost (supervisor restart/reload, child killed without failing its
pending turns, dropped completion), the lease sat in the registry forever:
_own_live_lease_ids vouches for it so the orphan sweep never reclaims it,
and the concurrent-session cap treats the dead session as active - new
sessions could not send until a backend restart.
Track when each lease was deferred and have the reaper tick force-release
deferred leases past a generous 30-minute ceiling (the longest legitimate
isolated turn is the compression ceiling, minutes). Normal settlement
still releases immediately and clears the age entry.
Fixes#62823 (zombie-concurrency-slot half; cross-window queue-sync was
fixed by #122953 and visual-merge by #70986 - not touched here)
A correction delivered to the provider mid-compression aborts the
compression (explicit_interrupt) - the user's follow-up kills the very
turn that would have answered it. The channel-side busy path already
demotes interrupt->queue for this exact reason (gateway/run_busy.py,
#56391); mirror that demotion for the local RPC path so TUI, Desktop and
classic chat share the Discord-gateway contract: prompt.submit busy
steer/interrupt and session.steer/session.redirect queue instead, and
the follow-up drains when compression finishes.
Fixes#61042
Accounts Connected now follows token validity instead of access-token
presence. Windows removal uses unambiguous PowerShell, a clear that
removes nothing is an error instead of a success toast, and the
connected-row terminal control runs disconnect.
Mapping the llamacpp aliases to custom in hermes_cli.models sent the
managed local runtime's /model validation down the custom branch before
the staged-library check, so a downloaded-but-not-running GGUF lost its
recognized verdict and an unstaged one was accepted. Keep local and vllm
on custom there, leave the llamacpp aliases on their runtime id, and
move the orphaned 'local' display label to 'custom'.
The parity test now reads the aliases from the custom provider profile
instead of a hand-written list.
Local OpenAI-compatible server aliases (local, vllm, llamacpp, llama.cpp,
llama-cpp) were normalized inconsistently across the three provider-name
tables: hermes_cli.providers mapped them to the orphan "local" id,
hermes_cli.models left them unmapped, while hermes_cli.auth mapped them to
"custom". Align all three on the generic "custom" provider so routing,
the model picker, and credential resolution agree, matching what
resolve_provider already did for vllm/llamacpp. Add a parity contract test.
Background MCP discovery does bursty CPU work (SDK/pydantic imports, JSON-RPC
schema parsing, tool registration) that at the default 5ms switch interval rides
the GIL convoy effect and starves the concurrent agent-build thread and the main
event loop for tens of seconds after a serve restart — _wait_agent(30s) times
out with 'agent initialization timed out' (error 5032).
Lower the interpreter switch interval while discovery runs (restored afterwards)
so its CPU bursts are sliced finely enough that waiters stay responsive.
Fixes#60371
[ -f ] and [ -e ] follow symlinks, so a dangling link probed as missing and
read_file_raw reported not_found: V4A Add followed the link and created its
target, Move replaced the link, both reporting success. The shell size
probe, the compound read probe and the native stat now classify any
symlink entry as not a regular file.
The fence drops backend output around the read, not output printed while
the transport runs: a BASH_ENV DEBUG hook firing for base64 alone puts
text inside the payload that still decodes (TERM is b"LDL"), and the edit
paths wrote it back. Emit wc -c in its own fenced segment and hand bytes to
a writer only when their length matches; the od fallback is checked the
same way.
Consolidates #120559 by @JoaoMarcos44 into this PR. Both PRs landed on the same
shape for #120514 — read_file_raw must be a byte-preserving mutation boundary,
framed so the backend's merged stdout is never decoded as payload — but the
sibling one step in FRONT of it was still unfenced here.
_sample_file_bytes whitespace-joined the whole reply before base64-decoding it,
so a backend that announces something on connect had that noise decoded into the
sample: "TERM" is four base64 characters and prefixes the sample with b"LDL".
That sample is the binary-admission gate read_file_raw consults, so noise there
decides whether a file is editable at all and what a refusal reports about its
bytes. Same class as the read this PR already fixed, one call earlier.
_fenced_read() is that framing extracted once — sentinel, payload segment, and
the body's own exit status in its own trailing segment — and the three byte
transports (sample, base64, od fallback) now share it instead of repeating the
split/parse/decode triplet, with _failed_read() for the diagnostic they all
surfaced by hand.
The mock double for the sample transport moves to the fenced reply shape the
other doubles in that file already build, so it stays honest about the fence.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Validation treated every read_file_raw error as "the path is free" for an
Add target and a Move destination; only _apply_add checked not_found. With no
byte transport (native reads off, no base64 or od) an `Add new-src` then
`Move new-src -> dst` validated, and the mv replaced the existing dst while
the patch reported success.
_occupied() frees a path only on not_found (overlay-aware), and _apply_move
re-checks the destination right before mv, since validation's answer is stale
once earlier ops of the patch have applied.
Three defects in the byte-exact read this PR introduced.
base64 is not on every backend (busybox, distroless). The sample path
already degrades when it is missing; this one returned the read as a
failure, and read_file_raw is also _apply_add's existence check, which
treated any error as "the path is free". A backend with cat but no base64
turned `*** Add File` over an existing file into a silent overwrite that
reported success:
main: refused, "file already exists — use Update File"
PR head: b'KEEP ME -- months of work\n' -> b'clobbered by Add'
So: fall back to od (POSIX, in busybox) when base64 exits 127, and when
neither exists report a transport error. ReadResult grows not_found, set
only where the path is genuinely absent, and _apply_add refuses unless it
sees that flag — a read that FAILED can no longer pass as an absent path.
getattr keeps a producer without the field failing closed rather than
raising. The doubles in test_patch_parser that meant "absent" now say so.
The native fast path stat'd the path and then opened it, two lookups on a
name. A swap to a FIFO in between blocks the thread, and nothing times out
that. One O_RDONLY|O_NONBLOCK open, fstat on THAT descriptor, then read, so
a non-regular file is handed to the shell path and its timeout instead.
Tests: the od fallback round-trips byte-exactly and patches; an Add over a
file the backend cannot read is refused with the file intact; the native
read hands a FIFO to the shell rather than opening it. The first two go red
if their fix is removed. The third pins the property, not the race — with
one descriptor there is no window left to swap into, so reverting to
stat-then-open does not flip it.
Not addressed here: inserted text is still encoded UTF-8 regardless of the
file's declared encoding, so an edit adding non-ASCII to a latin-1 file
writes mixed bytes. That is the write half, it predates this PR, and it
needs source-encoding detection.
The byte-exact read whitespace-joined the whole reply before decoding, so
anything a backend prepends to its merged stdout became part of the base64.
A remote shell announcing TERM is four base64 characters, which decodes to
b"LDL" and lands at the head of the file the edit paths then write back:
patching HEADER\nVERSION=1\n left LDLHEADER\nVERSION=2\n on disk.
Fence the payload between a per-call sentinel, the shape the compound read
probe already uses and for the same reason. _new_sentinel's underscores are
outside the base64 alphabet, so noise that lands inside the fence fails
validation instead of decoding into bytes, and noise outside it is dropped.
The read's exit status rides in its own trailing segment, so a failed read
is still told apart from an empty file without && chaining or exit.
stderr stays merged rather than silenced: on a failed read that segment
carries the backend's own diagnostic ("No such file or directory"), which
the callers surface, and a reply with no fence at all hands the backend's
text straight back so a wrapper cd failure still reports itself.
The two doubles this PR added keyed off the old command string; they now
build a fenced reply from the sentinel in the command they were handed, so
they stay honest about the fence the transport depends on.
The new test drives the noisy backend through the real transport: the
existing tests use LocalEnvironment, which takes the native read path on
POSIX and never exercises base64 at all.
Review follow-up (ehz0ah, teknium1). Keying the cleanup on the
__HERMES_FENCE_ marker still rewrote a real line that contains that
text, and nothing has emitted the wrapper since the spawn-per-call layer
(d684d7ee7e), so the cleanup can only ever eat the file's own bytes.
The same class also hit mode="replace", the default: patch_replace (and
V4A past the 1000-byte binary sample) read the file through the text
transport, which decodes with errors="replace", so any byte UTF-8 cannot
decode came back as U+FFFD and was persisted on lines the edit never
touched, while the diff and the post-write check (same lossy read) showed
nothing.
_read_exact_bytes reads natively on the local POSIX host (regular files
only) and over base64 elsewhere; read_file_raw, patch_replace and
_verify_patch_persisted decode it with surrogateescape, which write_file's
encode already inverts (#79178), so every untouched byte is written back
exactly and no readable file becomes a read error (V4A's Add/Move/Delete
existence checks are unchanged). A garbled transport reply refuses
instead of writing stray output into the file. The Python linter parses
bytes so a declared legacy coding still lints clean. Local edits spawn
two fewer shells (replace 5 -> 3, V4A 9 -> 7).
read_file_raw feeds the V4A write-back, and it ran the whole file through the
terminal fence-leak cleanup, which deletes every OSC sequence and BEL byte. A
prompt script's title escape or a beep line was silently rewritten on apply, and
the diff and post-write hash both started from the stripped text, so nothing
showed it. A leaked wrapper always carries the fence marker: clean only those
lines when the content is going to be written back.
The four CI reds from #41225:
- persist_on_release is a background-only terminal modifier: join the
sandbox blocked sets (precedent: heartbeat, 9acd0d33b6) in
_TERMINAL_BLOCKED_PARAMS and the stub-drift tests' mirrors.
- The gateway shutdown sweep passes source="gateway_shutdown" so
persisted jobs are still killed on host exit; the two shutdown tests'
kill_all fakes now accept and assert that kwarg instead of raising
TypeError that _quiet_step silently swallowed.
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.
Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
(_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
explicit operator stops (process_manage kill, /stop slash + RPC mirror,
CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
the host is going away and survivors would become PPID=1 orphans
Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
On Windows, subprocess text=True without an explicit encoding decodes
child output with the ANSI code page (e.g. 'gbk'); non-ASCII bytes then
raise UnicodeDecodeError inside subprocess._readerthread, killing the
Hermes backend before it becomes ready and surfacing as the desktop boot
timeout.
Sweep every hermes_cli text=True subprocess call to encoding='utf-8',
errors='replace', and add an AST-based regression test that fails when a
future text-mode call omits the encoding.
Fixes#55658
* fix(update): stop reading gateway identity off the restart watcher's argv (#107002)
The detached restart watcher is spawned as
`python -c <watcher source> <old_pid> <python> -m hermes_cli.main gateway run`.
Its trailing argv is the command it will spawn LATER, but the canonical matchers
read identity straight off the joined command line, so the watcher itself was
classified as a live `gateway run` process — the documented "never infer process
identity from argv substrings" bug class, on the exact surface `hermes update`
uses to verify a post-update relaunch.
Also budget the post-relaunch liveness poll against the watcher's own deadline:
the watcher respawns the gateway only after the PID it was handed exits, so a
30 s window can expire before the relaunch it verifies was scheduled to start.
* test(windows-live): run the gateway-ancestor harness parent from a script file
A `python -c <src>` parent is an interpreter running inline source and carries
no readable Hermes identity, so it is no longer a gateway to any classifier —
the harness's own comment already said a realistic gateway argv is not a -c blob.
* test(windows-live): share one sleeper SCRIPT across the live process-topology fixtures
Four live Windows E2E files stood processes up as `python -c "sleep" <hermes argv tail>`.
That shape no longer carries a readable Hermes identity, so the fixtures stopped standing
in for the gateways they simulate. One shared sleeper script replaces the -c spelling.
* fix(gateway): drop the duplicated _INLINE_SOURCE_FLAG_RE definition
The constant was emitted twice around command_line_runs_inline_source. Same
pattern both times, so behaviour is unchanged — but one definition is enough.
* test(stderr-timestamp): run the gateway-lookalike children from script files
Both lookalikes stood a gateway child up as `python -c <src> <gateway tail>`.
That shape no longer carries a readable Hermes identity (#107002), so the
wrapper correctly stopped treating them as gateway spawns and the tests failed.
A `-c` tail is data for a program the inline source may spawn LATER, never the
child's own identity — the real wrapper child is `python -m hermes_cli.main
gateway run`, which has no `-c`. Running the stand-ins from a real script file
restores what the tests mean to assert without depending on the misread.
* test(windows-live): restore the tempfile import dropped with the local sleeper helper
* test(windows-live): wait on the sleeper SCRIPT name, not its source text
The live fixtures proved argv visibility by waiting for `time.sleep(120)` in the
spawned process's command line. That string only ever appeared there because the
sleeper was spelled `python -c "import time; time.sleep(120)"`; now that it runs
from a file the source is in the file, so the probe timed out ("sleeper argv
never visible") even though the argv was perfectly visible.
Wait on the script name instead, exported as SLEEPER_MARKER next to the script
so the probe and the spelling cannot drift apart again.
* fix(tests,gateway): keep the live-system guard blocking -c-wrapped gateway spawns
The #107002 identity fix made _gateway_command_subcommand return None for
'python -c <src> … -m hermes_cli.main gateway run'. tests/_fixtures/live_system_guard.py
shares that matcher, so the autouse guard stopped blocking the detached restart
watcher: real gateways leaked out of the e2e run and squatted the webhook port.
Add gateway.status.gateway_spawn_intent_subcommand — the spawn-intent mirror of the
identity matcher, peeling the inline-source wrapper token-wise and re-running the same
canonical matcher on each suffix (still no substring matching) — and point the guard at
it. Read-only subcommands stay spawnable.
* fix(gateway): make the inline-source option walk value-aware so -X utf8 -c is not read as a gateway
The walk that decides whether a command line is an interpreter running inline
source (`python -c <src> ...`) treated every token starting with `-` as a flag
and the first non-flag token as the end of the option block. CPython options
that take a SEPARATE operand (-X/-W/-Q, --check-hash-based-pycs, --jit) break
that model: the operand was mistaken for the end of the block, so the walk
never reached the -c behind it and the watcher was read as a live gateway
again -- exactly the #107002 misclassification, one shape further out.
- Reuse the canonical operand sets from hermes_state_holders rather than
hand-rolling a second copy (AGENTS.md: parser-derived flag sets).
- Walk case-preserving tokens: operand-taking -Q/-W/-X must not be conflated
with operand-less -q/-b, so callers no longer lowercase before the walk.
- Handle clustered short options precisely (-uc is inline source, -Xc is -X c).
- Replace the ad-hoc _INLINE_SOURCE_FLAG_RE rescan in
gateway_spawn_intent_subcommand with the index the same walk returns; the
regex could not find spellings the walk accepts and would raise
StopIteration.
Reported by an automated review on PR #121635 and reproduced here.
Refs #107002
---------
Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
The removed heredoc row was the only case exercising the old "+1 when
heredoc" allowance. Keep the row with ok=False so the probe is proven to
reject exactly the drift the allowance used to accept, now that
_embed_stdin_heredoc delivers stdin byte-exact.
A heredoc body always ends in a newline, so on the heredoc-stdin backends
(Modal, Daytona, Vercel) every stdin payload arrived with one extra byte.
write_file and patch verify an exact sha256 of the written file, so even
with the command grouped every write failed verification and left
content + "\n" on disk.
Feed the heredoc through a process substitution that re-emits the body
minus that last character. The heredoc stays outside <( ) because bash
3.2 mis-parses heredoc bodies inside it, and the reader tolerates read's
EOF status under an inherited set -e.
The tool-result spill no longer needs its +1 byte allowance for heredoc
mode, so the size check is exact on every backend again.
When a desktop build fails because esbuild's platform binary
(@esbuild/<platform>) was never staged - typically because
ignore-scripts=true skips esbuild's postinstall - print an actionable
diagnosis with the fix instead of leaving only the raw build error.
Fixes#53082
PR #49037 (projects paradigm) broke the folder->session->sidebar flow (#53004):
- The right sidebar's Files pane gated the tree on $currentCwd, which is ''
for global/detached sessions, so those sessions hit a dead-end 'No project
open' pane with no way back into a folder. Restore an affordance in the
empty state: 'Open folder' runs the existing open-folder-as-project flow
(upsert + enter project + fresh session anchored at the picked folder),
decoupled from $currentCwd.
- The sidebar overview listed every auto-promoted repo on disk, session or
not. Auto projects with zero sessions now stay out of the sidebar until
they own a session (they reappear the moment work lands there); explicit
projects and the Home bucket always render.
UI-affecting: coordinator should hold auto-merge for visual review.
Fixes#53004
NVIDIA driver 580+ breaks ANGLE's EGL probing (Invalid visual ID requested),
killing the GPU process at startup on Ubuntu 24.04 and similar hosts. Detect
the driver major from /proc/driver/nvidia/version and, on Linux (not WSLg,
not a remote display), pre-launch appendSwitch('use-angle', 'swiftshader').
Deliberately avoids app.disableHardwareAcceleration(): on 580.173.02 +
Electron 40 that path SIGKILLs the renderer, so the closed#40119 approach
is unsafe here.
Overrides: HERMES_DESKTOP_NVIDIA_SWIFTSHADER=1 forces the fallback on,
=0 opts out; HERMES_DESKTOP_DISABLE_GPU=0 keeps the GPU untouched.
Fixes#40077
Every desktop window boots $queuedPromptsBySession from the same
localStorage key and then never syncs again: no storage-event listener,
and every save writes the window's whole snapshot back, so windows
clobber and resurrect each other's queued prompts (#46732).
- listen for storage events on the queue key (and key===null full clear)
and reload the atom; the event only fires in non-writer windows, so
there is no self-echo
- writeSession/migrateQueuedPrompts merge over the live persisted map
instead of the in-memory atom, closing the same-frame race where a
save reverts a write that landed between sync events
The session-switch half of #46194 (drain routing to the wrong session)
is already covered on main by the per-session queue keys, the
isBackgroundQueueDrain guard and the stored/runtime binding verification
in submit.ts; this PR removes the remaining shared-storage leak in the
same state cluster.
Based on #57516 by @furancis (closed unmerged) — the storage-listener and
merge-over-live-storage approach is reused and reworked onto current main.
Fixes#46732Fixes#46194
project_create (desktop_project tool / CLI) persists to the per-profile
projects.db but writes nothing to state.db, so sessions.changed never fires
and the desktop Projects sidebar goes stale until a manual refresh.
The gateway change watcher now watches projects.db and broadcasts
projects.changed when it moves; the desktop live-sync routes the event to a
new $projectsChangeTick that refetches the project list + tree.
Fixes#56757
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:
- batch_runner._process_single_prompt: one agent per prompt, N prompts
per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.
Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.
Fixes#50197
uv treats `foo_bar` and `foo-bar` as the same extra (PEP 685), so an
exact comparison would drop a recorded extra whose spelling differs from
its declaration. Compare normalized names; the recorded spelling is still
what reaches uv.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The venv ledger unions its recorded extras into every sync. When a source
update removes an extra (hindsight, 73c598e319/285768fdfb), installs that
recorded it pass `--extra hindsight` to `uv sync` forever and fail with
"Extra `hindsight` is not defined in any project's optional-dependencies
table". Every CLI launch then retries the failed source-update completion,
and gateways keep booting the previous generation's code.
Drop recorded extras the checkout's pyproject no longer declares, in both
the sync target and the currency probe. An explicitly requested unknown
extra still fails loudly.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A turn that falls out of the loop after a tool result with no follow-up
assistant text left the durable transcript ending at a raw tool row and
returned a silent result: Desktop/TUI showed a ready composer (or kept
spinning) with no final message. The pending_tool_result explainer copy
already existed but nothing minted the exit reason.
finalize_turn now detects the non-interrupted tool tail, fails the turn
with turn_exit_reason=pending_tool_result, and synthesizes the visible
assistant close before persistence so the durable tail is alternation-
safe. Stream-recovered turns (#95514) and interrupted tails keep their
existing paths.
Fixes#55316Fixes#54756
Co-authored-by: blakehermes9 <blakehermes9@users.noreply.github.com>
Detect the host sleep by the divergence between the wall clock and the
monotonic clock since arming (the monotonic clock does not advance
while asleep), instead of an absolute monotonic deadline — a timer
that fires without elapsed wait time (tests, spurious dispatch) shows
zero divergence and reaps normally. Re-arm through
_schedule_ws_orphan_reap with the fired timer as the expected one so
the fresh closure's identity guard matches the entry it installs.
threading.Timer's wait elapses in wall-clock time on platforms without
a monotonic condvar (macOS lacks pthread_condattr_setclock), so a system
sleep made the WS-orphan reap timer fire at the instant of wake — before
the Desktop app's WS reconnect or session.resume could re-bind a
transport — and every sleep/wake cycle longer than the 20s grace 404'd
the open session. Record a monotonic deadline when arming the reap and
re-arm for the remaining awake time when the wall-clock timer fires
early; a reconnect still cancels the chain as before.
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
read_preview followed the global right-rail tab. follow() copied the
interacted zone into that id, so one group overwrote an explicit open
in another. Resolve from the hovered or focused zone, and skip that
copy when the explicit open lives in a different group. When more
than one preview is mounted, include active_tab_id and the open tabs.