A bare version request lets uv pick an emulated x86_64 CPython on
Windows arm64 hosts ("support for the native architecture (aarch64) is
not yet mature"). The bootstrap interpreter then ran as win-amd64.
Request cpython-<minor>-windows-<arch>-none from the machine
architecture the scripts already detect, in install.ps1,
setup-hermes.ps1 and setup-hermes.sh (win32 only; POSIX keeps the bare
version so uv still picks the right libc variant).
npm caches the tarball it installs. The cache was written next to the
archive, inside the store's fetch-<sha> entry. Removing that entry after
publication then had to delete a tree, and on Windows it failed while
Defender still held the fresh tarball copy:
[WinError 145] The directory is not empty: ...\fetch-<sha>\.npm-cache\...
Give npm a throwaway temp cache whose cleanup cannot fail the install,
so the download entry holds only the archive.
The current source checker reports updateAvailable with behind null when GitHub
compare cannot count staged commits; the app-update predicate demanded an integer
behind and refused every HEAD->NEXT leg. Require behind > 0 only for the historical
shape without updateAvailable.
v2026.6.19's DMG bootstrap runs the same install.sh that writes .install_method,
so its macOS leg saw a dirty tree. Share the Linux driver's guarded exclude through
source-driver.sh and apply it before every app-driven macOS update.
A venv records its base interpreter path in pyvenv.cfg, not just a version.
The generation identity hashed only the pinned Python artifact, so the same
pin realized in a different tools store (a fresh HERMES_RUNTIME_DIR, a CI
setup-pm store) reused a venv whose interpreter no longer existed or belonged
to another store. Key the identity on the resolved selected interpreter and
build against it explicitly; a failed replacement still keeps the prior
generation selected.
Review of #121499: a slow-but-healthy dashboard install, partial pricing,
a nested permissions schema error, and a refusal that omits the holder pid
no longer XFAIL as the pinned bug.
PowerShell hands child processes a console with
ENABLE_VIRTUAL_TERMINAL_PROCESSING off, so conhost drew the progress
line's ESC[2K as a glyph: every install.ps1 status line printed as
'<-[2K✓ ffmpeg'. The progress stream now opts its console handle in to VT
before use, and a handle that is not a console (NUL reports isatty()
too) gets plain lines instead.
Root cause of the CI-only reds of test_resizes_keep_each_transcript_line_once_in_tmux_scrollback
("'t1w059' never appeared", provider saw 0 requests): the test typed its first question two
seconds after "Welcome to Hermes", but that line is printed before prompt_toolkit's app starts.
On a starved runner the app was not yet up, so the keys landed in the still-cooked tty: the
kernel echoed them (the plain "question zq1q please" row above the status bar), the echo
satisfied wait_for(question), and the later Enter completed the line in the kernel buffer. When
the app took raw mode it read text + Enter in one batch, and the rapid-input guard (#10994)
kept the Enter as a pasted newline: the question sat unsent in a two-line draft, "ctx --".
- Wait for the app-painted status bar before typing, and for the question on the "❯" composer
row (the app's own render, never the tty echo) before the Enter.
- The Windows ConPTY composer test wrote "\r" the instant the echo painted, inside the same
50 ms rapid-input window: give the Enter 0.5 s after the last typed key.
- Keep remain-on-exit so the SIGABRT faulthandler dump of the diagnostics stays readable.
Reproduced locally by delaying CLI startup 3 s after the welcome line: the pre-fix harness
fails with the exact CI screen and main=0; the fixed harness passes. No network is on the
turn path: every outbound call (update check -> api.github.com, tirith install -> GitHub
releases) runs on a daemon thread; black-holing all non-loopback DNS leaves the first turn
at 1.4 s (tmux) / 2.7 s (chat -q).
test_resizes_keep_each_transcript_line_once_in_tmux_scrollback typed the
question, slept a fixed 0.5 s, then sent Enter. The classic CLI treats an
Enter processed within 50 ms of the last buffer change as a pasted newline
(_RAPID_INPUT_ENTER_WINDOW_S, #10994), and that clock is processing time,
not arrival time. On a starved runner the CLI can still be chewing on the
typed batch when the Enter lands in the PTY; it then processes the Enter
right after the last character, inserts a newline, and the question is
never submitted. CI (64-worker e2e run) failed exactly so: the composer
held "question zq1q please" plus a blank row and 't1w059' never appeared.
Repro: a scratch copy of hermes_cli/cli_tui_mixin.py that blocks the event
loop for 1.0 s when the composer's first key is processed. Base fails 4/4
with the CI signature; the fix passes 12/12 (stall hit on all 3 turns).
Fix: wait for the typed question to render (every key processed) before
the 0.5 s gap and the Enter, the same handshake tests/e2e/core/terminal/
_pty.py::submit already uses.
The desktop-installer legs intermittently leave Hermes-Setup.exe on
LAUNCHING after the Launch click: no app window, no desktop.log, and the
installer never exits. Its tracing log is buffered and dies with the job,
so the stall left no trace of where it blocks.
On a driver failure, while the installer is still alive, record the
Hermes/webview/python/uv/git/node process table (was Hermes.exe ever
started, and by whom), the installer's thread states, and a full memory
dump of the installer into proof/driver-failure/. Capture failures are
logged and never replace the driver's own assertion.
Off macOS, ctrl folds to mod, so a Ctrl+B voice default is the sidebar
chord. Ship Ctrl+Alt+V instead, and let a binding be cleared to an empty
combo so sidebar mod+b can be unbound. Point the Voice settings hint at
the voice conversation action, not dictation.
Revealing the terminal restored the layout but left the composer holding
the keyboard. Terminals stay mounted while hidden, so the activation
focus effect never re-runs. Ctrl+backtick, the palette row, and the
statusbar pill now claim focus on the reveal direction and re-assert if
the composer steals it. Hiding does not. No second focus chord.
remote-secondary: a bot on a remote secondary connection (connections.json)
with the local backend as primary; a client-only image must reach the remote
as image.attach_bytes. startRemoteBackend({hide}) covers client folders with
an empty tmpfs in a private mount namespace (unprivileged userns; annotated
'fidelity' when unavailable), used by both remote specs. remote-topology:
the post-restart state.db poll (which passed before the new backend opened
the DB) is replaced by a new-pid check + the restarted backend serving the
title after reload.
lineage-sidebar's #121148/#121088 steps never compacted (/compress was
refused as would_grow on a tiny transcript), so they could not fail.
lineage-rotation drives real auto-compaction with compression.in_place:
false (state.db-verified sealed parent + continuation) and asserts one
sidebar row per lineage after rotation, reload and cold relaunch; the live
nested-branch shape is KNOWN #121148. lineage-compaction-prompt compacts
inside an acknowledged redirect turn so the refresh passes through
preserveLocalPendingTurnMessages; the duplicate prompt is KNOWN #121088.
known.ts: expectNoSymptom marks a test expected-failing at run time only on
the bug's own assertion (a fix merging first stays green). lineage-sidebar
keeps the branch/switch scenario, drops the #121096 claim and now checks the
in-page duplicate sampler.
A second open while the bar is already visible rewrote EMPTY, which
cleared the query, and the focus effect only watched active, so it
never ran again. Bump focusRequest instead, keep the typed query and
the captured scope, and select the existing text on refocus.
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
A manual /compress whose session.compress reply is `status: 'pending'`
(#97948 — the gateway's compute-host wait expired while the host kept
compressing) shows one 8s "still running in the background" toast and
returns. The summary and the success toast are rendered only in the
synchronous result branch that `return` skips, so when the host finishes
the `compacted`/`ready` edge merely clears the spinner: no transcript
line, no completion toast, and hydrateFromStoredSession swaps the
transcript underneath the user with no acknowledgement it ever ran.
Claim the session on the pending reply and announce the completion on the
terminal edge, reusing the notice id the handler already notified under so
the pending toast is replaced in place rather than stacked. The transcript
line is appended only after the hydrate resolves — it replaces the
transcript wholesale and would otherwise drop the line. The claim is
consumed once, so a later auto-compaction cannot replay the notice and an
auto-compaction the user never asked for stays silent.
implicitSlashAcceptIndex() returned null immediately when no command name
had been typed yet (!typed), so the activeExplicit branch was never
reached. Pressing Enter after arrow-key selection on a bare '/' sent a
bare '/' instead of the highlighted command (#98535).
Fix: move the activeExplicit guard above the !typed early-return. A
deliberately arrowed highlight always wins regardless of query content —
'Enter means I want this one' should hold even before any characters are
typed.
Regression tests added:
- bare '/' + arrow-key → returns the highlighted index
- bare '/' + no arrow-key → still returns null (auto-accept not engaged)
CI (run 35996862679) failed test_catalog_matrix_0.py::test_provider_row[deepinfra] with rc -9
"killed 8.0s after CONNECT api.deepinfra.com:443 with no inference at the fake"; it passed locally.
Root cause is the harness, not the product and not a cache leak:
- The child homes are identical locally and on CI: no models_dev_cache.json anywhere, models.dev
CONNECT refused by the sentinel in both places (checked the probe home + egress timeline).
- deepinfra CONNECTs its own vendor host once during startup (catalog fetch, refused at once and
neg-cached), then does ~2 s of CPU-bound init (imports, config load, scratch prune over
psutil.process_iter, system prompt, title write) before its first POST to the fake. On the
64-way CI runner that gap stretched past the 8 s grace and the watchdog killed a turn that
would have succeeded. Scaled repro: grace 1.5 s locally reproduces the exact CI red (same six
cells, same five local-probe GETs, rc -9).
- The watchdog was added to bound the auto-recovery ladder, which write_home already disables
(agent.auto_recovery_cycles: 0): xai now ends on its own (7 refused CONNECTs, rc 2, ~18 s) and
its KNOWN signature (fake_inference=0 + api.x.ai CONNECT) is unchanged.
run_hermes keeps only the TURN_TIMEOUT hard bound (proc.wait(timeout=...)); the fallback file
drops the same watchdog arguments (deepinfra as a fallback had the identical latent race).
Unused imports in _catalog_helpers.py removed.
Catalog suite (6 files): 3x serial + 2 parallel copies, XDG caches pointed at an empty scratch
dir: 125 passed / 9 skipped / 0 failed every run, xfails unchanged (matrix_0 1, oauth 2,
listing 37, fallback 2).
Merge-order check with #121398 (picker honours base_url) showed the stricter cell turned red on a
correct fix: minimax lists at <base>/models under the /anthropic override, and deepinfra /
commandcode-anthropic post-filter the relay's generic ids. The KNOWN-gated cell now asserts only
the configured endpoint at an exact listing path and no vendor host; rendering the relay's own ids
is a separate cell on the rows that already list from the configured endpoint. Also encoding= on
the OAuth auth.json read/write (windows footgun).
Review fixes on the provider-catalog matrix:
- KNOWN bugs go through tests.e2e.core._pending_fixes.known_failure again (strict_known and the
"now green -> drop KNOWN" asserts are gone), so a fix PR landing first simply turns its cells
green. Each KNOWN is gated on the bug's OWN observed signature via a dedicated CatalogGap raised
only at the gated assertion: xai = no inference at the fake + api.x.ai CONNECT; listing = vendor
host hit with no probe error; nebius switch = vendor-host validation error; fallback = fallback
fake untouched + its vendor host CONNECTed; OAuth patterns anchored. Unlisted red cells still fail.
- CatalogFake answers only exact routes (configured base path + dialect endpoint / listing path),
404 otherwise; reached_own_endpoint and the listing cells assert the exact path.
- Turns run with agent.auto_recovery_cycles: 0 (the documented post-exhaustion ladder parked the
xai/fallback rows for minutes: 14 CONNECTs over 118 s, bounded by design, not a retry bug) and a
watchdog kills a child 8 s after a vendor-host CONNECT with no inference at its fake.
- An unknown api_mode fails the row instead of skipping; only explicit auth types and the named
keyless provider skip.
- Usage cell asserts the exact sum the fake reported for answered main-turn calls.
- Listing 404 / hang degradation cells for the rows that already list from the configured endpoint.
- Shard body moved into the helper; fallback key check uses the dialect's auth header; the three
unconditional-skip OAuth params dropped (kept under NOT COVERED).
Rows come from real plugin discovery (providers.list_providers in a child), sharded by
name hash into 3 files. Strict KNOWN entries: #121347 (xai base_url), #121359 (fallback
base_url for anthropic/openrouter), #121387 (picker listing ignores base_url), #121388
(nebius switch validation). OAuth cells switch to a strict message-gated helper.