The admission gate read `db.db_path`, which made the refusal depend on whatever object
`SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with
AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from
`_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the
scan targets the same file in production and stays reachable regardless of the db object.
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
Review follow-ups on the new `hermes sessions set-journal-mode` verb:
- The header probe used os.pread, which does not exist on Windows, while the subparser is
registered unconditionally — the command died there with an uncaught AttributeError. It now
reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
(apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
bail in the command's own error style; every open/read is guarded.
The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.
The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).
Co-authored-by: webtecnica <webtecnica@gmail.com>
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
EmptyCompletionError is raised by complete_task; without a handler the
CLI and the tool would surface a traceback instead of the actionable
refusal (add --result/--summary). Same shape as the HallucinatedCards gate.
Part of #117483
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace. Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:
- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
also covers everything beneath it); a matching root never spawns, logged once at INFO. A
non-list value fails closed — WARNING naming the expected shape, every root skipped — because
silently excluding nothing would re-pay the stall the key was meant to avoid.
Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
The catch-up rewrote the receipt's fleet matrix before checking whether that
actually cleared the debt, so an exit-1 run (marker still owes a gateway,
blocked serve_restart_pending storage) left latest.json mutated —
tests/hermes_cli/test_update_scoped_reconciliation.py pins the receipt as
byte-identical on the failure path. Decide on the in-memory settled copy and
write only when the pending restart is discharged (#117051).
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).
In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.
Fixes#117051
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").
Part of #117051
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].
Fixes#117246
`list_profiles()` parsed three YAML files per profile on every call — config.yaml through
`read_user_config_raw` ("no cache", by contract) plus profile.yaml and distribution.yaml through
`_load_yaml_dict`. It is the shared body of `GET /api/profiles` and `profiles.list`, which the Bots
roster polls every 5s per connection, so an idle machine re-parsed every profile's YAML twelve
times a minute to produce the same handful of strings.
config.yaml is the expensive one: the installer seeds it by copying the annotated template verbatim
(`cp cli-config.yaml.example config.yaml`), 119,268 bytes / 2,267 lines of mostly comments, for
`model.default` and `model.provider`.
Measured, 9 profiles, no gateway locks so the liveness path stays out of it: 16.1ms -> 1.6ms with
the installer template, 6.4ms -> 1.6ms hand-trimmed; 27 YAML parses per pass -> 2, and those two
are the default profile's ABSENT files, correctly not cached. The pass is now independent of config
size.
The memo keys on each file's (mtime_ns, size, inode) — the atomic writers rename a temp file into
place, so a rewrite always lands a new inode even inside one mtime tick — and caches only the small
DERIVED values. `read_user_config_raw` and `_load_yaml_dict` keep their uncached contract: callers
write those documents back, and a stale read there could overwrite a newer file. `read_profile_meta`
hands out a copy so a future mutating caller cannot poison the next reader.
Regressions drive the real readers: an edited config/profile/distribution file is seen on the very
next read, `write_profile_meta` round-trips, a missing file is not cached and is picked up when
created, mutating a returned meta does not poison the next reader, an unchanged file is parsed once
across three reads, and `read_user_config_raw` still re-reads. Ignoring the signature fails four.
Fixes#117378
request_review recorded trow["assignee"] as the implementer on the
review_requested event and on the synthesized run. That is correct while a
worker holds the card, but wrong when the card was created already assigned to
its reviewer -- `kanban create --assignee <reviewer>` followed by
`request-review`. Both provenance records then name the reviewer, and
request_changes routes a rejection back to the profile that wrote the findings.
Derive the implementer after the reviewer is resolved, from the active run's
profile, falling back to the assignee only when it is not the incoming
reviewer. When no actor can be established the field is left unset, which
request_changes already refuses on rather than misrouting.
Refs #117229. Follow-up to #111064 / #111459, which fixed the run attribution
for the distinct-reviewer case and left the derivation itself unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Mechanical, behaviour-neutral extraction along the existing cli_*_mixin.py
pattern: the ten _init_* constructor phases become CLIInitMixin
(hermes_cli/cli_init_mixin.py) and the fourteen _tui_* run-loop phases
(input dispatch, after-turn, startup banner/prewarm/maintenance, application
build, signal handlers, shutdown) become CLITuiRuntimeMixin
(hermes_cli/cli_tui_runtime_mixin.py). __init__ and run() stay in cli.py as
the orchestrators. Every moved body is AST-identical to the base copy; cli.py
module names are resolved lazily through cli so monkeypatch seams survive.
cli.py 4835 -> 3960 lines.
Mechanical extraction of the service-definition cluster (generate_systemd_unit,
systemd_unit_is_current, refresh_systemd_unit_if_needed and their
_systemd_*/_service_*/_ld_library_path/_temp_home helpers) into
hermes_cli/gateway_service_unit.py. Behaviour-neutral: moved bodies read facade
helpers (get_python_path, _profile_arg, _run_systemctl, load_gateway_config, ...)
and each other through _gw() — a late import of hermes_cli.gateway — so every
seam tests and callers patch on the facade keeps intercepting the moved code;
the facade keeps one import block re-binding the names.
hermes_cli/gateway.py: 6808 -> 6496 lines.
Part of #116913
register_plugin_provider setdefault-ed aliases regardless of provenance, so a
$HERMES_HOME plugin alias colliding with an existing row kept resolving to the
old row while providers.get_provider_profile already followed the user profile;
override_registry_row left the display name untouched on a same-name
replacement. mirror_aliases applies the provider_source ownership rule that
e5e7fbcd27 introduced for endpoints.
Fixes#116668
`_suppress_mouse_residue_early()` runs at import time, before
`_apply_profile_override()` sets HERMES_HOME, so the first read of
`display.interface` always saw the default home. The result was memoised
in a cache keyed on nothing, so every later caller — the Termux fast
paths and the TUI launch decision, which all run after the profile is
applied — was pinned to the default home's interface for the whole run.
`hermes -p coder` with `display.interface: tui` in that profile booted
the classic REPL, and the mirror case booted the TUI for a profile that
asked for cli.
Key the cache on the config path it read. The hot path still parses the
YAML once per home, and the first call after the process is re-homed
re-reads the file it should have read.
Fixes#116902
`_hermes_home_for_pid` fell back to the INSPECTING process's `Path.home()`
when the target's environment carried neither HERMES_HOME nor HOME, which
is the normal shape for a systemd/launchd unit with a scrubbed
environment. The dashboard was then attributed to whichever user ran the
command, so `hermes update` and `--stop` read the wrong profile root and
could act on the wrong backend.
A process without HOME resolves its own default home from the password
database, so resolve the target's owner the same way (psutil, then the
/proc owner) before falling back. An unreadable owner or entry keeps the
previous behaviour rather than resolving to nothing.
Fixes#116906
The 5-branch if/elif ladder in CLICommandsMixin._handle_browser_command
becomes a _BROWSER_SUBCOMMANDS word -> handler table; adding a subcommand
is one row plus a usage line. Subcommand word is case-insensitive, the
argument keeps its case (CDP URL paths are case-sensitive).
Drop the duplicated _GRANT_DEAD_CODES frozenset and the relogin_required
plumbing: the pool's plugin recovery already treats exc.code in
hermes_cli.auth._OAUTH_GRANT_DEAD_CODES as terminal, so _token_http_error
imports that set late and just sets code=<json error>.
A spent refresh token returned HTTP 400 invalid_grant, and the helper mapped that to oauth_refresh_failed so the pool EXHAUSTED the row. Parse the JSON error and DEAD it. hermes auth add with an alias wrote the pool under the typed name, so status against the profile name looked logged out.
fallback.min_switch_reset_seconds is the only schema-surfaced fallback field; without the merge the dashboard grew a one-field orphan tab (test_no_single_field_categories red on CI).
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
(production entry) on a real AIAgent with a one-entry chain, so dropping the
reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
rate-limited primary's declared reset is sooner than N seconds,
try_activate_fallback returns False and no cooldown is armed; docs row added.
Review follow-up for #116905: bot_mode_probe was the third private copy of
the regex (profiles, update_restart_recovery) with a sync test policing the
drift the copy created. One canonical object next to named_profile_is_live;
the three modules import it and the sync test is gone.
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
Review follow-up: get_managed_system, the container/chmod-skip check and
the HERMES_UID/GID chown (_resolve_hermes_uid_gid/_chown_to_hermes_uid) now
live only in hermes_constants — the import-safe module apply_secure_dir_policy
already needed them in — and hermes_cli.config re-exports them, so there are
no 'keep in sync' copies and _secure_file skips on the same canonical
_detect_container signal as _secure_dir/get_scratch_dir. The dead config
copies are deleted and test_ensure_hermes_home_uid.py drives the new
symbols; one invariant test pins the single implementation.
get_scratch_dir unconditionally chmod'ed <home>/cache/scratch to 0700,
bypassing the policy _secure_dir() applies everywhere else: managed
installs are left alone, containers only apply an explicit
HERMES_HOME_MODE, and HERMES_UID/HERMES_GID ownership is honored. On a
group-shared home (setgid 2770) that stripped the setgid bit and broke
the shared-group inheritance the operator configured (#77579, #117347).
Move the policy primitive to hermes_constants.apply_secure_dir_policy()
(import-safe, stdlib-only) and route both get_scratch_dir() and
hermes_cli.config._secure_dir() through it, so one implementation
serves all callers including the cron.jobs delegation chain.
Fixes#117347
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by
cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback
logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482
sibling: 'Primary auth failed' for a 429 on the gateway surface).
- Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the
flag is only True when _ensure_runtime_credentials returned False).
- Fold the new tests: 2 parametrized production-path tests + 1 gateway test.
A Codex quota wall at startup was labeled "Primary auth failed" and
exited 1, so Kanban counted a failure and operators went looking for
a bad token. Use is_rate_limited_auth_error() for the fallback notice
and EX_TEMPFAIL (75) when HERMES_KANBAN_TASK is set.
Fold _profile_owns_catalog and _profile_owned_catalog into _profile_catalog ->
(catalog, authoritative) so _validate_live_listing has a single owned-catalog
branch; drop the redundant catalog-hit test (covered by the owns-catalog accept
tests in test_setup_provider_catalog.py).
The doctor "Browser tools (agent-browser)" row audits the root workspace
tree (npm audit --workspaces=false), whose versions come from the committed
package-lock.json. hermes update reinstalls that exact state via npm ci
(_run_npm_install_deterministic never mutates the lockfile), so the
previously prescribed local `npm audit fix --workspaces=false` was wiped on
the next update and the finding returned - a fix/reintroduce loop (#116774).
- Bump the vulnerable root lockfile pins so the shipped tree audits clean:
browserslist 4.28.6 -> 4.29.0 (high: GHSA-c83g-rgw3-j3cx,
GHSA-73wf-gq98-2v4g), baseline-browser-mapping 2.10.43 -> 2.11.25
(moderate: GHSA-w5vr-8v7q-w6rv), plus updated transitive deps
(caniuse-lite, electron-to-chromium, node-releases,
update-browserslist-db); also syncs the stale bootstrap-installer
lockfile entry to its manifest version. npm audit --workspaces=false now
reports 0 vulnerabilities.
- Rework the npm-audit remedy for every doctor row: the durable fix is an
upstream lockfile bump; never prescribe a local mutating fix command
(workspace-scoped rows already omitted it; the root row prescribed the
doomed one).
- Invariant tests in tests/hermes_cli/test_doctor_audit_remedy.py: the
remedy must not prescribe a mutating npm audit fix (proven red on base)
and must name the lockfile-bump remedy.
A pin is validated as 40 hex, but that does not make it a commit: a tag
object has a sha of its own, and an entry recorded from 'git rev-parse <tag>'
names the tag object rather than the commit it points at. Git resolves the
ref and detaches at the commit, so the revision guard compared a commit
against a tag-object sha and refused a CORRECT checkout — the entry was
uninstallable, with no way through (see the kiro-acp catalog entry).
Peel the pin to its commit before comparing. The guard keeps its teeth: HEAD
must still be exactly the commit the pin names, and a checkout that lands
anywhere else is still rejected. An unresolvable pin falls back to the raw
string, so a genuinely wrong sha fails exactly as before.
Tests: an annotated-tag pin installs at its commit (and records the commit,
not the tag object, so 'plugins update' drift checks stay exact); a
checkout that lands on another commit is still rejected. Red on base:
'Checked-out revision "..." does not match requested commit "..."'.
(cherry picked from commit 66baa043b70c4f66ae59c0a1147e74ba694140ef)
_clear_terminal_on_exit() fell back to os.system('cls'/'clear') when the
escape sequence was rejected. os.system() spawns a shell: on Windows a
conhost window flashes before the process exits, and in minimal POSIX
containers without `clear` on PATH the shell fails silently while the
except-pass hides it. The repo already standardises on windows_hide_flags()
(CREATE_NO_WINDOW) for short-lived argv spawns; use it here too.
- Windows: subprocess.run(['cmd', '/c', 'cls'], creationflags=windows_hide_flags())
(cls is a cmd builtin, so probing a `cls` exe would be wrong).
- POSIX: shutil.which('clear') first; skip the spawn entirely when absent
instead of letting a shell swallow a 127.
- stdin=DEVNULL so a child that reads can never block CLI shutdown.
- Regression tests: parametrised nt/posix argv + creationflags assertions
(RED on the os.system path), plus the no-`clear` skip case.
Review follow-up for #117363: the classify-in-place branch returned before
the expected_run_id guard, so a worker whose run had ended could type a
breaker-parked card. Refuse when expected_run_id is given; the run is over.
Tighten the salvaged classify-in-place path to the reporter's contract:
only block_kind IS NULL and current_run_id IS NULL qualify, recurrence
accounting starts at 1, the audit event is the ordinary `blocked` kind so
existing consumers see the reason, and a typed block is never re-typed.
Drops the unrelated VALID_SORT_ORDERS deletion and trims tests to two
invariants (one drives the real enforce_max_runtime breaker path).
block_task() refuses to set block_kind on a card that is already
status=blocked because the WHERE clause only matches running/ready.
The circuit breaker (enforce_max_runtime) parks cards blocked with
block_kind=NULL, making them permanently unclassifiable.
Add an early-return path: when the card is already blocked and kind
is supplied, update block_kind in-place and append a block_classified
event. Idempotent when the same kind is already set.
Fixes#117363
Review follow-up for #117009: the British -ise family now mirrors the -ize
family (e|es|ed|ing|ation); the curated-vocabulary table keeps one
prose-negative per false-positive stem and one auth-positive per family.
The curated auth alternation from #117009 dropped authenticates/
authenticating, authorizes/authorizing, and the British authorise/
authorised/authorisation spellings, so those genuine auth-failure
strings would miss the respawn guard and retry immediately instead of
auto-blocking. Extend the alternation to cover them, per review.
_RESPAWN_BLOCKER_RE used auth\w* to detect quota/auth failures in
last_failure_error, which also matches ordinary English words like
"author", "authored", "authoring", "authoritative" and "authors" in
worker progress prose. Because the guard is re-derived from an
unchanged string on every dispatch tick with no terminal state, a
single such word permanently parks an otherwise-healthy ready card
(observed: 437 respawn_guarded events over 7.3h on one card, #117009).
Replace the open stem with a curated set of real auth-failure tokens
(auth, authenticate/authenticated/authentication,
authorize/authorized/authorization, authz) so the false-positive
words stop matching while genuine auth failures still do.
Review follow-up: the boot lock's mkdir/os.open ran in the context manager's
__enter__, outside the body's try/except, so a read-only runtimes dir or a
foreign-owned boot.lock raised OSError out of ensure_local_runtime — the
function main documented as never raising into session start. Catch OSError
around acquisition, warn, and yield unlocked (the contention-timeout shape).
Two Hermes backends starting within the same second (e.g. default +
non-default profile) each see no server.json yet and each spawn a
llama-server router on the stable port — the existing per-process
_SUPERVISOR singleton can't rule this out since every process starts
with its own copy. Wrap the state-check/spawn sequence in an OS-held
file lock (fcntl/msvcrt, mirroring managed_uv.py's install lock) so the
second caller waits for the first to publish its state and adopts it
instead of racing to spawn its own.
Fixes#116682
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>