Commit Graph

8564 Commits

Author SHA1 Message Date
teknium1
dfaf016dbc fix(sessions): resolve the held-store scan path from the store resolver, not the db object
The admission gate read `db.db_path`, which made the refusal depend on whatever object
`SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with
AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from
`_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the
scan targets the same file in production and stays reachable regardless of the db object.
2026-09-20 20:13:33 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
12fb5eb476 fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb:

- The header probe used os.pread, which does not exist on Windows, while the subparser is
  registered unconditionally — the command died there with an uncaught AttributeError. It now
  reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
  admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
  running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
  (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
  silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
  header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
  bail in the command's own error style; every open/read is guarded.

The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
2026-09-20 20:13:25 -07:00
teknium1
96da5d97fc feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.

The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
2026-09-20 20:13:25 -07:00
Austin Pickett
1a1f4a59e2 fix(auth): explicit-provider gate uses the credential resolver's reader
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-20 22:46:49 -04:00
xxxigm
4ea57fa61b fix(auth): count profile .env keys in the explicit provider gate
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
2026-09-20 22:46:49 -04:00
teknium1
c746249861 fix(kanban): CLI and kanban_complete tool report an empty-completion refusal
EmptyCompletionError is raised by complete_task; without a handler the
CLI and the tool would surface a traceback instead of the actionable
refusal (add --result/--summary). Same shape as the HallucinatedCards gate.

Part of #117483
2026-09-20 19:03:47 -07:00
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
teknium1
442636bf44 fix(update): settle latest.json only when the live fleet discharges the pending restart
The catch-up rewrote the receipt's fleet matrix before checking whether that
actually cleared the debt, so an exit-1 run (marker still owes a gateway,
blocked serve_restart_pending storage) left latest.json mutated —
tests/hermes_cli/test_update_scoped_reconciliation.py pins the receipt as
byte-identical on the failure path. Decide on the in-memory settled copy and
write only when the pending restart is discharged (#117051).
2026-09-20 18:52:45 -07:00
teknium1
29c4ec8fa7 fix(update): settle a stuck failed receipt once the live fleet serves the checkout code
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).

In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.

Fixes #117051
2026-09-20 18:52:45 -07:00
teknium1
d6b0d37ece fix(update): pending fleet restart leaves gateways already on the checkout code alone
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").

Part of #117051
2026-09-20 18:52:45 -07:00
teknium1
eaef5ec717 feat(update): hermes update --list-venv-holders prints the venv guard's holders as JSON, exit 3
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].

Fixes #117246
2026-09-20 18:51:36 -07:00
John Paul Soliva
55a8f42329 perf(profiles): the listing re-reads a profile's YAML only when that file has changed
`list_profiles()` parsed three YAML files per profile on every call — config.yaml through
`read_user_config_raw` ("no cache", by contract) plus profile.yaml and distribution.yaml through
`_load_yaml_dict`. It is the shared body of `GET /api/profiles` and `profiles.list`, which the Bots
roster polls every 5s per connection, so an idle machine re-parsed every profile's YAML twelve
times a minute to produce the same handful of strings.

config.yaml is the expensive one: the installer seeds it by copying the annotated template verbatim
(`cp cli-config.yaml.example config.yaml`), 119,268 bytes / 2,267 lines of mostly comments, for
`model.default` and `model.provider`.

Measured, 9 profiles, no gateway locks so the liveness path stays out of it: 16.1ms -> 1.6ms with
the installer template, 6.4ms -> 1.6ms hand-trimmed; 27 YAML parses per pass -> 2, and those two
are the default profile's ABSENT files, correctly not cached. The pass is now independent of config
size.

The memo keys on each file's (mtime_ns, size, inode) — the atomic writers rename a temp file into
place, so a rewrite always lands a new inode even inside one mtime tick — and caches only the small
DERIVED values. `read_user_config_raw` and `_load_yaml_dict` keep their uncached contract: callers
write those documents back, and a stale read there could overwrite a newer file. `read_profile_meta`
hands out a copy so a future mutating caller cannot poison the next reader.

Regressions drive the real readers: an edited config/profile/distribution file is seen on the very
next read, `write_profile_meta` round-trips, a missing file is not cached and is picked up when
created, mutating a returned meta does not poison the next reader, an unchanged file is parsed once
across three reads, and `read_user_config_raw` still re-reads. Ignoring the signature fails four.

Fixes #117378
2026-09-20 18:45:07 -07:00
fangliquanflq
1a4c68ce97 fix(kanban): expose task runtime limit in JSON 2026-09-20 18:43:59 -07:00
chadhouser
79c4c363c6 fix(kanban): derive review implementer from the run, not the card's assignee
request_review recorded trow["assignee"] as the implementer on the
review_requested event and on the synthesized run. That is correct while a
worker holds the card, but wrong when the card was created already assigned to
its reviewer -- `kanban create --assignee <reviewer>` followed by
`request-review`. Both provenance records then name the reviewer, and
request_changes routes a rejection back to the profile that wrote the findings.

Derive the implementer after the reviewer is resolved, from the active run's
profile, falling back to the assignee only when it is not the incoming
reviewer. When no actor can be established the field is left unset, which
request_changes already refuses on rather than misrouting.

Refs #117229. Follow-up to #111064 / #111459, which fixed the run attribution
for the distinct-reviewer case and left the derivation itself unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 18:43:35 -07:00
teknium1
aaafa80577 refactor(cli): move HermesCLI init phases and TUI run-loop phases into mixin siblings (#116911)
Mechanical, behaviour-neutral extraction along the existing cli_*_mixin.py
pattern: the ten _init_* constructor phases become CLIInitMixin
(hermes_cli/cli_init_mixin.py) and the fourteen _tui_* run-loop phases
(input dispatch, after-turn, startup banner/prewarm/maintenance, application
build, signal handlers, shutdown) become CLITuiRuntimeMixin
(hermes_cli/cli_tui_runtime_mixin.py). __init__ and run() stay in cli.py as
the orchestrators. Every moved body is AST-identical to the base copy; cli.py
module names are resolved lazily through cli so monkeypatch seams survive.
cli.py 4835 -> 3960 lines.
2026-09-20 18:43:11 -07:00
teknium1
70fceb80ab refactor(gateway): move the systemd service-unit cluster out of hermes_cli/gateway.py
Mechanical extraction of the service-definition cluster (generate_systemd_unit,
systemd_unit_is_current, refresh_systemd_unit_if_needed and their
_systemd_*/_service_*/_ld_library_path/_temp_home helpers) into
hermes_cli/gateway_service_unit.py. Behaviour-neutral: moved bodies read facade
helpers (get_python_path, _profile_arg, _run_systemctl, load_gateway_config, ...)
and each other through _gw() — a late import of hermes_cli.gateway — so every
seam tests and callers patch on the facade keeps intercepting the moved code;
the facade keeps one import block re-binding the names.

hermes_cli/gateway.py: 6808 -> 6496 lines.

Part of #116913
2026-09-20 18:33:07 -07:00
teknium1
5aeddb6c85 fix(auth): user provider plugin owns its aliases and display name in the auth registry
register_plugin_provider setdefault-ed aliases regardless of provenance, so a
$HERMES_HOME plugin alias colliding with an existing row kept resolving to the
old row while providers.get_provider_profile already followed the user profile;
override_registry_row left the display name untouched on a same-name
replacement. mirror_aliases applies the provider_source ownership rule that
e5e7fbcd27 introduced for endpoints.

Fixes #116668
2026-09-20 18:29:22 -07:00
rodricksz4h5
f50c6f43a2 fix(cli): let the early interface probe follow the profile's home
`_suppress_mouse_residue_early()` runs at import time, before
`_apply_profile_override()` sets HERMES_HOME, so the first read of
`display.interface` always saw the default home. The result was memoised
in a cache keyed on nothing, so every later caller — the Termux fast
paths and the TUI launch decision, which all run after the profile is
applied — was pinned to the default home's interface for the whole run.
`hermes -p coder` with `display.interface: tui` in that profile booted
the classic REPL, and the mirror case booted the TUI for a profile that
asked for cli.

Key the cache on the config path it read. The hot path still parses the
YAML once per home, and the first call after the process is re-homed
re-reads the file it should have read.

Fixes #116902
2026-09-20 18:26:14 -07:00
rodricksz4h5
b3ae818261 fix(cli): attribute a scrubbed-env dashboard to its owner's home
`_hermes_home_for_pid` fell back to the INSPECTING process's `Path.home()`
when the target's environment carried neither HERMES_HOME nor HOME, which
is the normal shape for a systemd/launchd unit with a scrubbed
environment. The dashboard was then attributed to whichever user ran the
command, so `hermes update` and `--stop` read the wrong profile root and
could act on the wrong backend.

A process without HOME resolves its own default home from the password
database, so resolve the target's owner the same way (psutil, then the
/proc owner) before falling back. An unreadable owner or entry keeps the
previous behaviour rather than resolving to nothing.

Fixes #116906
2026-09-20 18:25:48 -07:00
teknium1
fafa2e215c refactor(cli): route /browser subcommands through a dispatch table (#116914)
The 5-branch if/elif ladder in CLICommandsMixin._handle_browser_command
becomes a _BROWSER_SUBCOMMANDS word -> handler table; adding a subcommand
is one row plus a usage line. Subcommand word is case-insensitive, the
argument keeps its case (CDP URL paths are case-sensitive).
2026-09-20 18:23:02 -07:00
teknium1
f76bdb6e8f fix(auth): PKCE token errors reuse the canonical grant-dead code set
Drop the duplicated _GRANT_DEAD_CODES frozenset and the relogin_required
plumbing: the pool's plugin recovery already treats exc.code in
hermes_cli.auth._OAUTH_GRANT_DEAD_CODES as terminal, so _token_http_error
imports that set late and just sets code=<json error>.
2026-09-20 18:21:34 -07:00
Adolanium
5eb0659177 fix(auth): mark PKCE invalid_grant dead and store alias logins canonically
A spent refresh token returned HTTP 400 invalid_grant, and the helper mapped that to oauth_refresh_failed so the pool EXHAUSTED the row. Parse the JSON error and DEAD it. hermes auth add with an alias wrote the pool under the typed name, so status against the profile name looked logged out.
2026-09-20 18:21:34 -07:00
fangliquan
7df20a670d fix(cli): fail closed without bang shell sanitizer 2026-09-20 18:20:44 -07:00
teknium1
dcae0d81d0 fix(dashboard): fold the new fallback config section into the agent tab
fallback.min_switch_reset_seconds is the only schema-surfaced fallback field; without the merge the dashboard grew a one-field orphan tab (test_no_single_field_categories red on CI).
2026-09-20 17:00:43 -07:00
teknium1
f568b860d7 fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
  (production entry) on a real AIAgent with a one-entry chain, so dropping the
  reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
  rate-limited primary's declared reset is sooner than N seconds,
  try_activate_fallback returns False and no cooldown is armed; docs row added.
2026-09-20 17:00:43 -07:00
teknium1
348568ddaf fix(profiles): hoist the profile-id grammar into hermes_constants.PROFILE_ID_RE
Review follow-up for #116905: bot_mode_probe was the third private copy of
the regex (profiles, update_restart_recovery) with a sync test policing the
drift the copy created. One canonical object next to named_profile_is_live;
the three modules import it and the sync test is gone.
2026-09-20 17:00:11 -07:00
teknium1
d325e53116 fix(kanban): drop the dead edit_completed_task_result shim and route dashboard priority edits through edit_task
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
2026-09-20 16:33:49 -07:00
fangliquanflq
3f399c0bd4 fix(cli): align kanban edit and paused cron listing 2026-09-20 16:33:49 -07:00
teknium1
891c1d79c6 fix(constants): one home for managed, container and HERMES_UID policy; drop the config twins
Review follow-up: get_managed_system, the container/chmod-skip check and
the HERMES_UID/GID chown (_resolve_hermes_uid_gid/_chown_to_hermes_uid) now
live only in hermes_constants — the import-safe module apply_secure_dir_policy
already needed them in — and hermes_cli.config re-exports them, so there are
no 'keep in sync' copies and _secure_file skips on the same canonical
_detect_container signal as _secure_dir/get_scratch_dir. The dead config
copies are deleted and test_ensure_hermes_home_uid.py drives the new
symbols; one invariant test pins the single implementation.
2026-09-20 16:32:50 -07:00
liuhao1024
a5605724c2 fix(constants): honor managed/shared-home permission policy in get_scratch_dir
get_scratch_dir unconditionally chmod'ed <home>/cache/scratch to 0700,
bypassing the policy _secure_dir() applies everywhere else: managed
installs are left alone, containers only apply an explicit
HERMES_HOME_MODE, and HERMES_UID/HERMES_GID ownership is honored. On a
group-shared home (setgid 2770) that stripped the setgid bit and broke
the shared-group inheritance the operator configured (#77579, #117347).

Move the policy primitive to hermes_constants.apply_secure_dir_policy()
(import-safe, stdlib-only) and route both get_scratch_dir() and
hermes_cli.config._secure_dir() through it, so one implementation
serves all callers including the cron.jobs delegation chain.

Fixes #117347
2026-09-20 16:32:50 -07:00
teknium1
265e68d784 fix(auth): one primary_failure_wording() helper labels quota vs auth failure at all three fallback surfaces
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by
  cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback
  logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482
  sibling: 'Primary auth failed' for a 429 on the gateway surface).
- Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the
  flag is only True when _ensure_runtime_credentials returned False).
- Fold the new tests: 2 parametrized production-path tests + 1 gateway test.
2026-09-20 16:23:22 -07:00
teknium1
75430ff290 refactor: one fallback notice call, label chosen up front
Fold the duplicated render_notification branches into a single call and re-point the
source-slicing test marker at the flattened block.
2026-09-20 16:23:22 -07:00
686f6c61
89096ba56c fix(cli): treat credential-resolution 429 as quota, not auth
A Codex quota wall at startup was labeled "Primary auth failed" and
exited 1, so Kanban counted a failure and operators went looking for
a bad token. Use is_rate_limited_auth_error() for the fallback notice
and EX_TEMPFAIL (75) when HERMES_KANBAN_TASK is set.
2026-09-20 16:23:22 -07:00
teknium1
214623389a fix(models_validate): one profile-catalog helper decides accept/reject before the generic listing
Fold _profile_owns_catalog and _profile_owned_catalog into _profile_catalog ->
(catalog, authoritative) so _validate_live_listing has a single owned-catalog
branch; drop the redundant catalog-hit test (covered by the owns-catalog accept
tests in test_setup_provider_catalog.py).
2026-09-20 16:22:25 -07:00
Konstantin Khlopkov
57a3bbbcc1 fix(models_validate): a profile's own catalog endpoint decides, not the generic /models listing 2026-09-20 16:22:25 -07:00
joaomarcos
d76c77beed fix(doctor): stop prescribing audit fixes that updates revert (#116774)
The doctor "Browser tools (agent-browser)" row audits the root workspace
tree (npm audit --workspaces=false), whose versions come from the committed
package-lock.json. hermes update reinstalls that exact state via npm ci
(_run_npm_install_deterministic never mutates the lockfile), so the
previously prescribed local `npm audit fix --workspaces=false` was wiped on
the next update and the finding returned - a fix/reintroduce loop (#116774).

- Bump the vulnerable root lockfile pins so the shipped tree audits clean:
  browserslist 4.28.6 -> 4.29.0 (high: GHSA-c83g-rgw3-j3cx,
  GHSA-73wf-gq98-2v4g), baseline-browser-mapping 2.10.43 -> 2.11.25
  (moderate: GHSA-w5vr-8v7q-w6rv), plus updated transitive deps
  (caniuse-lite, electron-to-chromium, node-releases,
  update-browserslist-db); also syncs the stale bootstrap-installer
  lockfile entry to its manifest version. npm audit --workspaces=false now
  reports 0 vulnerabilities.

- Rework the npm-audit remedy for every doctor row: the durable fix is an
  upstream lockfile bump; never prescribe a local mutating fix command
  (workspace-scoped rows already omitted it; the root row prescribed the
  doomed one).

- Invariant tests in tests/hermes_cli/test_doctor_audit_remedy.py: the
  remedy must not prescribe a mutating npm audit fix (proven red on base)
  and must name the lockfile-bump remedy.
2026-09-20 15:57:25 -07:00
tobenwarrior
dff92f3d78 fix(plugins): peel an annotated-tag pin before verifying the checkout
A pin is validated as 40 hex, but that does not make it a commit: a tag
object has a sha of its own, and an entry recorded from 'git rev-parse <tag>'
names the tag object rather than the commit it points at. Git resolves the
ref and detaches at the commit, so the revision guard compared a commit
against a tag-object sha and refused a CORRECT checkout — the entry was
uninstallable, with no way through (see the kiro-acp catalog entry).

Peel the pin to its commit before comparing. The guard keeps its teeth: HEAD
must still be exactly the commit the pin names, and a checkout that lands
anywhere else is still rejected. An unresolvable pin falls back to the raw
string, so a genuinely wrong sha fails exactly as before.

Tests: an annotated-tag pin installs at its commit (and records the commit,
not the tag object, so 'plugins update' drift checks stay exact); a
checkout that lands on another commit is still rejected. Red on base:
'Checked-out revision "..." does not match requested commit "..."'.

(cherry picked from commit 66baa043b70c4f66ae59c0a1147e74ba694140ef)
2026-09-20 15:56:55 -07:00
Aaron08140
8e4c943477 fix(cli): clear-screen fallback spawns no shell and no console window (#116904)
_clear_terminal_on_exit() fell back to os.system('cls'/'clear') when the
escape sequence was rejected. os.system() spawns a shell: on Windows a
conhost window flashes before the process exits, and in minimal POSIX
containers without `clear` on PATH the shell fails silently while the
except-pass hides it. The repo already standardises on windows_hide_flags()
(CREATE_NO_WINDOW) for short-lived argv spawns; use it here too.

- Windows: subprocess.run(['cmd', '/c', 'cls'], creationflags=windows_hide_flags())
  (cls is a cmd builtin, so probing a `cls` exe would be wrong).
- POSIX: shutil.which('clear') first; skip the spawn entirely when absent
  instead of letting a shell swallow a 127.
- stdin=DEVNULL so a child that reads can never block CLI shutdown.
- Regression tests: parametrised nt/posix argv + creationflags assertions
  (RED on the os.system path), plus the no-`clear` skip case.
2026-09-20 15:56:33 -07:00
teknium1
31017b7344 fix(kanban): a caller asserting run ownership cannot classify a parked card
Review follow-up for #117363: the classify-in-place branch returned before
the expected_run_id guard, so a worker whose run had ended could type a
breaker-parked card. Refuse when expected_run_id is given; the run is over.
2026-09-20 15:55:14 -07:00
teknium1
9850010ff5 fix(kanban): classify only an untyped, run-free breaker block; restore sort order
Tighten the salvaged classify-in-place path to the reporter's contract:
only block_kind IS NULL and current_run_id IS NULL qualify, recurrence
accounting starts at 1, the audit event is the ordinary `blocked` kind so
existing consumers see the reason, and a typed block is never re-typed.
Drops the unrelated VALID_SORT_ORDERS deletion and trims tests to two
invariants (one drives the real enforce_max_runtime breaker path).
2026-09-20 15:55:14 -07:00
allo
c80ab1d8f1 fix(kanban): allow block_task to classify an already-blocked card
block_task() refuses to set block_kind on a card that is already
status=blocked because the WHERE clause only matches running/ready.
The circuit breaker (enforce_max_runtime) parks cards blocked with
block_kind=NULL, making them permanently unclassifiable.

Add an early-return path: when the card is already blocked and kind
is supplied, update block_kind in-place and append a block_classified
event. Idempotent when the same kind is already set.

Fixes #117363
2026-09-20 15:55:14 -07:00
teknium1
87f892b8db fix(kanban): cover authorising/authorises in the respawn auth pattern and trim the table
Review follow-up for #117009: the British -ise family now mirrors the -ize
family (e|es|ed|ing|ation); the curated-vocabulary table keeps one
prose-negative per false-positive stem and one auth-positive per family.
2026-09-20 15:53:55 -07:00
Jony
3899401aa6 fix(kanban): ignore auth words in crashed worker output 2026-09-20 15:53:55 -07:00
chelsealong
857709d63b fix(kanban): cover -es/-ing/British spellings in the auth blocker pattern
The curated auth alternation from #117009 dropped authenticates/
authenticating, authorizes/authorizing, and the British authorise/
authorised/authorisation spellings, so those genuine auth-failure
strings would miss the respawn guard and retry immediately instead of
auto-blocking. Extend the alternation to cover them, per review.
2026-09-20 15:53:55 -07:00
chelsealong
6a52c9e41d fix(kanban): curate the respawn guard's auth pattern instead of an open auth\w* stem
_RESPAWN_BLOCKER_RE used auth\w* to detect quota/auth failures in
last_failure_error, which also matches ordinary English words like
"author", "authored", "authoring", "authoritative" and "authors" in
worker progress prose. Because the guard is re-derived from an
unchanged string on every dispatch tick with no terminal state, a
single such word permanently parks an otherwise-healthy ready card
(observed: 437 respawn_guarded events over 7.3h on one card, #117009).

Replace the open stem with a curated set of real auth-failure tokens
(auth, authenticate/authenticated/authentication,
authorize/authorized/authorization, authz) so the false-positive
words stop matching while genuine auth failures still do.
2026-09-20 15:53:55 -07:00
teknium1
56b7a0fb37 fix(local-runtime): an unwritable boot lock degrades to an unlocked boot, never an OSError
Review follow-up: the boot lock's mkdir/os.open ran in the context manager's
__enter__, outside the body's try/except, so a read-only runtimes dir or a
foreign-owned boot.lock raised OSError out of ensure_local_runtime — the
function main documented as never raising into session start. Catch OSError
around acquisition, warn, and yield unlocked (the contention-timeout shape).
2026-09-20 15:53:25 -07:00
chelsealong
4352a7d053 fix(local-runtime): serialize managed runtime boot across processes
Two Hermes backends starting within the same second (e.g. default +
non-default profile) each see no server.json yet and each spawn a
llama-server router on the stable port — the existing per-process
_SUPERVISOR singleton can't rule this out since every process starts
with its own copy. Wrap the state-check/spawn sequence in an OS-held
file lock (fcntl/msvcrt, mirroring managed_uv.py's install lock) so the
second caller waits for the first to publish its state and adopts it
instead of racing to spawn its own.

Fixes #116682
2026-09-20 15:53:25 -07:00
teknium1
85564321be fix(windows): route every bare-bash spawn through _find_bash and surface silent interpreter failures
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-20 15:50:26 -07:00