Commit Graph

39639 Commits

Author SHA1 Message Date
teknium1
2ca53fc386 refactor(cli): move cli.py module-level helper clusters into topical siblings; cli.py lands under 2,000 lines (#116911)
Mechanical, behaviour-neutral extraction along the existing hermes_cli/cli_*.py
pattern. Six clusters of module-level helpers leave cli.py:

- cli_config_load.py     prefill messages, reasoning/service-tier parsing, terminal
                         env mirroring, CLI defaults + YAML merge, logging bootstrap
- cli_render.py          reasoning-tag stripping, ANSI/skin colours, light-mode
                         detection, markdown/final rendering, output-history
                         recording, _cprint, ChatConsole, compact banner, panel wrap
- cli_terminal_input.py  file drops/attachments, bracketed-paste patch, extended
                         Enter keys, CPR guards, TUI input height, query images
- cli_shutdown.py        process session-id sync, exit watchdog, cleanup steps,
                         session-finalize notifications, one-shot finalize
- cli_single_query.py    kanban goal loops, exit-code mapping, quiet -q runner,
                         image routing, signal handlers, single-query mode
- cli_auto_maintenance.py state-db / checkpoint startup maintenance

Every moved body is AST-identical to the base copy modulo two mechanical
edits: cli-level names are late-bound with a call-time `from cli import ...`
(the cli_init_mixin pattern) and mutable cli state is read through `_cli().NAME`
(the gateway_service_unit `_gw()` pattern), so every monkeypatch seam on the
`cli` facade still intercepts moved code and no read snapshots stale state.
Mutable module state and every `global`-writing function stay in cli.py; the
facade re-exports the moved names in one import block.

test_bracketed_paste_timeout AST-loads the paste helper from its new home.
cli.py 3959 -> 1820 lines.
2026-09-20 20:42:17 -07:00
teknium1
1142a5b372 test: fold residency-cap tests into two parametrized invariants
Same cases (tight card, roomy card, ceiling, no device, refused model, boot wiring,
broken probe), within the salvage bar of two tests.
2026-09-20 20:42:05 -07:00
finn763
eaa5523f11 fix(local-runtime): cap resident models by the hardware budget, not a count
Residency was bounded by `local_runtime.models_max` alone, so a second model was admitted
against an already-full card. On Windows/WDDM that over-commit is not refused: the
allocation is paged to host memory and the child decodes at about a third of its speed for
the rest of its life — no error, no UI hint, and ejecting the incumbent afterwards does not
repair it (only a clean reload does).

The router now gets a cap priced from the same capacity budget the presets are priced
against: the count rises above one only while the largest staged model still fits twice, so
any pair of staged models is inside the card by construction and llama.cpp evicts its LRU
before an incoming child allocates. `models_max` stays a ceiling (a user's smaller number is
honoured) and an unpriceable budget or model keeps today's count.

Refs #116078

(cherry picked from commit d7abad7984d16d23ab2442b92d623b2270776586)
2026-09-20 20:42:05 -07:00
teknium1
e77c516e73 fix(models): DeepSeek picker drops retired ids and labels V4.1 Flash (#117516, salvage #117525)
Native `deepseek` is curated-only: `deepseek-flash` / `deepseek-v4-pro`, in that
order. models.dev still indexes the retired `deepseek-v4-flash*` ids, and with
`deepseek` in `_MODELS_DEV_PREFERRED` the registry union re-added them
registry-first whenever the live /models fetch was unavailable. The Desktop
label for `deepseek-flash` now reads "DeepSeek V4.1 Flash".

Trim of the contributor diff: the deepseek-specific branch inside
`_merge_with_models_dev` was unreachable once deepseek left the preferred set,
so it is dropped; the test drives `provider_model_ids("deepseek")` instead of
the helper.
2026-09-20 20:41:54 -07:00
Forkbert
a839530cd8 fix(models): clean up DeepSeek picker catalog
(cherry picked from commit f0b854bd067573e33f6ca8e1c5f68b184f4fef90)
2026-09-20 20:41:54 -07:00
teknium1
7756d04e40 fix(desktop): export Switch from the four remaining plugin-sdk test mocks
group-chat-view now renders the hold-detection Switch, so every sibling test that wholesale-mocks
@hermes/plugin-sdk must return it; only the render test's mock had been updated.
2026-09-20 20:41:43 -07:00
teknium1
fb9c77e064 docs(bot-mode): held messages are replayed after release; Detect stop directives toggle
User-visible behaviour from #117488/#117472 and the quoted-stop-word rule from #117605.
2026-09-20 20:41:43 -07:00
fangliquanflq
49c27fa50e test(desktop): keep group view SDK mock complete 2026-09-20 20:41:43 -07:00
fangliquanflq
d633255a0c fix(desktop): preserve explicit group stop cancellation 2026-09-20 20:41:43 -07:00
fangliquanflq
a273c2c031 fix(desktop): preserve messages skipped by group holds 2026-09-20 20:41:43 -07:00
teknium1
ba43baba41 test: keep the 1M-window cap test in scope after the #116467 pin scoping
config_context_length_for_runtime now scopes model.context_length to the configured default route,
so the bare test runtime no longer inherits the 1M window; the test is about the cap, not the scoping.
2026-09-20 20:41:31 -07:00
teknium1
aa667cf13f fix(tui-gateway): drop the dead import guard around the default threshold cap 2026-09-20 20:41:31 -07:00
liuhao1024
b06052ecd9 fix(tui-gateway): restore the default threshold cap when the live config key is absent
The per-turn live compression sync reads the UNMERGED user config (a missing
key means "unset"), so a config.yaml without compression.threshold_tokens
coerced the cap to None — wiping the 256K default the ctor had installed from
the merged config read. The next preflight re-derived the uncapped ratio
trigger (500K on a 1M-window model), so compaction stopped firing at 256K
while the compression telemetry kept reporting the capped figure.

Restore the DEFAULT_CONFIG value on key removal, matching what a fresh agent
build installs; an explicit `threshold_tokens: null` still means ratio-only.

(cherry picked from commit 415f8e4482e387e627da335924b39816ae624648)
2026-09-20 20:41:31 -07:00
teknium1
3a281fa545 test: held-store gate asserts no --db on the gated actions only
main went red when `hermes sessions set-journal-mode --db PATH` (96da5d97) and the
held-store gate test (968422e7) merged from branches that never saw each other. The
test scanned every `hermes sessions` subparser for --db; set-journal-mode is an
offline command with its own holder scan and takes --db legitimately, and it never
goes through the held-store gate. Scope the invariant to _HELD_STORE_ACTIONS, which
is what the gate actually protects.
2026-09-20 20:36:42 -07:00
teknium1
b6a87db515 fix(desktop): onboarding and model pill share one provider label table
The salvaged providerDisplayName duplicated onboarding PROVIDER_DISPLAY titles;
keep one table in the lib and let onboarding keep only its featured order.
2026-09-20 20:25:34 -07:00
Forkbert
3235206c5e fix(desktop): label xAI OAuth model selections
(cherry picked from commit c3f708e93b4777f1808fe189a0331aae7148712f)
2026-09-20 20:25:34 -07:00
teknium1
4c2c20f6e7 feat(website): plugin catalog shows when each entry was added/updated and can sort by it
extract-plugins.py derives addedAt/updatedAt per entry from one `git log
--name-status` over plugin-catalog/ (committer dates, so rebase-merged PRs
are stamped when they land; renames carry the original addedAt). A shallow
checkout or missing git yields null dates instead of wrong ones, and
deploy-site.yml now checks out with fetch-depth: 0 so the deploy has the
history.

The /docs/plugins page gains a Sort control (Most starred / Newest /
Recently updated) and an "Added … · Updated …" line on every card; Updated
is hidden while it equals Added.
2026-09-20 20:18:58 -07:00
teknium1
e089be80a1 docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
2026-09-20 20:14:20 -07:00
teknium1
56af2bdf7e fix(desktop): idle renderer stops ticking once a decode placeholder resolves
The DecodeText scramble primitive replayed forever by default, so every
quiet-surface placeholder that uses it (the empty pane zone's "HERMES"
mark, the contrib LOGS/HERMES panes) kept a 45 ms setInterval + setState
alive for as long as it was on screen. Measured with the perf harness on a
seeded 200-message session with no turn running, that ticker alone held the
idle renderer at ~9.6 React commits/s (DecodeText: 16 renders/s); with it
resolved once, the same workspace idles at ~0.7 commits/s.

`loop` now defaults to false; the boot "CONNECTING" overlay — the one
progress surface that should replay while it waits — asks for it explicitly.
The vitest drives the real component under fake timers and is red on base
(the tail keeps scrambling after the hold).

Part of #98394
2026-09-20 20:14:20 -07:00
teknium1
28ff18740b test(desktop): pin the suppressed-rotation aftermath and trim the close-out scenario
Review follow-ups on the rotation-foreground gate:

- Assert the residual a suppressed rotation leaves behind. Nothing
  re-points the primary, so its selection and route keep the
  PRE-rotation stored id; that is benign only because the lineage row
  still resolves it to the tip, which is now asserted with the same
  cachedSessionRow() lookup every sidebar/route resume goes through.
- Note in the suppression cases that the "hash route" shape IS the
  pop-out/secondary-window shape: isSecondaryWindow() starts that
  renderer with $sessionTiles empty and $layoutTree null, so
  $focusedStoredSessionId collapses to the selection and the route is
  the only voice left. A separate pop-out case would be a byte-identical
  store state.
- Trim wrong-session-closeout.test.tsx to the one invariant this fix
  owns: the real producer wired to the real route-follow consumer must
  not navigate, re-select, or move focus. The queued-drain fence and the
  attachment recovery it also composed are pinned by their own suites
  (use-background-queue-drain, use-prompt-actions), so coupling this
  fix's red signal to them only misattributes failures.
2026-09-20 20:14:10 -07:00
teknium1
806a86ae7d test(desktop): pin the wrong-session close-out scenario across the mapped producers
Composes the producers mapped on #86106 into one regression against the real
cache, session-actions and prompt-actions hooks plus the production
owner-routing dispatcher (gateway edge mocked, no injected bindings):

  A active -> B queued send -> user focuses tile C -> B runtime reaped ->
  A's delayed stored-id rotation -> B drain (text + image) -> B recovers

On origin/main the rotation consumer navigates the primary to A-next while
the user is on tile C (route/focus steal); with the #86359 gate the user stays
on C, B's attachment and prompt reach connection-B on B's recovered runtime,
and every stored id maps to exactly its own runtime.
2026-09-20 20:14:10 -07:00
teknium1
c2b94ca273 test(desktop): consolidate the rotation-foreground suite into two invariants
Fold the six #86359 cases into two it.each tables: one negative over the
three surfaces that can name the on-screen session (stored selection,
HashRouter route, focused tile) and one positive over the shapes that still
belong to the rotating lineage. Same coverage, one assertion body per
invariant, so a future change to the foreground rule has one place to read.
2026-09-20 20:14:10 -07:00
Christopher
9aab00844c test(desktop): satisfy import ordering in session state cache test 2026-09-20 20:14:10 -07:00
Christopher
6d7a8005d0 fix(desktop): gate rotation follow on focused surface and hash route 2026-09-20 20:14:10 -07:00
Christopher
32bb86ff72 fix(desktop): cover route-based foreground and unselected A-next rotation
isSessionInForeground now treats a missing route and store selection as
the on-screen chat, so a fresh session can still emit A -> A-next.
Tests drive the primary-route branch via history.pushState and pin the
no-route/no-selection path.
2026-09-20 20:14:10 -07:00
Christopher
ced5f36b2a fix(desktop): prevent background session rotation from stealing foreground route (#86106)
When a session stored id is rotated (e.g. by auto-compression), the
rotation event was emitted whenever the active runtime id matched,
ignoring whether the user had already switched focus to a different
session. This let a fast A -> B switch followed by A delayed rotation
pull the foreground route back to A new tip.

Add isSessionInForeground() in session-states.ts that checks the
current route/selected stored id against the rotating session lineage.
Apply the guard both in handleTransition() and in the no-op-updater
path inside useSessionStateCache.ensureSessionState() so all rotation
emission sites re-validate the user current focus.

Include a regression test that sets selectedStoredSessionId to an
unrelated session while the active runtime rotates, expecting no
activeSessionStoredIdRotation event.
2026-09-20 20:14:10 -07:00
teknium1
b4345b4e57 fix(gateway): suppress the completion send only when the queued lane's send actually landed
The queued-follow-up lane treated "the send call did not raise" as "the text reached the
chat". `_deliver_queued_first_response` returns normally when the adapter reports
`SendResult(success=False)` (flood control, retries exhausted): `_send_queued_final_text`
only logs the failure. The result was then marked `already_sent`, the completion path
suppressed the only remaining send, and the user got nothing at all — the opposite failure
to the duplicate this lane fixes, on the same code path.

`_deliver_queued_first_response` now reports whether the text is in the chat, and the mark is
gated on it. A refused send returns False (and skips the attachment upload, so the completion
send replays text and files together); a connector egress DECLINE still returns True, because
that destination is not approved and must not be re-sent.

Attachments were also uploaded twice on the single-delivery path: the queued lane uploads the
response's MEDIA: files, then the completion path's `already_sent` rescan uploaded them again.
The result now records `media_already_delivered` and the rescan skips them.

Live (real gateway, real IRC adapter, stand-in server/model, temp home): follow-up refused by
the reference guard with the first send refused — pre-fix head delivered nothing, this head
delivers the answer once.
2026-09-20 20:14:04 -07:00
teknium1
ddd3c5429a fix: queued-follow-up lane marks the first response delivered so the completion path never re-sends it
When a message is queued while a turn runs, `_run_agent_deliver_first_response` sends the
first turn's final itself (with streaming off — the shipped default — the stream consumer
never confirms delivery, so this fallback fires on every queued follow-up). Nothing recorded
that send, so every early `return result` in `_run_agent_queued_followup` after it (follow-up
text refused by the @-reference guard, stale goal continuation) handed the first turn's
result back to `_handle_message_with_agent`, whose normal completion send delivered the
identical text a second time (#81052).

`already_sent` on the result dict is the flag the completion path (`_hmwa_deliver_turn_response`)
already consults for streamed finals; the queued lane now sets it on both result objects it
may return, right after its own send succeeds. Live: real gateway + real IRC adapter over a
stand-in server, streaming off, second message queued mid-turn — each final delivered once,
follow-up still runs.
2026-09-20 20:14:04 -07:00
teknium1
e7e1928174 fix(update): the invoking profile's launchd restart is claimed only on a pid that changed
`hermes update` on macOS printed "✓ Restarted ai.hermes.gateway" for the
profile running the update whenever launchd reported *any* pid for the label —
including the pre-update process the restart was supposed to replace.
f29ee96dd3 closed this for SIBLING labels only
(`_wait_for_launchd_service_pid(label, old_pid=old_pid, ...)`); the invoking
profile kept `wait_for_launchd_gateway_supervision` →
`_launchctl_label_supervising_process`, a predicate that is true for "launchctl
list exits 0 and some positive pid" and never compares against the pid seen
before the restart.

- `_launchctl_supervised_pid()` exposes the pid `_launchctl_label_supervising_process`
  already parsed; the boolean is now a thin wrapper over it.
- `wait_for_launchd_gateway_supervision(..., old_pid=...)` accepts a fresh pid
  only, mirroring the sibling loop's contract. `old_pid=None` keeps the old
  "any supervised pid" meaning for callers with no pre-restart observation.
- `_restart_launchd_gateway_after_update()` snapshots the pid before
  `launchd_restart()` and passes it in; the failure line now says launchd is
  not supervising a NEW process. The snapshot is verification-only, so the
  no-`launchctl list`-gating invariant of #74973 is untouched.

Also from round-2 review of this PR:
- the supervised-serve discharge test parametrized 5 supervisors, 3 of which
  the inventory writer can never put on a serve/dashboard row; cut to the two
  it can emit and documented the parity-only members.
- `_marker_only_restart_obsolete` now names who still owns a discharged
  launchd-supervised serve (update-time `report_unaccounted_runtimes`, exit 1).
2026-09-20 20:13:57 -07:00
teknium1
e9197afd6d fix(update): a supervised serve backend no longer pins the fleet-restart warning on forever
A `fleet_restart_pending` marker records the pre-update plan's runtimes as its inventory,
including `serve`/`dashboard` rows. `_marker_only_restart_obsolete` rejected every non-gateway
inventory row outright, so on any host that runs a dashboard or the Desktop `serve` backend the
marker could never discharge: `hermes version` kept printing "a previous `hermes update` ... did
not restart running gateways" after every gateway was already current on the pulled SHA.

Receipts already draw this boundary (`_receipt_owed_gateways`, #115090) and so does the restart
phase for the Desktop backend (#111494); the marker path was the one place left counting a
supervisor-owned serve row as evidence against the gateways it does not cover. A manual-serve row
still needs its durable handoff and an unclassified backend still fails closed.

Also corrects `_owed_stale_serve_rows`'s docstring, which justified the Desktop exclusion by the
"armed forever" consequence this change removes.
2026-09-20 20:13:57 -07:00
teknium1
a5d561b1ea fix: state the real failure mode of the model_config sanitiser
The docstring claimed every `json_extract` reader (`hermes sessions list`, the
dashboard chain CTE) raises on the recovered database. Both are guarded:
`_sql_json_extract` wraps each extract in `CASE WHEN json_valid`. The verified
symptom is the write path — `reopen_session` rewrites reset-child markers with
`json_set` and raises `OperationalError: malformed JSON` on the first resume.

The test drops its two mechanism assertions (`model_config == '{}'`,
`model_config_reset == 1`) and keeps `reopen_session('parent')`, the invariant that
is red on base; a different valid repair must not fail a working store.
2026-09-20 20:13:41 -07:00
teknium1
411a75a562 fix(gateway): auto-archive every served profile's store, not just the launch home
`hermes serve`'s opportunistic auto-archive now defers to `_check_gateway_running`
(the canonical per-profile predicate `_maybe_run_skill_maintenance` already uses)
instead of a hand-rolled `gateway.lock` probe. The lock is per PROCESS HOME, so a
multiplexed secondary's store still got a second writer from serve; the predicate's
multiplexer rung covers a profile that owns no gateway.pid of its own.

Deferring is only correct if the gateway actually sweeps that store, and it did not:
the "Auto-archive tick" chore resolves `acquire()`/`load_config()` through
`get_hermes_home()` and ran unscoped, so a served secondary was archived by nobody.
It now rides `profile_scoped_chore`, like the curator and sync ticks.

Invariants: the serve-side guard is proven against a real flock holder with a live
`gateway run` command line and a gateway.pid record (kernel state + process identity,
no patched predicate), and the tick is proven to sweep A's and B's state.db.
2026-09-20 20:13:41 -07:00
teknium1
a263df8dcd fix: backup import refuses to publish over a deleted-but-held state.db
`_import_db_member` guards the replace path with `_safe_restore_db`, but its
"target is missing" branch treated absence as proof of no holders. A gateway or
dashboard that had the database open when it was unlinked keeps writing the
deleted inode, so an atomic publish there re-creates exactly the #90950 split
brain the replace path exists to prevent.

Consult `_foreign_db_holder_pids` before the publish and raise an OSError
naming the PIDs, which the caller already reports as a skipped file.

Credit: @ngpestelos (#110179)
2026-09-20 20:13:41 -07:00
teknium1
7c2914becf fix: recovery no longer ships a store that raises on the first resume
`integrity_check` never inspects column contents, so a source whose
`sessions.model_config` JSON was truncated by the damage recovers "clean" and
the recovered store then raises `OperationalError: malformed JSON` the first
time a caller resumes a parent session — `reopen_session` rewrites reset-child
markers with `json_set`, which refuses unparseable input.

Reset unparseable blobs to '{}' in `_finalize_derived_metadata`, the one
boundary both the SQL-copy lane and the `.recover` lost_and_found lane pass
through; the count lands in the recovery report.

Credit: @efe-arv (#101679)
2026-09-20 20:13:41 -07:00
teknium1
b4f8d83e12 fix: hermes serve stops opening state.db writable under a live gateway
`_maybe_auto_archive_for_profile` ran 90s after bind (and on every dashboard
session-list request) and opened a WRITABLE SessionDB even when a gateway owned
the store. The gateway's own housekeeping already runs that sweep, so the serve
copy only added a second writer to a database another process is archiving.

Return early when the profile's gateway runtime lock is held.

Credit: @isair (#110405)
2026-09-20 20:13:41 -07:00
teknium1
968422e7c5 fix(gateway): keep the operator restart tail on the home-channel storage notice
The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).

Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
2026-09-20 20:13:33 -07:00
teknium1
f88ba2ea80 fix(console): accept the --force override the held-store refusal advertises
The Desktop/dashboard console printed `hermes sessions optimize --force` as the override, but
_sessions_optimize rejected every argument — with a gateway running the command could only
ever refuse. It now parses --force itself (and the hint names the console form).
2026-09-20 20:13:33 -07:00
teknium1
1dc881b67a fix(state): automatic maintenance skips the VACUUM while another process holds state.db
maybe_auto_prune_and_vacuum() runs the same store rewrite as `hermes sessions optimize`
(VACUUM + TRUNCATE checkpoint) from CLI startup and the gateway constructor, with no holder
scan — so the manual command was gated while the automatic producer of the same #110054
failure was not. The VACUUM branch now runs the same foreign_state_db_holders admission and
SKIPS (debug log + a vacuum_skipped_holders count in the result) when a sibling writer holds
the store or a WAL sidecar. Housekeeping never refuses a turn; it only defers the rewrite to
the next run.
2026-09-20 20:13:33 -07:00
teknium1
dfaf016dbc fix(sessions): resolve the held-store scan path from the store resolver, not the db object
The admission gate read `db.db_path`, which made the refusal depend on whatever object
`SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with
AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from
`_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the
scan targets the same file in production and stays reachable regardless of the db object.
2026-09-20 20:13:33 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
12fb5eb476 fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb:

- The header probe used os.pread, which does not exist on Windows, while the subparser is
  registered unconditionally — the command died there with an uncaught AttributeError. It now
  reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
  admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
  running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
  (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
  silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
  header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
  bail in the command's own error style; every open/read is guarded.

The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
2026-09-20 20:13:25 -07:00
teknium1
96da5d97fc feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.

The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
2026-09-20 20:13:25 -07:00
teknium1
95b1f4c855 fix(cron): scope the yield predicate's claims to what the record proves
Review follow-ups on the stale-code tick yield gate:

- `_gate_mocks` patches `_current_gateway_code_sha` with `raising=False` so the
  base tree (where the symbol does not exist yet) fails only on the two new
  invariants instead of erroring the whole module out with AttributeError.
- Note at the pid comparison that `get_running_pid(pid_path=None)` can fall back
  to `get_runtime_status_running_pid()`, which derives the pid from the same
  record — the equality is tautological in that branch and the real proof there
  is `runtime_status_is_stale` + `code_sha`.
- Docstring says what the record actually is: `gateway_state.json` is per-HOME
  and last-writer-wins, the pid equality binds it to the lock holder, and the
  `--replace` takeover window fails open by design.
2026-09-20 20:13:19 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
fangliquanflq
153b5f0145 fix(cron): compare full gateway code identities 2026-09-20 20:13:19 -07:00
fangliquanflq
bc04c35f6a fix(cron): verify fresh gateway before yielding ticks 2026-09-20 20:13:19 -07:00
Austin Pickett
1a1f4a59e2 fix(auth): explicit-provider gate uses the credential resolver's reader
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-20 22:46:49 -04:00
xxxigm
4ea57fa61b fix(auth): count profile .env keys in the explicit provider gate
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
2026-09-20 22:46:49 -04:00
xxxigm
050ea53aba fix(desktop): honor Applies-to on Custom Endpoints
The page only showed a read-only active-profile note, so endpoint
saves followed the left-rail Bot instead of the Settings chips
Accounts and API keys already share.
2026-09-20 22:46:49 -04:00
teknium1
8ba5fe9c16 test: assert composer-paste guard on result.message/warnings, not the expanded flag
The new test read result.expanded (a bool) as if it were the expanded
message, so it raised TypeError in CI instead of exercising the guard.
Check the body on result.message and the refusal on result.warnings;
still red without the production allow-root and green with it.
2026-09-20 19:45:30 -07:00