6331 Commits

Author SHA1 Message Date
teknium1
7a06cb50d7 feat(telemetry): shared-metrics opt-in offer on every surface
One answer per profile, asked once, with the Desktop strip's three equal
choices (Send to Nous / Local only / No thanks) now offered by:

- `hermes setup`: at the end of every flow (Quick, Full, Blank Slate,
  Portal, --quick), via the setup-completed record they all funnel into.
- `hermes` / `hermes --tui`: once before an interactive chat on a profile
  that never answered. Skipped for -q, piped/JSON output, spawned actions
  (HERMES_NONINTERACTIVE), Desktop-hosted panes and managed installs.
- Web dashboard: a banner for the profile being managed, backed by
  GET/PUT /api/shared-metrics/consent.

"No thanks" is the terminal default so Enter never opts anyone in; Esc and
the banner's X leave the question open.

Also fixes the answer that could not stick: a CLI "no" equals the shipped
defaults, so save_config stripped it and every surface (Desktop strip
included) kept asking. The answer is now written to config.yaml directly,
and `decided` has one definition shared by the RPC, the CLI and the router.
2026-09-30 21:15:08 -07:00
Adolanium
5a2a300d99 fix(dashboard): remove session files when deleting empty sessions
DELETE /api/sessions/empty called delete_empty_sessions() without a
sessions_dir, so the rows were removed but their files in sessions/ stayed
on disk. Pass the profile's sessions dir, the same way the single delete
route does.
2026-09-30 23:14:09 -05:00
Brooklyn Nicholson
036c577c86 fix(desktop): surface staged skill-write review on the desktop and TUI
With skills.write_approval on, every skill write is staged to the
file-backed pending store, but no non-CLI surface could review it:
the desktop resolver marked the whole /skills family as the sidebar's
(desktop="settings") and discarded the subcommand token, and
tui_gateway command.dispatch had no route for /skills or /memory —
a failed slash worker turned both into the routing refusal, so staged
writes accumulated silently (#98330, #123315, #118812).

- hermes_cli/commands.py: new CommandDef.desktop_subcommands field;
  /skills declares its write-approval review slice
  (pending/approve/reject/diff/approval) and drops desktop="settings".
  command_desktop_meta serializes the allowlist (empty tuple = deny
  all); gateway contract regenerated.
- apps/desktop: the resolver threads the typed arg —
  desktopSubcommandAllowlist, an exec gate
  (desktopSubcommandUnavailableMessage) that renders an actionable
  refusal for hub mutations instead of forwarding them, completion
  filtering so the popover never suggests a dead end, and the
  isDesktopSlashCommand(name, arg) threading at the runExec gate.
  An explicit /skills spec keeps a current Desktop able to review
  staged writes even against an older backend catalog.
- tui_gateway command.dispatch: _cmd_skills answers only the
  registry-declared review slice under _session_home_scope (an
  unlisted subcommand keeps the 4018 routing refusal); _cmd_memory
  routes the shared handler with the live agent's store or a fresh
  on-disk one and a profile-scoped config setter. Gate-off still
  answers while staged writes exist (gateway parity, never stranded).
- tools/write_approval: _staged()'s review hint is surface-aware —
  slash-capable surfaces (CLI/TUI/desktop/chat gateway) keep
  "<command> pending"; headless contexts (cron, kanban, api_server,
  webhook delivery) get the pending directory path instead of a
  command nobody can deliver.

Batch skill writes also rendered "(batch on ...)" for /skills diff —
skill_pending_diff had no action='batch' case; ops now diff in order
against what earlier ops of the same batch left (finn763's #123861).

Supersedes #98794 (registry-driven desktop_subcommands + resolver
allowlist), #105773, #118821 (command.dispatch handlers + arg
threading), #123861 (batch diff renderer). #107799 is superseded by
the same subcommand-aware resolution.

Co-authored-by: foma-agent <foma@bokonon.ai>
Co-authored-by: Jason Havenaar <35441875+jasonHav@users.noreply.github.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: chelsealong <chelsealong@126.com>
2026-09-30 23:05:43 -05:00
liuhao1024
8f628e2e81 fix(plugins): date a treeless checkout by its files when history is unreadable
The no-lazy-fetch probe made ``in_tree_catalog_time()`` return ``None`` on
every tree:0 checkout — indistinguishable from "not a git checkout". That
``None`` feeds the frozen-copy rule in ``_prefer_in_tree_entry`` (no time,
no versions ⇒ live wins), so a live cache fetched before ``hermes update``
re-pinned the old sha on exactly the installs that had just bumped it —
the very regression ``load_catalog_live()``'s contract forbids (#125716
review).

On a git checkout whose path history is unreadable, fall back to the
newest checked-out catalog file's mtime: the checkout ``hermes update``
just wrote outranks the pre-bump doc again, while a checkout whose catalog
files predate the doc still defers to it. ``None`` now means only "no
usable in-tree time" (release installs, never-checked-out clones).
2026-09-30 22:30:18 -05:00
liuhao1024
f99518333d fix(plugins): keep the catalog date probe off lazy fetches on treeless clones
in_tree_catalog_time() dates the checkout with a pathspec'd `git log`,
which opens every commit's tree. On a tree:0 partial clone (the layout
`hermes update` produces) those trees are absent, so the probe lazy-fetches
each one against the remote: measured 40-80s per plugins.manage list,
past the desktop's 30s RPC limit (#125683, #123446). Scope
GIT_NO_LAZY_FETCH=1 to the probe so missing trees fail fast into the
documented None fallback, and resolve the pins/versions/titles maps from
ONE load_catalog_live() pass instead of three in _plugin_rows().
2026-09-30 22:30:18 -05:00
Brooklyn Nicholson
4b707a0f2b fix(cli): cancel /prompt compose on editor failure instead of shell fallback
_compose_in_editor fell back to shell=True with the raw $EDITOR value
unquoted when the argv invocation raised, so a crafted EDITOR value ran as
shell code. Drop the fallback: an unargv-splittable or unrunnable editor,
or a nonzero editor exit (which may leave abandoned buffer text), now
cancels the compose cleanly.

Co-authored-by: Adolanium <94890352+Adolanium@users.noreply.github.com>
2026-09-30 22:24:00 -05:00
Brooklyn Nicholson
7ee2840c06 fix(desktop): pin each voice operation to one owner and fence stale STT continuations
- Resolve the (connection, profile) once per dictation / conversation and use
  it for the warm-up, the transcription (client-direct config + relay) and the
  release; untagged halves stay untagged instead of re-reading ambient scope.
- Recorder and conversation drop any continuation (warm-up, transcript,
  submit) that settles after unmount, end(), or a newer start().
- syncSttLease: every settled-apart acquire is sent (no client freshness
  window); an acquire coalesces only onto a queued acquire that is still the
  latest intent, so on/off/on ends acquired and waits for its own acquire.
- Idle-unload watcher applies the stt.local policy of the caller that last
  loaded/used the shared model, re-read in that caller's context each cycle.
2026-09-30 15:16:40 -05:00
Brooklyn Nicholson
e40f293796 fix(desktop): carry STT warm-up readiness into transcription and enforce lease ownership
Addresses the four unresolved P2 review threads on #106002.

- Readiness barrier: the recorder stop() and the conversation loop turn
  paths await the warm-up issued at mic-open before submitting audio to
  transcription, so a cold model load settles inside the lease request own
  budget instead of eating the transcription request 180s decode deadline
  (#105955 review).
- Warm-only idle eviction: warm_stt_provider() registers the warmed model
  with the existing idle-unload watcher, so stt.local.unload_after_idle_seconds
  is enforced for a model loaded only by warm-up (mic open then cancel) and on
  reload after a prior watcher exited.
- Readiness is validated, never remembered: the lease dedupe now separates
  registration from warm-up validation — an HTTP 200 action "error" body
  clears readiness and retries the next acquire, and a validated warm-up is
  only trusted for a short window so an idle eviction mid-conversation
  re-warms on the next listening start. Cloud noop stays a valid outcome.
- Lease ownership: queued operations and dedupe state key on the owning
  connection/profile captured at intent time and carried through the wire
  call (setSttLease gains the same owner scope speakText uses), so switching
  gateway/profile mid-flight releases the lease on the backend that holds it
  and a new owner acquire warms its own backend.
2026-09-30 15:16:40 -05:00
brooklyn!
bda33b601d Merge pull request #128992 from NousResearch/bb/bot-focus
fix(bot-mode): make bot-to-bot messaging, Bot Chat opens, and lease handoff actually work
2026-09-30 14:40:57 -05:00
Brooklyn Nicholson
4546cd40ec fix(sessions): show a steer as the user's own words in REST history
session.resume already unwraps steer rows; the REST projection the
desktop hydrates from after a turn did not, so a correction sent during a
tool swapped its bubble for the model-facing OUT-OF-BAND marker text.
2026-09-30 12:01:29 -05:00
Brooklyn Nicholson
e0c0c8c68c test: keep one ordering contract for single-query finalize
The full-order test in test_oneshot_completion_linger plus the real-DB,
real-lease successor test pin the finalize order; drop the lease-before-linger
and memory-before-successor tests that re-asserted it, and let the wait-failure
test assert what it is about.
2026-09-30 10:32:09 -05:00
Zheqing Zeng
82c76be695 fix: run memory-provider session finalization inside ownership too
Review feedback (Enough1122): _run_cleanup -> _shutdown_agent_memory_provider
-> MemoryManager.on_session_end is session-keyed settlement (OpenViking commits
sessions/{sid} synchronously; Supermemory flushes pending turns stamped with the
session id), and at the previous head it ran AFTER the lease release — a
successor admitted during the linger could be mid-turn on the session while the
predecessor committed its remote memory session underneath.

Move the shutdown into phase 1 (after the durable flush — #88583's failure
ordering — before the finalize hook's slot in ownership). It is idempotent
(_memory_provider_shutdown), so _run_cleanup's post-release repeat is a no-op.

Also correct the docstring: phase-1 settlement is best-effort, not a proven
invariant — failures are logged and the release still happens, because an
exited process pinning the lease is worse than a dangling row (session_recovery
reaps stale rows).

Adds a discriminating test: real lease, recording agent stub, successor
acquires during the linger — memory_commit must precede release precedes
successor_acquired. Order tests now pin flush → finalize → memory → release →
linger → teardown.

(cherry picked from commit 06dc026126a473a08f182014d38f5213a842c8a3)
2026-09-30 09:57:20 -05:00
Zheqing Zeng
1d39a2d3d3 fix: settle session-owned state before releasing the lease
Review feedback (andrexibiza): the previous revision released the lease before
_flush_one_shot_session_store, whose unconditional end_session (WHERE id=? AND
ended_at IS NULL, no owner fence) could stamp a successor's open interval when
the successor acquired and reopened the session during the lingering process's
post-release flush.

New order: session settlement (flush + finalize hook) → lease release →
process-only linger + teardown. Early admission is preserved (release still
precedes the linger); the handoff race is closed because the predecessor has
no session-owned writes left once it releases.

Adds a two-owner regression test (real SessionDB + real lease): the linger is
patched to play a successor that acquires and reopens mid-linger; the
successor's open interval must survive the predecessor's finalize. Verified
discriminating against BOTH unsafe orders (previous PR head and origin/main).

(cherry picked from commit e0b4a34a981f6d2bbdcc97be977e64e867987b6a)
2026-09-30 09:57:16 -05:00
Zheqing Zeng
734ac474e5 fix(cli): release the one-shot session lease before the exit linger
A one-shot CLI run (hermes chat -Q, the Bot Mode delivery transport)
claimed the session lease at turn start and released it only via
atexit/finalize-tail, after the bounded exit linger for
notify_on_complete children. A run whose turn had ENDED but whose
process lingered alive-but-idle (observed: 434s drain wait, 0.016s CPU
over 20 min wall) kept the lease valid — its liveness check is
(pid, create_time) — so every new delivery into the session was refused
("Refused active session") for minutes after the turn finished, and
each refused delivery spawned its own lingering process tree
(#118826).

_finalize_single_query runs only after every turn of the run (main
turn, kanban goal loop, notify-completion follow-ups) is done, and none
of its remaining steps — the linger wait, the durable flush, cleanup —
runs a turn. The lease means "a turn may run on this session", so
release it FIRST: the boundary the one-shot path already declares ("the
exit linger is NOT part of the spawner's delivery", #113608). Cleanup
failure propagation is unchanged, and a cleanup failure can now never
skip the release.

(cherry picked from commit 438c68b80898efb9a287d39da32a94b772316d4e)
2026-09-30 09:57:11 -05:00
Austin Pickett
9c5da5a1c7 fix(secret-sources): revoke a removed source's value from os.environ
Removing or disabling the last plugin secret source rebuilt the per-home
snapshot but left the value it injected in os.environ, so single-profile
get_secret() kept serving it through its os.environ fallback until
restart.

The process-global apply path now records what each source wrote (and
the value it replaced). A later pass revokes writes whose source is no
longer registered/enabled, or that an enabled source stopped supplying,
but only while os.environ still holds exactly the written value: a value
another owner replaced (profile .env, managed .env, shell, later source)
stays. A source whose fetch failed keeps its last value (fail-open).
2026-09-29 22:57:48 -04:00
kshitijk4poor
27bbbda5c3 test(cron): keep three invariants for the routed-fire isolation
The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:

- a routed desktop-ticker fire runs under multiplex semantics for exactly
  its scope: a scope miss returns None instead of the launch credential and
  the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
  .env defined it or a launch external source supplied it (applied or lost
  to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
  profile's own value.

Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.

(cherry picked from commit 042b9830b4)
(cherry picked from commit 0b82557fbdbfa6968b2b795e4e63d484ffb582b1)
2026-09-29 22:57:48 -04:00
John Paul Soliva
65bd3eb336 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
(cherry picked from commit 62b4488cb5)
(cherry picked from commit d3319102a7e4dcce35ec2676fe22aed41613e893)
2026-09-29 22:57:48 -04:00
John Paul Soliva
66e3090a4f fix(cron): close three launch-residue leaks into a routed no_agent child
Review findings on d8c467f223, each reproduced through its production path.

Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.

Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.

Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).

Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.

(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
(cherry picked from commit 9d7de6c140)
(cherry picked from commit b0beb731b18edd884e36a30ab91c91687bc7a965)
2026-09-29 22:57:48 -04:00
DavidMetcalfe
57a22675ef fix: only unresolved gateway rows make successor evidence ambiguous
Review round 3 (Gemini 3.8 Flash): counting every planned gateway row per profile
made a dirty baseline exit 1 even when the fleet was fully accounted for. With an
orphan row the restart phase killed (killed_pids) plus an active row that was
restarted, the killed row's own verdict ("stopped") was already in hand, yet its
count still marked the profile ambiguous and blocked credit for the row that did
restart.

The baseline count now skips rows bookkeeping already resolves (killed /
relaunched / name-matched) and counts only rows that still need a verdict, so the
ambiguity guard fires on genuine ambiguity instead of on any multi-row plan. The
debug message no longer claims a successor exists when the probe returned none.
2026-09-29 21:50:16 -05:00
DavidMetcalfe
2594c5786f refactor: count planned gateways by kind, document the deliberate limits
Review round 3 (@GPT-OSS 120B): `kind not in _SERVE_KINDS` was read as
"excludes gateways" — the set is ("serve", "dashboard"), so the behaviour was
already gateway-only (a profile with a gateway plus a serve and a dashboard row
still credits the gateway), but the intent is now explicit: `kind == "gateway"`.
The served-kind test covers both non-gateway kinds for one profile to pin that.

Also documents what this fallback deliberately does NOT do, so the next reader
does not have to re-derive it:
- ambiguity resolves conservatively (pairing a successor to the planned process
  needs a start-time identity neither the plan record nor the fleet row carries);
- credit is not gated on restart bookkeeping (the reported failure already had
  bookkeeping; a self-respawned gateway has none, while the fleet row still
  proves the successor runs the new code).
2026-09-29 21:50:16 -05:00
DavidMetcalfe
f42fd748c8 fix: require an unambiguous baseline for successor-credited gateway restarts
Review (@ehz0ah): one live successor cannot prove that two planned same-profile
gateways both restarted. `_collect_gateway_runtimes()` de-duplicates by PID, not
by profile, so duplicate/orphan states can put two `coder` rows in the plan —
and while the post-restart fleet probe publishes at most one row per profile,
both planned rows were credited from that single successor, letting an untouched
sibling vanish behind it and suppressing the unaccounted tripwire.

Successor credit now requires the profile to hold exactly ONE planned gateway
runtime (baseline identity); an ambiguous profile keeps the name-path verdict and
stays armed. Logged at debug so the reason is visible.

Reported repro, before -> after:
  plan=[coder 76508, coder 76795], live_gateway_pids={"coder": {76796}},
  empty bookkeeping
  before: ['restarted', 'restarted']        tripwire: False
  after:  ['unaccounted', 'unaccounted']    tripwire: True
2026-09-29 21:50:16 -05:00
DavidMetcalfe
b7bb7685d3 fix: log missing successor evidence for a planned gateway
Review feedback (@kvnloo): when the post-restart fleet probe has no row for a
planned gateway's profile, the incarnation fallback silently does nothing and
the tripwire output reads identically to a genuinely missed restart. Log the
missing evidence at debug so the next report is diagnosable:

  No post-restart gateway evidence for profile 'coder' (planned pid 76508 via
  launchd) — reconciling on restart bookkeeping alone

Also make the call-site comment name the snapshot dependency the evidence
rests on (the same note #110343 carries at its wiring point).
2026-09-29 21:50:16 -05:00
DavidMetcalfe
d5c8097909 fix(update): credit a gateway restart by successor incarnation, not service name
`match_runtime_outcomes()` can only bridge a planned gateway runtime and the
restart bookkeeping by SERVICE NAME, and the two disagree whenever a service
serves a profile its name does not encode:

- macOS: the root-home LaunchAgent `ai.hermes.gateway` honours the install's
  sticky active profile (`hermes profile use <name>`), so the process declares
  `profile: coder` on its control socket while its label encodes the install
  root. `_gateway_service_matches_profile("coder", "ai.hermes.gateway")` is
  False for every naming rule, so the planned runtime is reported
  `unaccounted` → `restart.incomplete` → exit 1 after a restart that verified
  clean (fleet row: new PID, new SHA). The leftover `fleet_restart_pending`
  marker then warns on every later invocation.
- Linux: a hash-suffixed unit for a custom HERMES_HOME has the same shape of
  problem (#110238), which the open name-matching PRs only paper over for the
  default profile.

Reconcile gateways on incarnation evidence too — the rule serve/dashboard
runtimes already use via `stale_serve_pids`: the post-restart fleet snapshot
(control-socket identity per profile) is passed in as `live_gateway_pids`, and
a planned PID that is gone while a gateway answers for the same profile counts
as `restarted`. Requiring a successor keeps the tripwire honest: no successor,
or the planned PID still answering, stays `unaccounted` and still exits 1.
`down` rows are excluded from the evidence — they report the pre-restart PID.

Fixes #110333
2026-09-29 21:50:16 -05:00
Brooklyn Nicholson
3ec91de490 fix(logs): parse the update/handoff logs' ISO-8601 stamps; cover them in the LOG_FILES stamp contract
- _TS_RE now accepts the T-separated stamps update.log and
  desktop-update-handoff.log actually write (posix.sh date +%Y-%m-%dT%H:%M:%S%z,
  windows.ps1 yyyy-MM-ddTHH:mm:ssK, the update banner), so --since filtering
  works for hermes logs update/handoff instead of silently dropping everything.
- test_logs' LOG_FILES stamp contract gains real writer samples for update/handoff.
- Replace the PR's source-shape (inspect.getsource) upload tests with behavioral
  ones (monkeypatched uploader, --local capture); the upload/--local loops stay
  data-driven over bundle.items() as on main.
2026-09-29 21:48:25 -05:00
Halldrix
60f9d2abb1 fix(debug-share): capture update.log and desktop-update-handoff.log
When a Desktop-driven update fails at the Electron rebuild (updater
exit 6, 'Code and dependencies updated, but the Desktop app REBUILD
FAILED'), the only artifacts holding the root cause are
~/.hermes/logs/update.log (full stdout/stderr mirror of hermes update,
written by _UpdateOutputStream) and ~/.hermes/logs/desktop-update-handoff.log
(hand-off stages including the 'desktop --force-build --build-only' retry).
Both the updater's failure messages and the Desktop error box direct
users to 'hermes debug share', but the share captured neither file:
LOG_FILES and the debug.py capture chain covered only the classic
five logs, so update-failure issues (e.g. #100874) arrived with the
diagnostics that matter most missing.

Additionally, 'hermes logs list' (a directory scan) already listed
both files while LOG_FILES had no readable key for them, so
'hermes logs update' failed with 'Unknown log' against a file the
CLI itself displays.

Wire both files through the whole chain:
- hermes_cli/logs.py: add 'update' and 'handoff' LOG_FILES keys
  (fixes list-vs-read inconsistency; tail/follow/filters come free).
- hermes_cli/debug.py: _capture_default_log_snapshots,
  collect_debug_report (512 KB snapshot cap and the 100-line tail
  budget both apply; a hand-built snapshot dict without the new keys
  degrades to the old five-section report instead of KeyError),
  collect_share_bundle full logs, the paste-upload label list in
  build_debug_share, and the --local print list. The --nous bundle
  and privacy notice inherit the new keys; force-redaction applies
  via _capture_log_snapshot as for every other log.
- hermes_cli/subcommands/logs.py + website docs: name the new logs.

Tests (tests/hermes_cli/test_debug_share_update_logs.py): 12 new
tests — snapshot keys present with real content, missing files
degrade to the honest note without raising, report sections appear,
bundle keys omitted when absent, upload/--local label lists include
the new logs, LOG_FILES keys resolve, tail_log end-to-end. Sabotage
pass: 10 of 12 fail without the fix.
2026-09-29 21:48:25 -05:00
Brooklyn Nicholson
82bbc694b5 fix(desktop,gateway): proxy client-unreachable remote images through the gateway
An inline agent-generated image (FAL CDN URL) that the Desktop client
cannot reach renders as a collapsed frame with no usable fallback
(#74564): resolveMediaDisplaySrc passes inline https sources through
untouched, and a failed <img> load only records the broken-key.

- Gateway: GET /api/media/proxy fetches an allowlisted remote image URL
  on the client's behalf (the gateway generated the image and holds a
  working route to the CDN) and returns the same data_url shape as
  /api/media. Host-allowlist (fal.media/fal.run/storage.googleapis.com
  and subdomains), 20s timeout, image content-type check, 25MB cap.
- Renderer: gatewayImageProxyDataUrl wraps the new endpoint with
  owner-connection/profile pinning; useMediaImage retries a failed
  inline https load ONCE through the proxy and swaps in the returned
  data URL. The retry is latched per source: a failing proxied data URL
  or a remount never loops, and non-https sources never reach it.
2026-09-29 21:39:53 -05:00
Brooklyn Nicholson
063de859dd fix(auth): accept provider-verified session tokens in gated WS auth
Gated mode accepted only browser-minted ws-tickets and server-internal
credentials, but a token-mode Remote desktop connection bakes its
stored session token into ?token= (buildGatewayWsUrl) — so the WS
upgrade to /api/ws 403'd forever against a password-gated server even
though the same token authenticates REST fine. The REST bearer leg
already verifies session tokens via _verify_access_token; the WS leg
was the only inconsistent gate.

Gated ?token= is now verified against the dashboard auth session
providers (the same verify_session seam), stamps {user_id, provider}
identity, and audits success/failure (TOKEN_AUTH_SUCCESS/FAILURE). The
legacy in-process _SESSION_TOKEN stays rejected; a ProviderError (all
providers unreachable) rejects with token_unavailable instead of
crashing the upgrade; the bot-desktop display-ticket guard is kept.

Co-authored-by: salch-cred <141555468+salch-cred@users.noreply.github.com>
2026-09-29 21:39:17 -05:00
Brooklyn Nicholson
38d3dbce4a test(update): adapt handoff regression to the other_pid fixture 2026-09-29 20:58:59 -05:00
Jack Lau
e28991e9c8 fix(update): decide ancestry link by link so an unreadable /proc/1 cannot hide the orchestrator
`_is_ancestor_pid` asked psutil for the whole parent chain at once:

    any(parent.pid == pid for parent in psutil.Process().parents())

`Process.parents()` is eager: it walks all the way to the lowest pid
*before* returning a list, so the generator expression cannot
short-circuit when it reaches the matching ancestor. Its per-link
`parent()` tolerates only `NoSuchProcess`, so any process we may not
inspect anywhere above us raises `AccessDenied` out of the whole call,
and the `except Exception` guard discards every ancestor already found,
including the orchestrator one link down.

Under firejail with `ptrace_scope=1`, and in hardened containers,
`/proc/1` is exactly that unreadable process. The hand-off child then
failed to recognize its own parent's fresh marker, `UpdateLock.acquire`
treated a live orchestrator as a foreign concurrent updater, and the
in-app update exited 2 ("Another Hermes update is already running",
naming the hand-off script's own pid) on every attempt, indefinitely.

Walk the chain one link at a time and test each ancestor as it is
discovered, so a failure above a match can no longer erase it. A
failure encountered before any match still returns False, keeping the
conservative refusal for genuinely unprovable ancestry. The walk is
bounded by `_MAX_ANCESTRY_DEPTH` and skips a repeated pid so an
unexpected chain cannot spin.

Fixes #87514
2026-09-29 20:58:59 -05:00
Brooklyn Nicholson
55f616694c fix(update): restore cron agent-job prompts degraded during the update window
A writer in the update's mutation window replaced every agent-job
prompt with the job's own name while the job count stayed identical,
so the count-based restore_cron_jobs_if_emptied net passed the loss
undetected (#82990). Add a field-level safety net mirroring the config
model-settings net: restore only the prompt field of a live agent job
whose id matches a snapshot job and whose live prompt is blank or
collapsed to the job name — a legitimate user edit that merely
differs is never stomped. Wired into the post-migration restore
sequence for the active profile and every sibling profile.
2026-09-29 20:58:24 -05:00
finn763
d8dba2e687 fix(config): keep built-in model validation for settings-only provider blocks
A providers.<slug> block with no base_url/url/api of its own is settings,
not a user-defined endpoint. Only remap to the custom: validation branch
when the user-config entry declares a base_url; otherwise openai-codex and
other listing-less built-ins keep their normal catalog path.

Closes #120020
2026-09-29 20:57:57 -05:00
Bardioc1977
f284a05f1d fix(providers): a timeout-only providers.<name> block no longer shadows the builtin
providers.bedrock: {stale_timeout_seconds: 600} is the documented way to tune a
built-in provider's per-call timeout (agent/turn_recovery.py,
agent/thinking_timeout_guidance.py). resolve_user_provider() treated ANY dict
under providers.<name> as a full custom-endpoint definition, even with no
api/url/base_url — producing an empty openai_chat/api_key stub that shadowed
the real built-in overlay (bedrock_converse transport, aws_sdk auth).

For Bedrock this routed every /model switch through the generic
custom-endpoint /models probe against the Bedrock runtime host, which has no
such endpoint: the switch intermittently failed with "could not reach this
custom endpoint's model listing" depending on how fast the (irrelevant) probe
timed out.

resolve_user_provider() now requires a real endpoint (api/url/base_url) before
resolving an entry; a tuning-only block falls through to resolve_provider_full's
next rung, which picks up the built-in overlay correctly.
2026-09-29 20:57:57 -05:00
Brooklyn Nicholson
8b4b440ef3 fix(desktop,pm): stop the boot parking on a failed update receipt and bound completion retries (#122206)
Three gaps stranded the app through a legitimately long or failed
source-update completion:

- update-gate: a fourth gate signal 'failed-receipt' (hasFailedReceipt on
  UpdateGateDeps, optional) reports a FINISHED, failed update even when a
  stale marker survives; waitForUpdateClearance gains abandonOn so the
  boot path stops parking the full 20-minute budget on a receipt that
  already says failed and boots the current build instead.
- backend-ready: while the backend prints venv_sync's source-completion
  banners the 90s port-announce deadline is re-armed (5-min grace,
  30-min total cap, real wall clock) so a healthy multi-minute repair is
  not killed and re-run on every boot.
- venv_sync: relaunch-driven completion-tail retries are bounded — an
  attempts record beside the pending marker keeps the first retry
  immediate (the historical flaky-tail case), backs off after 2
  consecutive failures (15-min window), and caps at 6 attempts, past
  which the launch keeps the marker, skips the tail, and points at
  'hermes update'; success clears both the marker and the record.
2026-09-29 20:57:51 -05:00
Brooklyn Nicholson
eb8a1580c6 serve: retire desktop-owned local backends on code skew; guard every model-mutating path (#99859)
The code-skew watchdog only covered SSH-isolated backends; Desktop-owned
local serves share the same 'updater never restarts me' property, so they
served stale modules until the app happened to restart them (R1). Extend
start_code_skew_watchdog wiring to is_desktop_owned_backend() backends.
Centralize the skew refusal: POST /api/model/set and the tui_gateway
model.options / model.save_key RPCs now refuse with the same restart
message the picker and the gateway /model switch already return (R2).
2026-09-29 20:57:23 -05:00
Brooklyn Nicholson
0590ab25fe fix(marker): zombie-aware stale-marker self-heal parity across all readers
Rust live_marker_owner now REMOVES a stale marker (dead pid, past ceiling,
unparseable) on read — previously stale bytes were only ignored, so a crashed
updater whose Drop never ran refused every later acquire until the 20-minute
ceiling (#77259). pid_is_alive gains pid-0 and zombie handling (Linux /proc
state, macOS ps stat), closing the kill(pid,0) false-positive that kept a
dead-but-unreaped updater 'live'.

Python _pid_is_running gains the same zombie awareness via a new
_process_state probe (Linux /proc, macOS ps), which read_live_update inherits
through update_lock._pid_alive — so the CLI gate self-heals a zombie-owned
marker in seconds instead of 20 minutes (#120635, #125932).

Electron readLiveUpdateMarker consults a new posixProcessState probe
(injectable, fail-open to the signal-0 verdict) so the desktop boot gate no
longer parks on a zombie-owned marker.

Rust fix cherry-picked from #77885 (authorship kept); Python and Electron
are new parity work.
2026-09-29 20:57:17 -05:00
Brooklyn Nicholson
5c66b2e02f test(desktop): make the lock-before-handoff test hermetic
The test passed on the author's macOS but exited 1 on CI: with
--skip-build the build lock is never acquired (the probe asserted a
release that never happened), and the unmocked platform launch fixups
ran against the fake empty executable — on Linux,
_packaged_desktop_launch_command's sandbox helper is missing and sudo
is unavailable, so cmd_gui exits 1 before the handoff.

- Run the real acquire path (no --skip-build) so the lock is genuinely
  held, then probe its release at the Electron handoff.
- Pin HERMES_HOME to tmp so the checkout-keyed lock path is isolated
  from the runner's home.
- Stub _desktop_linux_sandbox_fixup and _installed_desktop_launch_target:
  environment-dependent platform fixups are not this test's subject.
2026-09-29 19:49:52 -05:00
Brooklyn Nicholson
749ef18b16 fix(desktop): serialize the mutable build preflight behind a cross-process lock
Two concurrent `hermes desktop`/`hermes update` invocations mutate the same
checkout-scoped node_modules and apps/desktop/release with no serialization:
the check-then-build preflight (_desktop_build_needed → npm install → pack) runs
unguarded, so the loser trips over the winner's half-installed tree
(ENOTEMPTY) or packs against a missing tsc.

Lock the whole mutable window behind DesktopBuildLock, a checkout-keyed flock
under the profile-common Hermes root (kept out of the checkout so read-only
prebuilt launches still work, released by the OS if a builder crashes):

- `hermes desktop` acquires it before the freshness check and refuses a
  contended build with an explicit message (exit 2) instead of racing.
- `hermes update`'s desktop rebuild (build_update_products) waits for an
  in-flight build rather than failing the update.
- the lock is released before the Electron handoff (source, packaged, and
  --build-only paths), so an open Desktop window never blocks a rebuild,
  and after a build failure so the next run isn't deadlocked.

Supersedes #94101 (lock design; its main.py wiring no longer applies after the
cmd_gui extraction) and #96581 (wait/exit-2 UX; not checkout-keyed).

Co-authored-by: worlldz <cryptoworlldz@gmail.com>
Co-authored-by: lepetitprince716-prog <lepetitprince716-prog@users.noreply.github.com>
2026-09-29 19:49:52 -05:00
Brooklyn Nicholson
6205ae1e19 test(desktop): keep the packaged-launch test out of the signature probe
The downgrade guard added a codesign probe inside the cmd_gui launch path,
which leaked /usr/bin/codesign calls into
test_packaged_launch_opens_the_refreshed_installed_app's captured
subprocess.run. The test pins the LAUNCH contract only — with no codesign
binary the signature checks are off and the swap proceeds, keeping it
focused (signature policy is pinned separately in
test_desktop_install_after_update.py).
2026-09-29 19:43:22 -05:00
Brooklyn Nicholson
65b3f43c02 fix(desktop): restore the quarantine-xattr strip before signing; pin the reconciled refusal contract
The rebase of the signing-downgrade guard dropped the unconditional
`xattr -cr` from _desktop_macos_relaunchable_fixup (the commit that claimed
to reorder it after the identity decision deleted it instead), so locally
built bundles kept their quarantine attributes and still tripped the
'Hermes is damaged' Gatekeeper path the strip exists to prevent. Restore it
unconditionally ahead of the signing attempt: in the stage-and-swap flow it
targets the STAGED bundle (discarded when the fixup refuses), and in-place it
is harmless hygiene on a bundle whose signature is never replaced.

test_relaunchable_fixup_configured_identity_failure_never_falls_back_to_adhoc
now models the reconciled policy from #123748/#121857: it pins a
publisher-signed (Team ID) replace-target — refusal, no ad-hoc, remedy named —
while the locally-signed target's identifier-pinned ad-hoc retry stays covered
by test_relaunchable_fixup_failed_identity_uses_pinned_adhoc_before_legacy.
2026-09-29 19:43:22 -05:00
Yuan Li
4bc126d038 fix(desktop): wire fixup refusal into promotion; name the remedy; order xattr after the identity decision
Review follow-up on the macOS signing-downgrade guard:

- _promote_staged_desktop_app no longer discards the fixup's False: a
  refused signing (configured identity failed / codesign missing) now
  discards staging and raises the same previous-app-kept RuntimeError
  the integrity check uses, so the live bundle is never swapped for one
  whose signature was never established. Regression drives the real
  promotion path end-to-end (the earlier tests only called the fixup
  directly).
- Every downgrade-refusal message now names the next action (publisher
  identity for the locally-signed branch, the desktop signing setting
  for team mismatch, bundle identifier alignment, codesign -vvv for
  strict-verification failures), matching the running-app message.
- The fixup reads the identity decision BEFORE the quarantine-xattr
  strip: a failed configured attempt no longer strips attributes off a
  bundle it then declines to modify.
2026-09-29 19:43:22 -05:00
Yuan Li
cf0291065c test(desktop): pin macOS signing-downgrade guard and no-ad-hoc-fallback contract
Red on upstream/main, green with the fix:

- _install_rebuilt_macos_bundles refuses to replace a Team-ID-signed
  install with a locally signed build, a different Team ID, a different
  bundle identifier, or an unverifiable rebuild; matching publisher
  identities and ad-hoc installs keep swapping.
- _desktop_macos_relaunchable_fixup keeps the existing signature when the
  configured desktop.macos_signing_identity fails instead of degrading
  to legacy ad-hoc.
2026-09-29 19:43:22 -05:00
686f6c61
b07bfd688a fix(desktop): pin ad-hoc signing when the macOS identity fails
A locked login keychain makes the configured identity unusable during an
SSH update. Retry the identifier-pinned signer before the cdhash-only
legacy deep sign so TCC grants and entitlements survive.
2026-09-29 19:43:22 -05:00
ryrenz
60800e87ca fix(desktop): preserve Electron Framework JIT entitlements
Nested frameworks were signed with no entitlements, so under the hardened
runtime Electron Framework lost com.apple.security.cs.allow-jit and V8
crashed on launch. Apply the inherit plist to every nested bundle, matching
what electron-builder's afterSign step does in a real build.
2026-09-29 19:43:22 -05:00
JoaoMarcos44
16b214e1a1 fix(config): make destructive replacements explicit
atomic_config_write() now refuses deletion-by-omission at any mapping depth:
atomic_roundtrip_yaml_save treats keys absent from the payload as deleted
recursively, while the old docstring described the call as a merge, so a
partial dict could silently erase unrelated config (#125107). A new
atomic_config_replace() names the deliberate full-state replacement contract
explicitly, and the owners that intentionally remove state (save_config,
provider reset/switch, credential mirror scrub, doctor key migration, profile
channel stripping, TUI full-state saves) route through it. Additive writers
stay on the guarded atomic_config_write.

Salvages #125152 (identical patch, cherry-picked; conflict in
_write_user_config resolved to pass atomic_config_replace through the new
telemetry wrapper) and carries the executable bit main expects on
scripts/check_config_yaml_writers.py.

Fixes #125107

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-29 20:24:19 -04:00
Brooklyn Nicholson
3ebbaf5243 test(serve): the orphan watchdog's graceful exit dies by SIGTERM, not os._exit(0)
The live-macOS watchdog test pinned returncode == 0, which was the old
os._exit(0) watchdog's signature. The graceful-shutdown change (#108601)
raises SIGTERM instead; the chaining exit-flush handler restores the
default disposition and re-raises, so the process dies by SIGTERM with
returncode -15 — the same exit a SIGTERM from a live desktop produces.
Pin that contract (and keep the prompt-reap window assertion).
2026-09-29 19:06:35 -05:00
Vadim Comanescu
1b64a72e25 fix(serve): shut down gracefully on parent death instead of os._exit
The parent-death watchdog ended its loop with `os._exit(0)`, which skips
signal handlers and `atexit`. The same serve process installs chaining
SIGTERM/SIGINT handlers that persist in-memory transcripts to state.db
before shutdown (#96095, `install_exit_flush_signal_handlers`), so the one
exit path that fires when the desktop dies uncleanly was the one path that
bypassed the flush those handlers exist for. It also skipped every `atexit`
hook reachable in a serve process (code kernels, browser tool, LSP, relay).

Raise SIGTERM in-process instead and back it with a daemon timer that
`os._exit(0)`s after a ceiling, the graceful-then-ceiling idiom `cli.py`'s
`_arm_exit_watchdog` already uses. The orphan reap stays guaranteed and
bounded: the ceiling sits past the 5s exit-flush budget and well inside the
old "leaks until killed" behavior. `raise_signal` is used rather than
`os.kill` because it reaches Python's handler on Windows too.

Tests: parent loss now runs this process's SIGTERM handler and takes no
`os._exit`; the ceiling timer is armed as a daemon bounded hard exit.
2026-09-29 19:06:35 -05:00
Brooklyn Nicholson
b9b1328b7d fix(cli,desktop): hide every console-subsystem child spawn on Windows
A console-less parent (pythonw backend, detached daemons) allocates a visible
console per bare spawn, so Windows startup flashes ~7-8 console windows
(#117781) and the GPU statusbar flashes two more every 5s (#101895). Threads
creationflags=windows_hide_flags() through every hermes_cli spawn site the
desktop reaches: gitlock (tasklist probe + git plumbing), the scheduled-task
PowerShell probe, the updater's _git_run/_no_prompt_git_kwargs/_GIT_TEXT_KW
runners, the safe-directory config probe, and consolidates the statusbar's two
nvidia-smi reads per poll into one cached hidden query. The statusbar also
stops polling while the document is hidden and resumes on visibilitychange
(#120262).

Supersedes #117805, #117833, #102047, #120285 (thanks @JoaoMarcos44, @Finn763,
@monerostar, @szzhoujiarui).
2026-09-29 18:49:58 -05:00
Brooklyn Nicholson
d2b9a49e63 fix(desktop): guarantee an exit path for a fullscreened preview pane
A preview guest that HTML5-fullscreens itself owns all input: the host
renderer can't see into the out-of-process webview, the macOS-only Close
menu and HUD chords aren't reachable, and Wayland has no xdotool/wmctrl to
break out from a terminal — so a page that swallows Esc locks the display
(#97213).

Three layers, all routed before the guest sees the key or fully outside it:

- before-input-event on every webview guest: Esc exits the host window's
  fullscreen (gated on the fullscreen state so normal-mode panes keep Esc),
  Ctrl/Cmd+Shift+W closes the pane via the existing close-preview channel
- hermes://close-preview deep link handled in the main process: restores
  and focuses the window, exits fullscreen, closes the pane
- `hermes desktop --close-preview` CLI flag forwarded to the packaged exe,
  riding the single-instance argv so a second invocation unlocks the running
  app

The permission handlers keep auto-allowing 'fullscreen' — the exit paths
are the fix, and scoping the grant would just break video fullscreen.

Refs #97213
2026-09-29 18:48:37 -05:00
liuhao1024
d67278cd88 fix(auth): detect the lock holder for real and route permanent lock failures
Review follow-up on the auth-store lock timeout copy (#124533):

- The 'likely holder' half of the timeout message was an unconditional
  literal, asserted even when no holder exists. Holders now stamp their
  pid into the lock file on acquire (cleared on release), and the timeout
  reads it back through a liveness probe: a live foreign pid is named,
  a stale pid or an empty file stays silent.
- Permanent lock failures (ENOSYS/EOPNOTSUPP on flock-unsupported
  filesystems, bad fd) used to burn the whole deadline and then blame a
  holder that does not exist. Non-contention errnos now propagate
  immediately; EACCES stays retried because msvcrt reports both real
  contention and ACL denials as EACCES.
- The shared Nous store lock raises through the same _file_lock on the
  same init path (resolve_nous_access_token) but its message did not
  match the routed prefix, so init fell back to the /model + hermes
  setup advice this PR exists to remove. Both prefixes now route to the
  contention copy, and the Nous timeout names its lock file too.
2026-09-29 18:47:44 -05:00
finn763
df126f6a05 fix(desktop): fail fast instead of hanging when sudo has no TTY
The Linux sandbox fixup shelled out to sudo with inherited stdio, no -n
and no timeout. A .desktop/autostart/detached launch that reaches the sudo
path (user-owned chrome-sandbox, no usable user namespace) hung forever on
the password prompt. Related to #123927.

This line is new in this diff, so the guard belongs here: sys.stdin is
None when the process has no stdin at all and a closed stream raises from
isatty(), and either exception escaped as a traceback before the caller's
--no-sandbox fallback could run.
2026-09-29 18:41:02 -05:00