Commit Graph

35951 Commits

Author SHA1 Message Date
teknium1
c055eee1fe refactor(checkpoints): slim the profile-rename rekey and name it as a checkpoint_manager sibling
Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):

- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
  repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
  idempotent rekey: every step overwrites and the old project metadata is removed last, so a
  mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
  A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
  192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
  (`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
  and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
  assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.

Fixes #112973
2026-09-16 14:34:22 -07:00
NanPan
8a2c9edac6 fix(profiles): preserve checkpoint history across rename 2026-09-16 14:34:22 -07:00
teknium1
059c0269e1 refactor(status): collapse the stopped-intent ladder into one condition
Same behaviour as the cherry-picked fix, one branch instead of a nested
if/else: a not-running gateway is "stopped" when the operator's durable
``desired_state`` says so or when the retained state is not one of the two
terminal states; ``exit_reason`` is dropped only on "stopped" so a live
``startup_failed`` (desired_state=running) keeps its diagnostic.

Live probe (real /api/status, temp HERMES_HOME): stopped default profile
with retained port-conflict failure -> unscoped ``gateway_state='stopped'``,
``gateway_exit_reason=None``; control ``?profile=worker`` (desired running)
-> still ``startup_failed`` + 'telegram: token rejected'.
2026-09-16 14:34:22 -07:00
KoNit-K
aa711728e8 fix(status): honor stopped gateway intent 2026-09-16 14:34:22 -07:00
teknium1
3946949646 test(profiles): prove delete_profile releases routed log handlers end to end
Replace the mocked delete_profile test (it patched release_profile_log_handlers
and only checked the call happened) with a real-path invariant: route a record
to a fresh profile — explicitly, as the Desktop cron startup does, and through
the setup_logging(hermes_home=<profile>) adoption path agent_init takes — then
delete_profile(name, yes=True) and assert no router still holds a handler or an
open stream for that home. Both parametrizations fail on origin/main (the
router's _profile_handlers keep two open streams into the removed directory)
and pass with the release in place.

A windows_only companion asserts the directory is actually removed: on win32
those streams are ConcurrentRotatingFileHandler .__agent.lock / .__errors.lock
handles, the WinError 32 the reporter hit. The Linux test covers the same seam
because rmtree of open files succeeds on POSIX and would hide the leak.

Shape follows #112543's fail-before tests by the reporter.

Refs #112538

Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>
2026-09-16 14:34:22 -07:00
KoNit-K
43e28551bb fix(profiles): release routed log locks before deletion 2026-09-16 14:34:22 -07:00
anhtahaylove
f27c719a93 fix(gateway): resolve the mirror sessions index at call time, not import time
`gateway/mirror.py` captured `sessions.json` under the home that was live at
import. Under the multiplexed gateway one process serves every profile, so the
pre-migration fallback in `_find_session_id()` resolved every profile's lookup
against the launch profile's index — a `platform`+`chat_id` pair present in both
indexes returns the launch profile's session id, mirroring a delivery into a
session that belongs to another profile.

Only the fallback is affected: the primary `state.db` query goes through a bare
`acquire()`, which follows `get_hermes_home()` at call time and is already
profile-correct. The access is read-only, so nothing was written to the wrong
profile.

Applies the resolve pattern the rest of the repo already uses (gateway/hooks.py,
gateway/sticker_cache.py, tools/process_registry.py, tools/checkpoint_manager.py):
keep the module constant as the monkeypatch surface, snapshot it at import, and
re-resolve only when unchanged.

The added test reloads the module under a launch home and then serves a lookup
for a second profile; without the fix it returns `sess_launch` instead of
`sess_active`.

Fixes #112844
2026-09-16 14:34:22 -07:00
Aniruddha Adak
6eef5ea0bb fix: fallback activation picks the OpenCode per-model wire (muse-spark → Responses)
A `fallback_providers` entry on opencode-go / opencode-zen / opencode-free (or a
custom provider hosted on opencode.ai) landed on `/chat/completions` unless the
user pinned `api_mode:` by hand: `_fallback_api_mode_resolved` re-detected the
wire for openai-codex, Nous, Anthropic URLs, Azure, direct OpenAI, gpt-5 and
Bedrock but never consulted `opencode_model_api_mode`, which the primary
`/model` path already uses. Responses-only models (muse-spark, gpt-*, grok-*)
500'd three times per activation and the chain advanced past a healthy entry;
Anthropic-wire models (minimax, qwen) on Go were misrouted the same way.

Resolve the OpenCode family through the same predicate the custom-provider
runtime uses (`_opencode_family_for_custom`: provider name or opencode.ai host)
and return the per-model table's mode. An explicit `api_mode:` still wins (the
resolver is only called for the chat_completions default), and chat-completions
OpenCode models keep their wire.

Fixes #102148
Salvages #102229 (@aniruddhaadak80)
2026-09-16 14:24:30 -07:00
teknium1
4a1daecc18 fix: seed Responses budget escalation from the observed ceiling when no cap is configured
With max_tokens unset the escalation armed nothing, so a default-config user hitting the
provider ceiling resent the same absent budget. The usage.output_tokens of the exhausted
response is that ceiling; seed from it (else 4096). Test control swapped to the reachable
invariant (non-empty fragment keeps today's replay) per the independent review.
2026-09-16 14:23:55 -07:00
teknium1
92c24d410d fix: Responses-wire length continuation raises the output cap and drops reasoning after an empty fragment
When a Responses-API call (muse-spark, gpt-*, grok-* routes) ends with
status=incomplete / incomplete_details.reason=max_output_tokens and no visible
text, reasoning consumed the whole output budget. continue_codex_incomplete
re-sent the identical request three times: same max_output_tokens, same
reasoning effort, plus a nudge. The model re-burned the same budget each time
and the turn ended as "Codex response remained incomplete after 3 continuation
attempts" with nothing for the user (#90393; measured in #103483 at xhigh:
out_tokens == max_output_tokens, reasoning_tokens == out - 3).

The chat-completions length path already handles this with two one-shot
overrides (_ephemeral_reasoning_off, _ephemeral_max_output_tokens); reuse them:
on a budget-exhausted empty fragment set reasoning off and double the cap
(2x, 4x, capped at 32768 / the configured cap) for the next attempt, and make
_build_codex_kwargs consume the ephemeral cap it never read before. A
reasoning-only status=completed response (Codex "still thinking") is untouched.

Live repro (fake Responses provider, real loop): before 2000/high, 2000/high,
2000/high -> partial; after 2000/high, 4000/no reasoning, 8000/no reasoning.
Chat-completions control unchanged (2000->4000->8000->16000, effort none).
2026-09-16 14:23:55 -07:00
teknium1
1086cc691f chore: map actualrat1984 contributor email
The login-only noreply form lacks the id+login shape the attribution gate auto-resolves.
2026-09-16 14:23:20 -07:00
teknium1
2b4ca80f5f test: trim Responses blank-carrier coverage to two invariants
Keep the fixture-driven test on the real failing turn shape from #103483
(no invented carrier, every reasoning item followed by its function_call,
call/output pairs intact) and the control that a lone reasoning item still
gets a non-empty follower. Drop the synthetic parametrized round: the fixture
already covers that shape once per tool round, and the repo caps a fix at two
invariant tests.

Part of #103483
2026-09-16 14:23:20 -07:00
actualrat1984
f33624c58a test(responses): cover real failing-turn shape from #103483
Fixture by @cristianbdev (content-redacted replay of the failing turn):
no invented blank carrier, every reasoning followed by its
function_call, all five call/output pairings preserved.
2026-09-16 14:23:20 -07:00
actualrat1984
1c7bc44eda fix(responses): avoid invented empty assistant turns before tool replay 2026-09-16 14:23:20 -07:00
teknium1
ba8b458c5b test: trim purge-identity tests to behaviour invariants
Drop the three cases that pin trivia the invariant tests already cover
(idempotence is exercised by the routing-store test; the empty-name /
no-store guard and the default-profile refusal are one-line argument
checks). Salvage bar: invariants only.
2026-09-16 14:21:14 -07:00
xielevi
aa90c01e9a fix(profiles): report a settle-pending delete as a typed partial success
A delete can end with the profile directory already removed and its durable
session/routing identity settlement still pending. delete_profile raised a bare
RuntimeError for that state, so a caller wanting to distinguish the completed
filesystem half from a real failure had only the message to go on — and the
dashboard's DELETE /api/profiles/{name} folded it into its generic 500, making a
client read the completed delete as "delete failed" (its retry then 404'd).

The partial settlement is now the typed ProfileIdentitySettlementPending — a
RuntimeError subclass carrying profile / path / retry_command:

- the CLI's delete handler keeps its non-zero exit unchanged: the type subclasses
  RuntimeError, so its existing catch tuple covers it;
- DELETE /api/profiles/{name} catches exactly that type and answers
  200 {"ok": true, "path": ..., "identity_settled": false, "settlement_pending":
  true, "retry_command": "hermes profile purge-identity <name>"};
- a genuine filesystem failure stays a plain RuntimeError and keeps the 500 — the
  API does not catch RuntimeError broadly.

Tests: the live-multiplexer delete test now pins the typed payload (profile,
retry_command, the removed path, RuntimeError subclassing); a new endpoint test
drives the real delete chain — settle-pending answers the partial success with the
directory really gone, and an rmtree failure still answers 500. Red on base (the
endpoint answered 500 for the completed delete; the type is absent) and green with
the fix: 347 passed / 0 failed across the 10 affected files.
2026-09-16 14:21:14 -07:00
xielevi
a41552fad4 fix(profiles): purge a deleted profile's session/routing identity on delete
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:

- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
  `_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
  also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
  namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
  (`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
  in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
  writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
  `_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
  has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
  partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
  profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
  `create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
  delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
  leaves of a profile's conversation record is `delete_profile`'s business — it removes the
  profile's own home, `state.db` included.

Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
2026-09-16 14:21:14 -07:00
Teknium
c89a1cbc11 chore: map contributor email for attribution check 2026-09-16 14:20:30 -07:00
PixelAlchemist4
b978eaab43 chore(catalog): add agent-log desktop plugin 2026-09-16 14:20:30 -07:00
teknium1
f2dc875a57 docs(gateway): standalone gateways bind the launch scope once hosted activation flips the guard
Records the #112878 rule in gateway/AGENTS.md so a new standalone path gates on _standalone_launch_scope rather than the config flag.
2026-09-16 14:20:17 -07:00
teknium1
51c4b6ba9d fix(gateway): standalone gateway binds the launch profile's own scope after hosted activation
A native hosted room running a second profile calls
tui_gateway.launch_profile_policy.activate_multi_profile_hosting() inside the
messaging gateway process, so get_secret() fails closed for every unscoped read
afterwards. A gateway with gateway.multiplex_profiles: false never bound a scope
(the config flag was the only gate), so its next ordinary turn died in
_resolve_session_agent_runtime with UnscopedSecretError / "Hermes could not read
this profile's API key" until restart (#112878).

Builds on the predicate from #112884: instead of re-entering the routed-profile
scope (which rebuilds credentials from .env alone and would drop a key injected
by systemd / `op run`), the standalone branch binds the launch profile's OWN
scope — launch_secret_scope() (.env + external sources over the env frozen at
activation) plus its terminal policy — exactly what the tui_gateway already
binds for launch-profile RPC bodies. One predicate, GatewayTurnMixin
._standalone_launch_scope(), no-op while the process is single-profile.

Whole class: every standalone entry point now runs under it — the foreground /
background turn wrappers, busy, goals and heartbeat-restore paths already
routed through _profile_scope_for_source, and the primary adapter's message,
busy-session and platform-event handlers (slash commands such as /model,
/status and /compress run inside _handle_message and were equally stranded).
Cron already binds its own per-fire scope. The guard is not weakened: reads
outside any scope still fail closed and a secondary never sees the launch env.

Tests: tests/gateway/test_standalone_gateway_launch_scope.py — real activation
helper, real server-bound secondary scope entered and left, then the standalone
turn and handler resolve both the .env key and the env-injected key (red on
origin/main with UnscopedSecretError); control: no activation → nullcontext and
no scope in the handler. The #112884 test moves here in trimmed form.

Co-authored-by: Lyti4 <205342405+Lyti4@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 14:20:17 -07:00
KoNit-K
99f68e730b fix(gateway): scope standalone turns after hosted activation 2026-09-16 14:20:17 -07:00
Teknium
50520d8fd8 chore: map contributor email for attribution check 2026-09-16 14:20:00 -07:00
Alberto Pittoni
aca6c1c84a feat(catalog): add usage-stats plugin entry
Multi-provider usage/balance for the Hermes Desktop status bar (OpenCode,
OpenRouter, DeepSeek, Kimi, NovitaAI, ZAI, Alibaba, Arcee, plus
gateway-native Claude/Codex/Cursor/Nous).

Pinned at ba0652c (includes validated home directory and config file
integrity checks). 38/38 tests pass. No agent tools, hooks, or middleware.
No env vars required at runtime.
2026-09-16 14:20:00 -07:00
LisandroNahuelH
edf4e2efc1 chore(plugin-catalog): add composer-modes
One installable package (agent half + desktop half) that puts a mode selector in
the desktop composer: ask / agent / plan / debug. The mode's operating note rides
the turn's model-facing bytes via pre_llm_call, so the durable row, the bubble and
the session title keep exactly what the user typed; ask mode is enforced read-only
through pre_tool_call. No core change, no renderer rebuild.

Pinned at 4f4e1c0b36165d84a9e2db7c71516f5a0c627d72 - the v2.0.0 package commit
(gates: 103 pytest cases, the desktop-half smoke test, hermes plugins validate,
and the repo's CI).
2026-09-16 14:19:34 -07:00
Teknium
547c3d49fc chore: map contributor email for attribution check 2026-09-16 14:19:07 -07:00
Steven Leggett
3a147d8037 Pin sugar entry to the v3.10.2 release merge on main
The plugin shipped in Sugar 3.10.2 and is now merged to main
(eade16940789bc1f927af090f8a3b475e537eddd). Structural validator and
the pinned-source clone + hermes plugins validate both re-verified
against the new pin.
2026-09-16 14:19:07 -07:00
Steven Leggett
b2096cd6fa Add sugar to the plugin catalog (memory)
Sugar is a local-first persistent memory provider: SQLite on the user's
machine with project and global scopes and guideline-aware search.

- Tier: community, category: memory
- Source pinned to 37cbe10b33e08f04a5d91826475716ab5c315a49
  (roboticforce/sugar, hermes-plugin subdir)
- No standalone tools/hooks/middleware; registers a memory provider
  with three agent tools (sugar_store, sugar_search, sugar_list_recent)
- Works on hermes-agent >= 0.19 (verified against PyPI 0.19.0 and
  main 0.21.3: plugins validate passes, full provider E2E passes on
  both versions)
- No network access, no API key, no self-updating code; all data stays
  in the user's ~/.sugar/ and project .sugar/ SQLite databases
2026-09-16 14:19:07 -07:00
teknium1
bd46352705 feat: catalog card image slot is 2:1; version pill reads as a pill
repo social banners are 2:1, so that shape fills the slot without a crop;
the recommended size is in the README/docs. Desktop AgentPluginRow gains
catalog_version to match the generated contract.
2026-09-16 14:18:39 -07:00
teknium1
034313e7cd feat: plugin catalog entries carry an optional version label and card image
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).

Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.

Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
2026-09-16 14:18:39 -07:00
brooklyn!
d5b3cad7b9 fix(mcp): isolate unsupported inferred app labels 2026-09-16 15:54:47 -05:00
brooklyn!
74ca4f28ba fix(onboarding): consume future catalog entries without custom metadata 2026-09-16 15:54:47 -05:00
brooklyn!
d9fedb066d revert: remove onboarding catalog additions 2026-09-16 15:54:47 -05:00
brooklyn!
4716ec0ba4 feat(desktop): suggest varied first tasks from real integrations 2026-09-16 15:16:09 -05:00
brooklyn!
12f495feac feat(mcp): curate first-task examples and add Blender bridge 2026-09-16 15:16:09 -05:00
brooklyn!
a38d544f1b feat(mcp): expose optional backend-local app discovery 2026-09-16 15:16:09 -05:00
brooklyn!
4e9d3c713a feat(desktop): browse catalogs as cards with a saved list option 2026-09-16 13:55:13 -05:00
brooklyn!
2e05fcf524 feat(ui): support icon-only segmented controls 2026-09-16 13:55:13 -05:00
brooklyn!
8b551a1ab3 docs(catalog): restore native browsing and install guidance 2026-09-16 13:55:13 -05:00
brooklyn!
cf59980e2c feat(catalog): restore website install links to Desktop 2026-09-16 13:55:13 -05:00
brooklyn!
cbd76e4ea3 feat(desktop): restore native skill and plugin catalogs
Restore the shared browser, protocol handling, and install confirmations from the original catalog work for further UI iteration.
2026-09-16 13:55:13 -05:00
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
Austin Pickett
a7691f41ec fix(desktop): explicit bot opens and session creates dial the pool at foreground priority (#113269)
* fix(desktop): foreground spawn priority through requestGatewayForAgent/retainGatewayForAgent and the session-create dials

The #102281/#104139 chain tagged every activation-door dial
(openGatewayForAgent/openGatewayForProfile/ensureGatewayForAgent/
sharedPrimaryRoute) as 'foreground' so a user-initiated open takes the
pool's reserved slot instead of queuing behind background roster
hydration. requestGatewayForAgent and retainGatewayForAgent never got a
spawnPriority, so they always dialed 'background', and they are the only
dial path createBackendSessionForSend (first send on a fresh chat) and
openNewSessionTile ("New session" / tab-strip "+") use. On a saturated
3-slot local pool the user's first message or new-tab click could wait
out the background slot timeout.

Add `{ spawnPriority }` options (main's requestGatewayForProfile shape,
5eb0ed4531) to both functions and forward it into every dial they make:
requestGatewayForProfile on the scope===key and primary-registry
branches, isAttachedSharedRemote, openSecondary, and the plain-profile
retain branch's gatewayForProfile(key, true, ...) which neither #105390
nor #110354 covered. The four session-create call sites pass
'foreground'.

Tests: gateway-spawn-priority.test.ts gains both polarities for
request/retain plus the plain-profile retain branch; the
use-session-actions and default-new-session matchers pin the tagged
dial. SpawnPriority is now exported for the SDK.

* fix(desktop): canonical Bot Chat lookup dials foreground through host.requestProfile options

Clicking a bot in the roster runs findExistingCanonicalChat ->
requestForBot(bot, 'session.list') -> host.requestProfile(route, ...)
-> requestGatewayForAgent with no intent, so the first RPC of the
gesture, the one that cold-spawns the bot's backend on a local pool,
dialed 'background'. With three roster backends already hydrating the
click produced no backend activity and the fail-closed lookup surfaced
as a "try again" toast (#105104). The later host.openSession is already
foreground, but it never runs until this lookup returns.

host.requestProfile gains a fifth `options?: { spawnPriority }`
argument; `timeoutMs` stays the fourth positional so no existing caller
changes shape, and the SDK keeps its exact call arity when no options
are given. requestForBot takes the same options and forwards them only
when set, so passive roster warming (profiles.list, ui_meta) still
dials background. The canonical `session.list` and the `session.create`
that follows on a first-ever open both pass 'foreground'.

Tests: profile-routing.test.ts (options form with and without a
timeout), routing.test.ts (tag forwarded, untagged call keeps three
args), canonical-chat-registry.test.ts (session.list and session.create
carry foreground; session.title does not).

---------

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
Co-authored-by: jxfjosh <282433633+jxfjosh@users.noreply.github.com>
2026-09-16 14:27:54 -04:00
teknium1
47685348ea test(update): receipt fixture sets HERMES_HOME, the seam _receipt_dir now reads
The fixture patched hermes_cli.config.get_hermes_home, which the receipt writer no longer
imports; locally run_tests.sh's temp HERMES_HOME hid the mismatch, CI counted files in the
wrong directory.
2026-09-16 11:06:52 -07:00
teknium1
5f071fb3ad fix(update): a failed dashboard cleanup no longer aborts fleet verification and the receipt
`_finish_dashboard_update_cleanup` runs pulled `dashboard_procs` inside the pre-pull
interpreter. A stale-symbol failure there (#112604: `AttributeError: module
'hermes_cli.main_dashboard' has no attribute '_loaded_launchd_backend_jobs'`, reproduced on
the maintainer's own box) propagated out of `_verify_fleet_after_update`, skipping the
fleet version matrix, plan-vs-execution reconciliation and the inner receipt finalize; the
command boundary then stamped the receipt `failed` with the traceback as stop_reason even
though the code update and gateway restart had succeeded.

The cleanup is now isolated like every sibling post-update step: the exception is
printed with a manual-restart hint, logged at WARNING, and recorded as a failed
`dashboard_cleanup` step on the receipt. A dashboard/serve left on pre-update code is still
escalated by the survivor probe → reconciliation (exit 1). Covers the git and ZIP paths.

Refs #112604 (residual class after #112753). Test is red on origin/main.
2026-09-16 11:06:52 -07:00
teknium1
1259b140c9 fix(update): the receipt survives a mixed sys.modules graph and a lost write is visible
`hermes update` writes its receipt from the PRE-pull interpreter after the post-pull
module purge. `_receipt_dir()` resolved the home through `hermes_cli.config`, so the
write re-executed the pulled `config.py` against whatever was still cached — on a pull
that added a symbol to a root-level module (`utils.file_signature`, `base_url_origin`)
that import raised and the whole receipt was dropped. The failure was logged at DEBUG,
which the updater's INFO log discards, so an activation run left no receipt while the
following no-op run wrote one, and `latest.json` kept pointing at an older run.

- `_receipt_dir()` uses `hermes_constants.get_hermes_home` (purge-protected, stdlib-only).
- A failed write prints `⚠ Update receipt not written: <exc>` and logs at WARNING.

Refs #112465, #112558 (Finding A). Both new tests are red on origin/main.
2026-09-16 11:06:52 -07:00
teknium1
26e9205653 refactor(delegation): carry the dispatcher's Context to stale finalization
Store the dispatching contextvars.Context on the record instead of a
home string, and run the monitor's forced _finalize under it. That keeps
the whole scope (home override + secret scope) rather than only the
home, and removes the set/reset override dance from
_push_completion_event, which now stays scope-agnostic.
2026-09-16 11:06:06 -07:00
teknium1
85ec40027d chore(api): drop unrelated docstring wording change from salvage 2026-09-16 11:06:06 -07:00
NanPan
c11b0a07e7 fix(agent): preserve periodic callback context
(cherry picked from commit a1a47efc07b3eaa0e71413ff9d5732a719e1d5a2)
2026-09-16 11:06:06 -07:00
NanPan
dceae1a766 test(agent): pin periodic scheduler context propagation
(cherry picked from commit 972a0444032ec7671ec30007ed883e0ef165e0f6)
2026-09-16 11:06:06 -07:00