DELETE /api/sessions/<id> removed only the state.db rows: the durable
channel->session routing index (gateway_routing table + sessions.json
mirror) survived, so the next Discord/Telegram message routed to the SAME
id and resurrected the deleted row, and the on-disk .json/.jsonl
transcripts plus request_dump files were never scrubbed because
sessions_dir was not passed to delete_session (#42422).
- SessionStore.remove_by_session_id: drop every entry pointing at the id
(one channel can hold several) and persist the drop to both durable
copies; the index is written back by the gateway process, so removing
only DB rows elsewhere is undone by the next whole-index save.
- The API delete handler now passes the request-scoped sessions_dir and
clears the routing entries through the runner's SessionStore.
- Deletes made out of the gateway process self-heal at routing time via the
stale-route guard once a missing row counts as ended.
Fixes https://github.com/NousResearch/hermes-agent/issues/42422
_slack_tools_loaded/_discord_tools_loaded run per turn via _ephemeral_change_key and
_get_platform_tools only reads the config, so the deepcopy in load_config() is pure
overhead on the event loop. Use load_config_readonly().
Partial salvage of #117993: kept the two gateway/session.py hunks, dropped
TestCapabilityProbeConfigLoader because it monkeypatched cfg.load_config to raise —
a change-detector on the symbol read, not an invariant.
(cherry picked from commit a487903180090c5a2fccd4749b05d302cf2740b1)
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.
Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
SessionStore.switch_session re-points a session key at another session id
for non-boundary reasons too (async-delegation re-pin, compression-tip
binding heal, CLI handoff, /branch), but _replace_route_locked built the
new SessionEntry without model_override, so the user's /model pin silently
reset to None on every such re-point (#119864).
Scoped to switch_session's kwargs, not the shared helper: reset_session
(/new) uses the same helper and is a deliberate boundary (30e947e0a0) that
must keep dropping the persisted pin, or the next turn's
_rehydrate_session_model_override resurrects it.
Salvaged from #119868 (moved from _replace_route_locked into switch_session).
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
After a restart the routing index rebuilt every lane from `SessionEntry.origin`,
which carries the runtime profile (key namespace) but not the bot that received
the conversation. Delivery then fell to `_is_shared_bot_satellite`: a lane owned
by a secondary bot whose runtime profile is ALSO a satellite of the default bot
was handed to the default bot, and authorization read the wrong allowlist.
- `SessionEntry.transport_profile` (routing JSON) + nullable
`sessions.transport_profile` (SCHEMA_SQL, reconciled by the existing column
path; `agent:main` keys untouched, standalone gateways write nothing). Stamped
from the pinned `RoutingIdentity` at create, reset/switch, DB recovery and
every peer refresh; compression forks inherit it like the other routing columns.
- `session_identity.restore_identity()` re-pins a `RoutingIdentity(transport=None)`
from the persisted transport profile; `authz_mixin._restored_source(entry)` is
the one seam every revive path uses (auto-resume, heartbeat restore, plugin
injection, background-process events).
- `_adapter_for_source` / `_adapter_profile_for_source` honour a restored identity:
the persisted bot's adapter or None — never the default bot by heuristic.
Entries written before the column exist keep the old chain.
Phase 5 of #88715.
- 15 `MERGE-CHECK:` conflict-resolution comments removed from prod code (two were
TODOs already done: the utf-8-sig sessions.json read lives in session_persistence,
the pm-aware cron script helpers in scheduler_script).
- 49 imports the branch left unused (ruff F401, none present at the merge base,
none inside PLUGIN-COMPAT blocks). update_cmd's frozen-surface re-exports are
trimmed to the names tests/compat/old_updater_surface.json actually lists under
hermes_cli.update_cmd; the rest resolve through hermes_cli.main.__getattr__.
- tools/environments/local_gitbash_probe.py: nothing imported it once _find_bash
delegated to pm.shell().
- Three try/except wrappers around calls that cannot raise (install_truststore,
get_hermes_home, and a duplicated except clause in supermemory).
When the session FTS (full-text search) index rebuild fails, the
gateway previously entered a cooldown and then permanently stopped
retrying once the cooldown expired, leaving search silently broken
for the rest of the process lifetime. This adds a proper retry after
the cooldown window elapses, and escalates to ERROR-level logging
(previously missing 'import logging' meant this path could not even
log the failure) so repeated rebuild failures are visible to
operators instead of being swallowed.
Related but out of scope: while testing the fallback paths we
observed that the JSONL transcript fallback writer can write zero
bytes when the primary DB write path fails partway through - flagging
this for maintainers as a separate follow-up, not fixed here.
Fixes#114266
The non-compression branch of _resolve_async_delegation_session awaited the
spawning-session row lookup and then unconditionally called switch_session(),
so a run invalidated (/stop) or a route replaced (/new, /resume) while the
lookup was pending was overwritten by the stale completion's pin.
- Snapshot the routing key's run generation before the await and re-check it
before mutating the route; invalidated -> drop the injection, route untouched.
- switch_session(expected_session_id=...) turns the /resume primitive into a
CAS for this caller, mirroring advance_compression_session: a route that
moved past the snapshot wins over the late completion.
- _current_session_run_generation extracted from _is_session_run_current.
Slimmer redo of #113692 (@KoNit-K) and #113716 (@wangtaotaotao95): same
direction, without the second routing-authority lock and authorize callbacks.
Fixes#113690
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: wangtaotaotao95 <wangtaotaotao95@users.noreply.github.com>
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:
- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
`_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
(`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
`_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
`create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
leaves of a profile's conversation record is `delete_profile`'s business — it removes the
profile's own home, `state.db` included.
Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.
Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop
* feat(gateway): /signin signs the free tier into a Nous account from a DM
* feat(cli): chat surfaces name /signin as the sign-in verb
* fix(auth): review follow-ups for the shared sign-in flow and /signin
* fix(i18n): carry the /status free-tier line in every locale catalog
* refactor(cli): the chat sign-in command is /login
* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
Merge upstream b1f003e186 while preserving PM runtime ownership and
Python 3.14 worker startup, Windows signing, and macOS wait recovery.
Keep retired runtime modules deleted. Port upstream updater preflight
checks into the checkout strategy and preserve live build logging.
Carry checkpoint filename handling and process recovery into the current
module layout. Regenerate locks and adapt incoming platform test markers.
Focused Python and JavaScript tests, desktop and root-test typechecks,
conflict-path lint checks, lock validation, and retired-import checks pass.
The full test suite and packaged release builds were not run.
Reuse the effective gateway config and shared session platform policy before hashing model-facing metadata. Preserve original routing state and cover enabled/disabled redaction across all busy injection routes.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
Re-applies the gateway compat removal byte-for-byte; see 92d0bd0d73 for the
full inventory (30 re-exports/aliases + 2 shim modules dropped, 3 shim-only
names re-removed, 24 callers + 34 test files repointed). No new changes.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:
- hermes_state.find_latest_gateway_session_for_peer: the fallback query
now requires COALESCE(s.profile_name, <store owner>) = <store owner>
(owner via SessionDB._own_profile_name). Handles NULL legacy rows and
the single→multiplex migration case a key-namespace fence would break;
stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
multiplexing no longer `return True` — the recovered row's agent:<ns>:
must match the REQUESTED key's namespace (the active profile is
meaningless when several profiles serve concurrently). Single-profile
behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
when parent and child agree on agent:<ns>: (or either is keyless), so a
default child forked from a sibling row is not durably mislabelled.
Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>