Commit Graph

773 Commits

Author SHA1 Message Date
teknium1
393ffff04f fix(providers): resolve the chatgpt alias in the /model parser and hermes auth login too
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.

Part of #95794
2026-09-19 09:38:01 -07:00
Teknium
99a6ed741a Merge pull request #115757 from NousResearch/fix/boa-codex-app-server-lifecycle-migration-mcp
fix(codex-runtime): same-name MCP servers no longer make ~/.codex/config.toml unloadable; adds `hermes codex-runtime migrate` (#79023, salvage #79182)
2026-09-19 09:35:41 -07:00
kshitijk4poor
1536e1d242 docs(discord): finish the free-response auto-thread surfaces
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
2026-09-19 21:43:59 +05:30
Xunjin ZHENG
774e070731 feat(discord): add free_response_auto_thread opt-in
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.

Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.

Default behavior is unchanged; all 291 existing discord tests pass.
2026-09-19 21:43:59 +05:30
Siddharth Balyan
8d4abc3ea6 fix(mcp): retire the n8n bridge catalog entry (#116048)
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.

Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
2026-09-19 17:51:30 +05:30
teknium1
d09183d8d3 feat(cli): add hermes codex-runtime migrate [--dry-run] [--json]
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.

Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
2026-09-18 23:38:13 -07:00
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
5e95050608 fix(agent): a server context rejection the transcript cannot explain is no longer "conversation too long"
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).

_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.

Fixes #114644
2026-09-18 19:18:57 -07:00
kshitijk4poor
1250a3e4eb fix(backup): an incomplete archive never prunes the last complete ones
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
2026-09-19 03:43:58 +05:30
leomcamilo
ccb3d968ce fix(backup): exit non-zero when a full backup is incomplete
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.

`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.

Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
2026-09-19 03:43:58 +05:30
kshitijk4poor
8925c70a1c docs(whatsapp): group access section says what the gateway admits; env reference rows; trim bridge tests
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
2026-09-19 03:15:40 +05:30
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
b7b203cda0 fix(skills): built-in name collisions show a note in /skills, /help skills and the palette
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).

Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
  "slash command /<name> unavailable — name taken by built-in; use /skill <name>"

- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
  `hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
  skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
  carries the notice in its existing `warning` field (rendered by the Desktop
  `/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
  listing (Telegram/Discord/…) appends one ⚠ line per colliding skill

Fixes #113560
2026-09-18 12:51:22 -07:00
teknium1
250e12e760 fix(config): provider switch via config set drops the previous provider's route (#113719)
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.

Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):

- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
  a base_url is: the target's registry/plugin host, a `providers:` /
  `custom_providers:` entry resolving to the target, or bare custom/local
  aliases (configured BY base_url) -> owned; another known provider's host
  or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
  alone, with no base_url, is old-route wire state and goes too); keeps an
  owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
  prints what was cleared and why, or warns that an unrecognised base_url
  still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
  (openai-codex + chatgpt.com), a custom entry with that URL, and bare
  `custom` are untouched.

Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
2026-09-18 12:49:40 -07:00
teknium1
668e505ac7 fix(mcp): OAuth discovery/registration carry a User-Agent; cancelled login frees its callback port
Widen the salvaged fixes to the whole class and add the pieces they missed:

- The default User-Agent for SDK-built OAuth requests moves from the manager's
  bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
  provider gets it too, and the manager's pre-flight metadata discovery (its own
  client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
  registration and token requests through client.send(), which never merges the
  client's default headers; www.tradingview.com's WAF answers a header-less GET
  with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
  /register and /authorize on the MCP host, and the login died with
  "Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
  ... != None" because no issuer was ever discovered).

- The callback listener now runs serve_forever() and is shut down before
  server_close(). A thread parked in handle_request()'s select() keeps the
  closed listening socket alive (the kernel holds the file for the duration
  of the poll), so a flow cancelled mid-wait left the port bound and the
  retry on the same pinned/cached port raised "OAuth callback port N is
  already in use" with no external collider.

- When every authorization-server metadata fetch failed, a registration error
  is re-raised leading with those statuses ("Could not read
  authorization-server metadata (403 from ...); dynamic client registration
  then fell back to a guessed endpoint on the MCP host and failed: ...").
  humanize_oauth_registration_error leaves that message alone so the 403 in
  it is not mistaken for a DCR allowlist refusal.

Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
2026-09-18 12:47:49 -07:00
teknium1
76c640ff11 fix(web): list, activate and delete legacy custom_providers entries on Custom Endpoints
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.

The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
2026-09-18 12:44:32 -07:00
teknium1
3e182f46c3 fix(doctor): flag a non-list custom_providers and legacy list entries with no providers: twin; name the edited profile on Custom Endpoints / Local Models
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.

Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).

Part of #114471 (items 2, 5, 6).
2026-09-18 12:44:32 -07:00
teknium1
8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00
teknium1
09cf1b926f fix(state): reset forks leave the Python lineage walk too; /resume ranks lineages by activity
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.

Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.

Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.

Part of #114271
2026-09-18 10:39:13 -07:00
teknium1
4590ef8b58 fix(profiles): polled profile lists never walk skill trees; vanished skill dirs no longer abort enumeration
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.

- list_profiles(lazy_skill_count=True): skill_count is the last known value;
  a missing/aged entry schedules ONE background _count_skills per profile per
  60 s recheck window, so the request thread does zero skill-tree I/O and the
  refresh cadence is decoupled from the poll rate. The two polled callers and
  the REST fallback entry use it; the sync default (CLI, detail views) is
  unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
  os.walk with excluded/support dirs pruned) instead of rglob — half the
  fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
  uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
  projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
  skill add/remove immediately on the next refresh).

Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
  before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
          profiles.list 2189 + 2015, projects/tree 2177 + 2015;
          a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
  after:  GET /api/profiles 0 skill-tree stats/scandir on the request thread
          (counts land from the background refresh by the next poll),
          profiles.list 0, projects/tree 0; _count_skills (detail/control)
          still reports 200 with 203 scandir; the vanished-dir case returns
          199 and list_profiles() enumerates all 5 profiles.

Fixes #114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:18:40 -07:00
teknium1
72360ae1d2 fix(doctor): detect a dead IPv6 route and name network.force_ipv4
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.

Docs: doctor reference and the network config section describe the check.
2026-09-18 09:56:28 -07:00
teknium1
fc94ff56e1 fix(cli): expand HERMES_HOME once at process entry; sweep raw readers
Follow-up to the salvaged hermes_constants expansion (#109212): the CLI has
~30 raw `os.environ["HERMES_HOME"]` readers (fast --version path, early
display.interface probe, profile re-home, dotenv loader, subprocess-home
helper) that never go through get_hermes_home(). A literal `~` (fish, or
any quoted value) left them resolving `~/.hermes` against cwd while the
resolver now expands it, so the process would disagree with itself.

Normalize the env var once, at the earliest point of hermes_cli/main.py
(stdlib-only, before the fast paths), and expand it in the one raw
hermes_constants reader (`_profile_home_path`). A relative value that is
not tilde/variable-shaped is deliberately left alone: rejecting it would
break `HERMES_HOME=./tmp-home` in tests and CI for no user-facing gain.

Tests: real-CLI subprocess with HERMES_HOME='~/.x' under a fake HOME asserts
`config path` lands in the fake home and that no literal `~` directory
appears under cwd (the reporter's acceptance criterion, #114353), plus a
unit test for the normalizer. Docs: environment-variables reference.
2026-09-18 09:49:46 -07:00
teknium1
b65ec6a3f0 docs: list the DISCORD_MISSED_MESSAGE_BACKFILL_* env fallbacks
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
2026-09-18 09:44:59 -07:00
teknium1
9c088a7a78 fix(config): fix container type for unseeded roots, top-level lists and bare-name list slots
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.

agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
2026-09-18 09:41:02 -07:00
teknium1
0972b4819a fix(config): config set refuses a wrong-shaped value for a list/mapping key
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.

Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.

Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.

Config-set atom of #114471.
2026-09-18 09:41:02 -07:00
teknium1
46500f9c79 fix(dashboard): --stop and the update sweep resolve backend ownership tri-state, never guess
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).

What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
  (spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
  else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
  with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
  `_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
  `/proc/<pid>/environ` never reflects it. Without this, the common default-home install
  (nothing exported) got a `--stop` that matched nothing — the regression flagged on
  #113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
  #113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
  only a foreign backend prints "No hermes dashboard processes running for this profile."
  instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
  hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
  their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
  the (now filtering) finder. Docs: `--stop` row in the CLI reference.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 09:38:01 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
Ritesh Patel
9927199998 docs: drop opencode-free references (provider removed)
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
2026-09-18 15:40:37 +05:30
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
98b4efbf48 fix: clarify cards that cannot render are re-asked as plain text, never reported as user inactivity
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could
fail without the user ever seeing a question, and the agent then waited out the full
clarify_timeout and reported "[user did not respond within Nm]" (#112684):

- the platform rejects the card at once -> the wait aborted with a delivery sentinel but the
  question was never re-asked;
- the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts,
  a stale-thread retry) -> the runner never looked at the send future again and blocked for
  the whole timeout for a card that never posted;
- no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported
  timed_out=True with an empty notice.

gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers
move out of the gateway/run.py facade) now retries every definitive card failure once through the
adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress
DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a
possibly-delivered send so a late failure releases the waiter with
"[clarify prompt could not be delivered]". The no-surface case reports
"[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already
treats as a non-answer).

Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s):
card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and
the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt,
after the numbered prompt is sent and "2" resolves to the choice.
2026-09-17 09:06:17 -07:00
teknium1
f3bdcd0877 fix: dashboard stop sweeps the wedged descendants that outlive the SIGKILLed backend
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.

_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.

Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.

The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).

Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
2026-09-17 09:04:11 -07:00
teknium1
c61a11474b fix(desktop): update check falls back to the gh CLI login before going anonymous
Closes the last atom of #112615. A Desktop app started from the Dock, Finder or a
desktop launcher inherits a minimal environment, so the GITHUB_TOKEN / GH_TOKEN rung
landed in #113190 never fires for exactly the users the issue is about (shared exit
IP, 60/hour anonymous budget exhausted by neighbours). The update check now walks the
same ladder as the Python GitHub client (tools/skills_hub_github.py::GitHubAuth):

1. GITHUB_TOKEN, then GH_TOKEN, from the launch env (unchanged).
2. `gh auth token` — execFile with an argv array (no shell), stdin closed, 3 s
   timeout, windowsHide. `gh` is resolved on PATH plus the GUI-safe install
   locations backend-env already appends for the backend (Homebrew, /usr/local,
   ~/.local/bin, nix profile, the Windows GitHub CLI installer dirs). The outcome
   (token or none) is cached for the process lifetime so gh runs at most once.
3. Anonymous.

A rejected credential (401) still retries anonymously; the log line names the
source (env vars vs gh login) and never the token, once per source per process.
A rejected gh token also drops the cache so a re-login is picked up by the next
check. envTokenRejected is renamed githubTokenRejected since both rungs use it.

Live: with PATH=/usr/bin:/bin:/usr/sbin:/sbin and no env token, the module found
gh, resolved a credential in 135 ms and api.github.com reported a 5000/hour core
budget (anonymous: 60). A logged-out gh and a hanging gh both resolved to null
(anonymous) in 5 ms / 3001 ms.

Docs: the credential ladder is documented on the Desktop "Updating" page and in
the environment-variables reference.
2026-09-17 08:54:13 -07:00
kshitijk4poor
f537076738 docs(config): config set docs describe the wrong-prefix-only refusal
The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.

Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.
2026-09-17 21:19:22 +05:30
fangliquanflq
b632f20484 docs(mcp): clarify transport keepalive defaults 2026-09-17 21:18:03 +05:30
teknium1
8a8bd93ef2 docs: -z exit codes differ from chat -q/-Q on purpose; real turn_exit_reason example
The two one-shot entry points map the same outcome to different codes
(failed/partial/budget -> 2 on -z, 1 on -q/-Q; completed-with-no-text -> 1 vs 0)
and the reference documented both tables 60 lines apart without saying so.
Also `iteration_limit` is a metrics end_reason, not a turn_exit_reason; the
finalizer emits `max_iterations_reached(N/M)`.
2026-09-16 18:12:34 -07:00
teknium1
391a7007fb fix(cli): hermes -z exits non-zero for partial, failed and interrupted runs even when text was printed
`hermes -z` judged the run by whether stdout was non-empty: a turn that hit the
iteration budget, ended `partial`, failed on the provider, or was interrupted
still exited 0 as long as it printed an explanation, so scripts treated a
half-done job (or a 401 summary) as success. The `-q`/`-Q` paths already got
one result-derived contract in 23863ccbafe0d; this is the same class on the
scripted one-shot entry point (#111770).

- `hermes_cli/oneshot.py::_oneshot_exit_code`: 0 only when the turn completed;
  130 interrupted; 2 failed / partial / completed:false; 1 a completed turn
  with no text (unchanged message).
- `--usage-file` now also reports `partial`, `interrupted` and
  `turn_exit_reason`, so a pipeline can tell a budget stop from a Ctrl-C.
- Docs: `hermes -z` exit-code table + usage-file fields in cli-commands.md.

Live probe (run_oneshot with a stubbed _run_agent): partial-with-text 0 -> 2,
budget-with-text 0 -> 2, interrupted-with-text 0 -> 130, failed-with-text 0 -> 2;
completed 0 -> 0 (control), partial-no-text 2 -> 2 (unchanged).
2026-09-16 18:12:34 -07:00
teknium1
6671365fb2 test: one invariant suite for mid-turn session-switch refusals + docs
Replace the two salvaged per-command test files with a single suite that drives
the real /branch, /resume and /sessions <id> handlers against a real SessionDB
and asserts the invariant that matters: while a turn is running, the session id
(CLI and agent) is unchanged and the current row is not ended; when idle,
/branch still proceeds. Red on origin/main (3 busy cases), green with the guards.

The /branch MagicMock fixture gains an explicit `_agent_running = False`: a bare
MagicMock attribute is truthy and would trip the new guard.

Docs: /branch and /resume rows in the slash-command reference say they are
refused mid-turn in the classic CLI and why.

Co-authored-by: sonderhq <144971773+sonderhq@users.noreply.github.com>
Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 17:49:48 -07:00
teknium1
b3aef45987 fix(clarify): batch result carries the surface's no-answer notice; trim tests
Follow-up to the #112725 salvage. When a gateway clarify batch stops at an
unanswered question, the batch payload only said `timed_out: true` with blank
answers — so a card the platform REJECTED ("[clarify prompt could not be
delivered]") read exactly like a user who walked away, which is the misreport
#112684 describes ("a timeout must not be presented as user inactivity when the
prompt was never delivered").

- gateway/run_turn_runner.py::_clarify_batch_sync: the unanswered question's
  own response text rides along as `notice`.
- tools/clarify_tool.py::_run_batch/_batch_result: pass `notice` through into
  the result JSON beside `timed_out` (only when present); schema description
  mentions it. CLI/TUI batch callbacks send no notice -> byte-identical output.
- tests: fold the single-question re-arm pin into the existing bracket-answer
  test (<=2 new tests), parametrize the end-to-end test over an undeliverable
  Telegram-shaped adapter so the delivery sentinel is pinned as `notice`.
- docs: tools-reference.md describes `notice`.

Part of #112684
2026-09-16 17:47:50 -07:00
teknium1
da942e4483 fix(config): hedge the config get phantom-key notice — the schema walk has false positives
`_validate_config_key` is a DEFAULT_CONFIG walk, and several live keys are
deliberately unseeded so that a stored value counts as an explicit user pick
(browser.cloud_provider, stt.provider, cron.misfire_grace_minutes,
gateway.proxy_url / relay_url, dashboard.ssh_isolated_idle_grace_s,
computer_use.backend). The runtime reads all of them and Hermes' own wizards
write some of them, so `hermes config get browser.cloud_provider` asserting
"Hermes does not read it" was a false claim.

Match the set-path notice ("Hermes may not read it"): the key still gets
flagged so a leftover misspelling cannot pass unnoticed, but the CLI no longer
states something it cannot know. Docs updated to the same wording; one test
pins the hedge on an unseeded live key (red before this commit).
2026-09-16 17:42:32 -07:00
teknium1
efc0359051 fix(config): inline the phantom-key notice into get_config_value, trim to 2 tests, docs
Slim follow-up to the salvaged commit from #112350 (@DavidMetcalfe): the two
helpers (`_is_unread_nested_key`, `_print_unread_key_notice`) collapse into
~10 lines at the single call site in `get_config_value`, reusing
`_validate_config_key` for both the verdict and the did-you-mean suggestion
instead of calling it twice. `is_env_key` is dropped: `.env` keys and
UPPER_SNAKE settings are never a known top-level section, so the guard
already excludes them.

Tests: keep the two invariants (phantom nested key -> notice on stderr while
stdout/--json stays parseable; known / custom top-level / open-subkey keys
are not flagged). Dropped `test_missing_key_exits_without_a_notice` — the
missing-key exit precedes the print and is pre-existing behaviour already
covered elsewhere.

Docs: one sentence each in cli-commands.md (`config get` row) and
configuration.md (the set-path did-you-mean paragraph).
2026-09-16 17:42:32 -07:00
teknium1
c7d4d32800 fix(desktop): env-only GitHub token for the update check, honest rate-limit copy
Trim the salvaged ladder to its env rung and make the 403 copy actionable.

- github-api-auth.ts keeps githubTokenFromEnv / githubApiHeaders; the
  `gh auth token` rung (resolveGithubToken, readGhCliToken, the per-process
  cache and the execFile import in main.ts) is dropped. Whether a GUI app may
  silently spend the user's gh CLI login on a passive background check is a
  product call, not a bug fix; the pure seam makes re-adding it a ~15-line
  follow-up.
- fetchGitHubApi reads GITHUB_TOKEN / GH_TOKEN from process.env per request
  (nothing stored) and attaches x-ratelimit-remaining / x-ratelimit-reset plus
  whether the call was authenticated to the rejected error.
- describeUpdateCheckFailure moves out of the unimportable main.ts into
  update-api-check.ts. A 403/429 with x-ratelimit-remaining: 0 now says the
  anonymous budget is 60/hour per network address shared with everyone behind
  the same connection, gives the real reset time and names GITHUB_TOKEN as the
  remedy; a 403 without rate-limit headers is reported as a plain HTTP 403
  instead of a rate limit the user did not cause.
- Docs: desktop.md "Updating" and the GITHUB_TOKEN row in
  environment-variables.md.
- Tests trimmed to two invariants per module.

Why: behind a shared exit IP "try again in an hour" is advice the user cannot
act on, and the old copy hid the one thing that does help.
2026-09-16 17:22:39 -07:00
teknium1
08aa67e473 docs: state that HERMES_DEBUG_INTERRUPT=0/false/off keeps interrupt tracing off
Both readers now share utils.env_var_enabled(), so the reference table
spells out the falsy values instead of implying "any value enables".
Also registers the contributor email mapping for the salvaged commit.
2026-09-16 17:21:14 -07:00
teknium1
f41ad3bc0a fix(gateway,cron): key Matrix handoff/cron thread seeds the way Matrix replies are keyed
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).

The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).

- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
  `get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
  use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
  `group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
  asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
  contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.

Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before  create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
        cron seed matrix🧵… ≠ inbound
after   create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)

Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
2026-09-16 17:18:44 -07:00
teknium1
adb4323e75 test(tools): platform-aware TZ assertion, Windows live-child contract, docs
tests/tools/test_code_execution.py::test_timezone_injected_when_set asserted
child_env["TZ"] == "America/New_York" unconditionally, which encoded the
Windows defect as intended behaviour; on win32 it now asserts TZ is absent and
keeps the POSIX assertion.

Add a windows_only live-child test to test_code_execution_windows_env.py: with
timezone: configured, a real Windows child's time.timezone and astimezone()
offset must equal the test process's own OS-zone view — the reporter's exact
witness (time.timezone == 0, +01:00 instead of -07:00, tzname still correct).

Docs (environment-variables.md, configuration.md): HERMES_TIMEZONE / timezone:
reaches execute_code children as TZ on Linux/macOS only; Windows children keep
the OS zone because the Windows C runtime parses only POSIX-form TZ strings.

Sibling env builders swept: the remote-backend prefix in
code_execution_tool.py (TZ=... before `python3 script.py`) targets POSIX
backends (Docker/SSH/Modal) and is unchanged; no other spawner forwards
HERMES_TIMEZONE into TZ. HERMES_TIMEZONE itself stays stripped from the child
(adf23550f5: under the multiplexed gateway it holds only the default
profile's value).

Refs #112233
2026-09-16 17:14:28 -07:00
teknium1
8899aeff53 fix(oneshot): --usage-file ledger reports auxiliary LLM spend
`hermes -z --usage-file` copied only the main-loop result, so title generation,
vision, compression, web_extract and background-review calls — recorded per task
in session_model_usage — never reached the pipeline ledger the flag advertises
as "so pipelines can always account for spend". The Insights page already folds
those rows in (#23270); the ledger is now consistent with it.

- SessionDB.auxiliary_usage_by_task(session_id): per-task sums over the
  session's compression lineage (aux calls bill to the id the turn started with
  while compression mints child ids mid-turn).
- oneshot snapshots aux usage before the turn and attaches the delta after it,
  so a resumed session's earlier runs are not re-billed.
- The report gains `auxiliary` (totals + `by_task`) and
  `total_including_auxiliary`; every existing key keeps its main-loop meaning.
- The auto-title upgrade runs on a daemon thread and can still be writing when
  the turn returns: title_generator tracks in-flight upgrade threads and
  oneshot joins them (bounded) before reading — no sleep, no eager read.

Fixes #112848. Direction shared with #112852 (@KoNit-K), which folded aux into
the headline counters; this keeps them backward compatible instead.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:03:25 -07:00
teknium1
30bde46510 fix(profiles): profile create never deletes a live marker-less dir; dangling symlink markers count as identity
Review follow-up on the identity predicate:

- create_profile: the rmtree guard is back to "tombstoned AND no identity
  marker". A live profiles/<name>/ without a marker is invisible to
  `profile list` but may still hold user files (skills/, memories/,
  cron/jobs.json) — main refused it with FileExistsError; the widened guard
  silently rmtree'd it. Now fails closed with an error naming the stray dir.
- named_profile_has_identity: `is_file()` follows symlinks, so a legacy
  profile whose only marker is a dangling symlinked config.yaml/.env
  (clone/migration leftover) vanished from list/serve/-p. A symlink is an
  identity claim: accept `is_symlink()` too.
- Docs: faq.md still said "each profile is just a directory under
  profiles/"; faq.md + user-guide/profiles.md + hermes_cli/AGENTS.md now
  state the identity-file contract.
2026-09-16 16:48:42 -07:00
teknium1
c055eee1fe refactor(checkpoints): slim the profile-rename rekey and name it as a checkpoint_manager sibling
Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):

- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
  repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
  idempotent rekey: every step overwrites and the old project metadata is removed last, so a
  mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
  A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
  192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
  (`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
  and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
  assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.

Fixes #112973
2026-09-16 14:34:22 -07:00
xielevi
a41552fad4 fix(profiles): purge a deleted profile's session/routing identity on delete
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:

- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
  `_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
  also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
  namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
  (`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
  in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
  writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
  `_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
  has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
  partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
  profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
  `create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
  delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
  leaves of a profile's conversation record is `delete_profile`'s business — it removes the
  profile's own home, `state.db` included.

Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
2026-09-16 14:21:14 -07:00
xielevi
81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00