Commit Graph

256 Commits

Author SHA1 Message Date
teknium1
8eda97cafb fix(webhook): make delivery mirroring opt-in and profile-scoped
Follow-up to the salvaged mirror commit.

- Opt-in (`mirror_to_session: true`, or `hermes webhook subscribe --mirror-to-session`),
  default off like cron's `mirror_delivery`. The mirrored text lands with user authority in
  the target chat, and on `deliver_only` routes it is the raw rendered payload, so a route
  author has to ask for it. Only a real boolean `true` opts in (a YAML string "false" no
  longer does).
- The mirror runs inside the routed profile's scope, so a `/p/<profile>/` delivery writes
  into THAT profile's state.db. A Telegram DM chat_id is the user's id on every bot; an
  unscoped mirror landed in the default profile's DM with the same person (live-probed).
- Tests trimmed to two invariants against a real state.db: an opted-in delivery lands in the
  routed profile's chat session and not the default profile's; a route without the opt-in
  (including `mirror_to_session: "false"`) never touches the target transcript.
- Docs: webhooks guide section "Replying to a delivery" (default, profile scope, trust note),
  CLI reference row, bundled hermes-agent skill reference.
2026-09-23 17:58:45 -07:00
teknium1
2fb564d2df fix(cli): move the hermes-agent argv layer into agent/legacy_cli
Follow-up to the salvaged wrapper: the argument layer now lives in a topical
sibling (agent/legacy_cli.py) instead of growing the run_agent facade, and
`python run_agent.py` (the installer's PATH launcher for hermes-agent) routes
through the same parser instead of fire, so `--version` and a bare invocation
no longer run the demo turn there either.

- --help/-h/--version use argparse's own exit path; options carry help text
- the runner is injected (`run=`) so `python run_agent.py` does not import
  run_agent a second time
- tests trimmed to two invariants that resolve the console-script target
  from pyproject exactly like pip's wrapper (red on main, green here)
- docs: hermes-agent section in the CLI reference
2026-09-23 10:40:39 -07:00
teknium1
c500dc6d99 fix: hermes update restarts only the gateways of the home it updates (#93349)
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.

Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
2026-09-23 05:36:30 -07:00
teknium1
9d799e0531 docs: hindsight installs from the plugin catalog
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
2026-09-23 01:16:41 -07:00
teknium1
9a2d96ca11 fix(config): flag quoted list/mapping values in doctor + startup, guard the list slots config set missed
A list/mapping slot holding ONE quoted string (`plugins:\n  enabled: '["a","b"]'`,
`model_catalog:\n  excluded_providers: '["openai-api"]'`) is skipped by every
isinstance-gated reader (`plugins_cmd._config_name_set` -> set(), `plugins._names`,
`inventory.excluded_providers` -> []) while `config get` echoes it back, so user
plugins silently unmount and provider exclusions silently lapse. Neither `hermes
doctor` nor the startup `print_config_warnings` banner said a word.

Reader/doctor half: one schema-aware pass in `validate_config_structure`
(`_validate_quoted_containers`) walks `DEFAULT_CONFIG` (sections included) plus
`_KNOWN_CONTAINER_TYPES` and warns when the user value is a string that parses to
a list/mapping. The message names the key, the quoted value and the remedy
(`hermes config set <key> '<literal>'`). Finding only: the file is never
rewritten. String-typed keys (`approvals.mode: "[off]"`), the `model: <name>`
shorthand and the `parse_config_string_list`-read slots
(`agent.disabled_toolsets`, `skills.disabled`) are not flagged. Feeds both the
doctor "Config Structure" section and the startup banner.

Writer half (audit of every path that can put a string in a container slot):
- `hermes config set` (`hermes_cli/config.py::set_config_value`): already parses
  bracket/brace literals (#88163) and refuses wrong-shaped values
  (`_refuse_container_type_mismatch`), BUT the guard only knew slots present in
  DEFAULT_CONFIG or `_KNOWN_CONTAINER_TYPES`. `plugins.enabled`,
  `plugins.disabled` and `model_catalog.excluded_providers` are deliberately
  absent from DEFAULT_CONFIG, so `config set plugins.enabled foo` /
  `plugins.enabled a,b` / `model_catalog.excluded_providers openai-api` stored a
  plain string (live repro on base). Added the three keys to
  `_KNOWN_CONTAINER_TYPES`: those writes are now refused with the literal hint.
- `cli.py::save_config_value` + callers (cli_*_mixin, gateway/slash_commands,
  gateway/run_busy): every caller passes a bool/enum string for scalar keys; no
  list-slot caller. Not reachable.
- `hermes_cli/plugins_cmd.py::_save_plugin_sets` / `_write_config_value` and
  `plugins_cmd_catalog` (via `_save_enabled_set`): write `sorted(set)` — real
  lists. Not reachable.
- Dashboard `PUT /api/config` (`web_routers/config_env.py::update_config` ->
  `_denormalize_config_from_web` -> `save_config`): schema-driven form;
  `web/src/components/AutoField.tsx` splits list-typed fields into a real array
  before the PUT. Not reachable.
- tui_gateway `config.set`: fixed `_CONFIG_SETTERS` table of scalar keys only
  (out of this lane's files anyway). Not reachable.
Conclusion: the quoted shapes on real machines are leftovers of pre-#88163
`config set` runs plus the DEFAULT_CONFIG-absent slots fixed here.

Live repro (fake HOME/HERMES_HOME, both quoted shapes in config.yaml):
  before: validate_config_structure() -> []; startup banner silent; doctor has
          no Config Structure section; _config_name_set('plugins','enabled') -> set()
  after:  two warnings naming plugins.enabled / model_catalog.excluded_providers
          with `hermes config set ... '["a","b"]'`; doctor prints them under
          Config Structure; running the remedy yields real lists,
          validate -> [], _config_name_set -> {'a','b'}
  writer: `config set plugins.enabled a,b` wrote `enabled: a,b` on base; now
          refused ("must be a list ... nothing was written").

Fixes #83308
Fixes #105706
credit: @fangliquanflq #105725 (schema-aware detection in validate_config_structure; slim redo)
credit: @Luna161 #83313 (first report + per-key warning in plugins_cmd)
2026-09-21 23:31:02 -07:00
teknium1
77f80d31ee fix(plugins): catalog provenance is installer-owned; kill list covers update/enable/load; re-pin keeps user files
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):

- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
  as catalog:official everywhere and marked the real entry "installed". Provenance
  now lives on the installer-owned .install-metadata.json record (catalog block
  written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
  never reads the tree. Pre-fix installs are adopted once when the installer
  record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
  canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
  gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
  list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
  the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
  data and user patches with no warning. Untracked/ignored files are carried
  into the new tree; edits to tracked files are copied to
  <HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
  The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.

repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
2026-09-21 23:17:05 -07:00
teknium1
91d1339a8d docs: multi-profile gateway pages match the multiplex-only topology
Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.

Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).

Completes #118273 (docs item). Part of #109417
2026-09-21 11:40:48 -07:00
teknium1
661d73f06b test+docs: trim doctor exit-status tests to two invariants; document the exit code
Keep the in-process return-code matrix (issues / manual / --fix full and
partial repair) and the real-process exit-status check; drop the exception
passthrough, --ack and --live cases, which pin behaviour this change does not
touch. Document 0/1 exit status under `hermes doctor` in the CLI reference.
2026-09-21 01:57:05 -07:00
teknium1
2520eb7c0d docs(backup): list the browser profile dirs hermes backup excludes 2026-09-21 01:44:42 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
85b2a3df6c feat(cli): hermes usage [--json] prints the /usage account limits without a session
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.

Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).

Co-authored-by: himanusia <himanusia@users.noreply.github.com>
2026-09-19 10:31:14 -07:00
teknium1
393ffff04f fix(providers): resolve the chatgpt alias in the /model parser and hermes auth login too
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.

Part of #95794
2026-09-19 09:38:01 -07:00
Teknium
99a6ed741a Merge pull request #115757 from NousResearch/fix/boa-codex-app-server-lifecycle-migration-mcp
fix(codex-runtime): same-name MCP servers no longer make ~/.codex/config.toml unloadable; adds `hermes codex-runtime migrate` (#79023, salvage #79182)
2026-09-19 09:35:41 -07:00
Siddharth Balyan
8d4abc3ea6 fix(mcp): retire the n8n bridge catalog entry (#116048)
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.

Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
2026-09-19 17:51:30 +05:30
teknium1
d09183d8d3 feat(cli): add hermes codex-runtime migrate [--dry-run] [--json]
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.

Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
2026-09-18 23:38:13 -07:00
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
kshitijk4poor
1250a3e4eb fix(backup): an incomplete archive never prunes the last complete ones
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
2026-09-19 03:43:58 +05:30
leomcamilo
ccb3d968ce fix(backup): exit non-zero when a full backup is incomplete
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.

`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.

Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
2026-09-19 03:43:58 +05:30
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
250e12e760 fix(config): provider switch via config set drops the previous provider's route (#113719)
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.

Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):

- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
  a base_url is: the target's registry/plugin host, a `providers:` /
  `custom_providers:` entry resolving to the target, or bare custom/local
  aliases (configured BY base_url) -> owned; another known provider's host
  or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
  alone, with no base_url, is old-route wire state and goes too); keeps an
  owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
  prints what was cleared and why, or warns that an unrecognised base_url
  still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
  (openai-codex + chatgpt.com), a custom entry with that URL, and bare
  `custom` are untouched.

Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
2026-09-18 12:49:40 -07:00
teknium1
76c640ff11 fix(web): list, activate and delete legacy custom_providers entries on Custom Endpoints
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.

The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
2026-09-18 12:44:32 -07:00
teknium1
3e182f46c3 fix(doctor): flag a non-list custom_providers and legacy list entries with no providers: twin; name the edited profile on Custom Endpoints / Local Models
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.

Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).

Part of #114471 (items 2, 5, 6).
2026-09-18 12:44:32 -07:00
teknium1
72360ae1d2 fix(doctor): detect a dead IPv6 route and name network.force_ipv4
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.

Docs: doctor reference and the network config section describe the check.
2026-09-18 09:56:28 -07:00
teknium1
9c088a7a78 fix(config): fix container type for unseeded roots, top-level lists and bare-name list slots
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.

agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
2026-09-18 09:41:02 -07:00
teknium1
0972b4819a fix(config): config set refuses a wrong-shaped value for a list/mapping key
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.

Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.

Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.

Config-set atom of #114471.
2026-09-18 09:41:02 -07:00
teknium1
46500f9c79 fix(dashboard): --stop and the update sweep resolve backend ownership tri-state, never guess
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).

What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
  (spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
  else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
  with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
  `_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
  `/proc/<pid>/environ` never reflects it. Without this, the common default-home install
  (nothing exported) got a `--stop` that matched nothing — the regression flagged on
  #113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
  #113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
  only a foreign backend prints "No hermes dashboard processes running for this profile."
  instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
  hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
  their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
  the (now filtering) finder. Docs: `--stop` row in the CLI reference.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 09:38:01 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
Ritesh Patel
9927199998 docs: drop opencode-free references (provider removed)
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
2026-09-18 15:40:37 +05:30
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
f3bdcd0877 fix: dashboard stop sweeps the wedged descendants that outlive the SIGKILLed backend
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.

_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.

Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.

The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).

Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
2026-09-17 09:04:11 -07:00
kshitijk4poor
f537076738 docs(config): config set docs describe the wrong-prefix-only refusal
The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.

Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.
2026-09-17 21:19:22 +05:30
teknium1
8a8bd93ef2 docs: -z exit codes differ from chat -q/-Q on purpose; real turn_exit_reason example
The two one-shot entry points map the same outcome to different codes
(failed/partial/budget -> 2 on -z, 1 on -q/-Q; completed-with-no-text -> 1 vs 0)
and the reference documented both tables 60 lines apart without saying so.
Also `iteration_limit` is a metrics end_reason, not a turn_exit_reason; the
finalizer emits `max_iterations_reached(N/M)`.
2026-09-16 18:12:34 -07:00
teknium1
391a7007fb fix(cli): hermes -z exits non-zero for partial, failed and interrupted runs even when text was printed
`hermes -z` judged the run by whether stdout was non-empty: a turn that hit the
iteration budget, ended `partial`, failed on the provider, or was interrupted
still exited 0 as long as it printed an explanation, so scripts treated a
half-done job (or a 401 summary) as success. The `-q`/`-Q` paths already got
one result-derived contract in 23863ccbafe0d; this is the same class on the
scripted one-shot entry point (#111770).

- `hermes_cli/oneshot.py::_oneshot_exit_code`: 0 only when the turn completed;
  130 interrupted; 2 failed / partial / completed:false; 1 a completed turn
  with no text (unchanged message).
- `--usage-file` now also reports `partial`, `interrupted` and
  `turn_exit_reason`, so a pipeline can tell a budget stop from a Ctrl-C.
- Docs: `hermes -z` exit-code table + usage-file fields in cli-commands.md.

Live probe (run_oneshot with a stubbed _run_agent): partial-with-text 0 -> 2,
budget-with-text 0 -> 2, interrupted-with-text 0 -> 130, failed-with-text 0 -> 2;
completed 0 -> 0 (control), partial-no-text 2 -> 2 (unchanged).
2026-09-16 18:12:34 -07:00
teknium1
da942e4483 fix(config): hedge the config get phantom-key notice — the schema walk has false positives
`_validate_config_key` is a DEFAULT_CONFIG walk, and several live keys are
deliberately unseeded so that a stored value counts as an explicit user pick
(browser.cloud_provider, stt.provider, cron.misfire_grace_minutes,
gateway.proxy_url / relay_url, dashboard.ssh_isolated_idle_grace_s,
computer_use.backend). The runtime reads all of them and Hermes' own wizards
write some of them, so `hermes config get browser.cloud_provider` asserting
"Hermes does not read it" was a false claim.

Match the set-path notice ("Hermes may not read it"): the key still gets
flagged so a leftover misspelling cannot pass unnoticed, but the CLI no longer
states something it cannot know. Docs updated to the same wording; one test
pins the hedge on an unseeded live key (red before this commit).
2026-09-16 17:42:32 -07:00
teknium1
efc0359051 fix(config): inline the phantom-key notice into get_config_value, trim to 2 tests, docs
Slim follow-up to the salvaged commit from #112350 (@DavidMetcalfe): the two
helpers (`_is_unread_nested_key`, `_print_unread_key_notice`) collapse into
~10 lines at the single call site in `get_config_value`, reusing
`_validate_config_key` for both the verdict and the did-you-mean suggestion
instead of calling it twice. `is_env_key` is dropped: `.env` keys and
UPPER_SNAKE settings are never a known top-level section, so the guard
already excludes them.

Tests: keep the two invariants (phantom nested key -> notice on stderr while
stdout/--json stays parseable; known / custom top-level / open-subkey keys
are not flagged). Dropped `test_missing_key_exits_without_a_notice` — the
missing-key exit precedes the print and is pre-existing behaviour already
covered elsewhere.

Docs: one sentence each in cli-commands.md (`config get` row) and
configuration.md (the set-path did-you-mean paragraph).
2026-09-16 17:42:32 -07:00
teknium1
8899aeff53 fix(oneshot): --usage-file ledger reports auxiliary LLM spend
`hermes -z --usage-file` copied only the main-loop result, so title generation,
vision, compression, web_extract and background-review calls — recorded per task
in session_model_usage — never reached the pipeline ledger the flag advertises
as "so pipelines can always account for spend". The Insights page already folds
those rows in (#23270); the ledger is now consistent with it.

- SessionDB.auxiliary_usage_by_task(session_id): per-task sums over the
  session's compression lineage (aux calls bill to the id the turn started with
  while compression mints child ids mid-turn).
- oneshot snapshots aux usage before the turn and attaches the delta after it,
  so a resumed session's earlier runs are not re-billed.
- The report gains `auxiliary` (totals + `by_task`) and
  `total_including_auxiliary`; every existing key keeps its main-loop meaning.
- The auto-title upgrade runs on a daemon thread and can still be writing when
  the turn returns: title_generator tracks in-flight upgrade threads and
  oneshot joins them (bounded) before reading — no sleep, no eager read.

Fixes #112848. Direction shared with #112852 (@KoNit-K), which folded aux into
the headline counters; this keeps them backward compatible instead.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:03:25 -07:00
teknium1
3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1
08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1
f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1
23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1
da9810387d docs(config): scope the .env routing claim to registered env settings
Only names in OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS and the platform *_HOME_CHANNEL /
*_ALLOWED_USERS suffix family route to .env; other documented ALL-CAPS names still land in
config.yaml as top-level scalars with a notice. Say so instead of 'every documented
environment variable' (#111848 stays open for the remaining names).
2026-09-15 18:28:49 -07:00
teknium1
0e63a1bc5c fix(config): refuse an unknown path under a known section before writing
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.

Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.

`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.

Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:28:49 -07:00
teknium1
d1bd778a5e fix(checkpoints): clear-legacy trims to two real-path tests, aligns failure line with the log
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.

Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:24:22 -07:00
teknium1
a9a8a3fa2e fix(config): hermes config get masks credentials on every path; --raw opts out
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.

`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.

Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).

Fixes #110758
Fixes #84106
2026-09-15 05:08:55 -07:00
teknium1
7bb52c0b74 feat: profile clone can opt into staying synced with its source (--sync-imports)
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.

`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
2026-09-15 04:44:09 -07:00
teknium1
8731bb91f5 refactor(gateway): move the supervised-restart handback into gateway_supervised_restart.py
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.

Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
2026-09-15 04:07:13 -07:00
Teknium
6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
Alan
1657a1ce2d feat(cli): add --format stream-json for structured JSONL output
Adds a --format flag to hermes chat single-query mode. stream-json
emits newline-delimited JSON events (init, text, tool_use, tool_result,
result envelope with token stats + exit code) to stdout for CI
pipelines and external tooling. Session ID stays on stderr.

Salvaged from PR #12278 by @ProDrifterDK onto current main, including
the follow-up commit enforcing the single-query contract (implies
quiet, rejects --tui, emits a final result record with exit code 130
on interrupt).
2026-09-15 03:53:13 -07:00
teknium1
931387dd42 test(cron): two invariant ESTOP tests for the fire webhook and misfire backstop; docs
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.

Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
2026-09-15 03:33:13 -07:00