Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
Two invariant tests (red on origin/main): a 401 rotation onto a Codex pool
row keeps the HERMES_CODEX_BASE_URL target, and model.base_url under
model.provider: openai-codex resolves for pool credentials. Adds the
previously undocumented HERMES_CODEX_BASE_URL row to the environment
variables reference so proxy users can find the knob and its reach.
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.
Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).
Co-authored-by: himanusia <himanusia@users.noreply.github.com>
Precedence is session /model > channel_overrides > config.yaml, so dropping the session
override after a --global write regressed chats whose channel_overrides names a model: the
confirmation said "switched to gpt-5.5" while the next turn resolved 'channel-model'. The
cleanup now runs only when no channel_overrides entry applies to the source; otherwise the
override stays (and is written through) so the confirmation stays true.
A failed set_model_override(key, None) was logger.debug only while the reply claimed a clean
save and memory had already popped the override — the durable stale copy would shadow
config.yaml on the next restart (the original #100314 symptom). The cleanup failure is now
returned like the config-write error: the in-memory override is kept, the confirmation carries
a warning line instead of "Saved to config.yaml", and a riding --reasoning stays session-scoped.
Tests (tests/gateway/test_model_picker_persist.py, both red on the previous head): the
channel_overrides case drives _handle_model_command with a real JSONL SessionStore and asserts
_resolve_session_agent_runtime(source) yields gpt-5.5; the cleanup-failure case asserts the
warning and the kept in-memory override. Docs note the channel exception and that the CLI/TUI
keep their per-session pin by design (resume restores the model that chat used).
A `/model <m> --global` switch (typed or picker) used to persist twice: the profile
config.yaml AND a per-session `model_override` in the session store. The session copy
has higher precedence on rehydration, so after a later global change (CLI `hermes model`,
another chat's `--global`) and a gateway restart the stale override silently won —
#100314 saw an explicit `gpt-5.6-sol-900k` resume as the base 272K `gpt-5.6-sol`.
Now `_record_model_switch` writes config.yaml FIRST and, on success, drops the redundant
session override from memory and the store. If the config write fails the switch stays a
truthful session override and the confirmation says "config.yaml not updated (...)" plus
the session-only hint instead of claiming "Saved to config.yaml"; a riding `--reasoning`
pin follows the same effective scope. `--session` and `--once` semantics are unchanged.
Tests: two invariants in tests/gateway/test_model_picker_persist.py (typed+picker clear
the durable override and a fresh SessionStore rehydrates nothing; failed config write keeps
the override and an honest reply), red on base. test_model_command_request_overrides now
points get_hermes_home at its own config so the --provider switch resolves session-scoped
as intended instead of the sandbox's fresh-install first-pick rule.
Fixes#100314
Supersedes #99825 (slim redo; the original wrapped the commit boundary through a
sys.modules-swapped mixin).
Co-authored-by: Andrex Ibiza, MBA <andrexibiza@gmail.com>
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.
Part of #95794
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.
Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.
Default behavior is unchanged; all 291 existing discord tests pass.
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.
Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.
- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
admits provider="custom" only when `codex_model_provider_id()` finds a
configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
`.modelProvider` (fields verified against the codex 0.147 app-server
schema). `_ensure_codex_session` fills them only for custom agents; codex
resolves base_url/env_key from its own `[model_providers.<name>]`, so the
Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
Desktop/TUI agent knows the provider id (it otherwise collapses to
"custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
bare-`custom` ineligibility.
Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).
Fixes#75186
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.
Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.
`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:
1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
generations, usage rows, system prompt) to the owning store, parents before
children so lineage survives, copy-then-delete so a crash leaves a duplicate
the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
-> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
import re-injects them into routing every boot).
Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).
Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).
Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).
_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.
Fixes#114644
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.
`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.
Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).
Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
"slash command /<name> unavailable — name taken by built-in; use /skill <name>"
- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
`hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
carries the notice in its existing `warning` field (rendered by the Desktop
`/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
listing (Telegram/Discord/…) appends one ⚠ line per colliding skill
Fixes#113560
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.
Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):
- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
a base_url is: the target's registry/plugin host, a `providers:` /
`custom_providers:` entry resolving to the target, or bare custom/local
aliases (configured BY base_url) -> owned; another known provider's host
or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
alone, with no base_url, is old-route wire state and goes too); keeps an
owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
prints what was cleared and why, or warns that an unrecognised base_url
still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
(openai-codex + chatgpt.com), a custom entry with that URL, and bare
`custom` are untouched.
Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
Widen the salvaged fixes to the whole class and add the pieces they missed:
- The default User-Agent for SDK-built OAuth requests moves from the manager's
bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
provider gets it too, and the manager's pre-flight metadata discovery (its own
client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
registration and token requests through client.send(), which never merges the
client's default headers; www.tradingview.com's WAF answers a header-less GET
with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
/register and /authorize on the MCP host, and the login died with
"Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
... != None" because no issuer was ever discovered).
- The callback listener now runs serve_forever() and is shut down before
server_close(). A thread parked in handle_request()'s select() keeps the
closed listening socket alive (the kernel holds the file for the duration
of the poll), so a flow cancelled mid-wait left the port bound and the
retry on the same pinned/cached port raised "OAuth callback port N is
already in use" with no external collider.
- When every authorization-server metadata fetch failed, a registration error
is re-raised leading with those statuses ("Could not read
authorization-server metadata (403 from ...); dynamic client registration
then fell back to a guessed endpoint on the MCP host and failed: ...").
humanize_oauth_registration_error leaves that message alone so the 403 in
it is not mistaken for a DCR allowlist refusal.
Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.
The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.
Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).
Part of #114471 (items 2, 5, 6).
Salvage follow-up to the previous commit (#114397 by @Finn763):
- Codex/Copilot rows went through cached_provider_model_ids directly, so a
cold cache on the non-blocking read path rendered an EMPTY Copilot row
(live repro: copilot:0). Route them through _live_or_curated_ids like
every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
_mark_catalogs_pending: no surface consumes it and it would have needed
a gateway contract regen. Drop the _spawn_background_warm wrapper: the
ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
every endpoint re-ran four chat-completion probes on every
credential-pool load (load_pool("zai") runs several times per picker
open; the reporter's logs show exactly these repeated POSTs). Memoize
the failure in-process for 5 minutes. Copilot already has the same
negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
open + row still renders; explicit refresh still probes) plus one for
the Z.AI negative cache; a rigid test fake gains **kw for the widened
cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.
Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.
Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.
Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.
Part of #114271
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.
- list_profiles(lazy_skill_count=True): skill_count is the last known value;
a missing/aged entry schedules ONE background _count_skills per profile per
60 s recheck window, so the request thread does zero skill-tree I/O and the
refresh cadence is decoupled from the poll rate. The two polled callers and
the REST fallback entry use it; the sync default (CLI, detail views) is
unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
os.walk with excluded/support dirs pruned) instead of rglob — half the
fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
skill add/remove immediately on the next refresh).
Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
profiles.list 2189 + 2015, projects/tree 2177 + 2015;
a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
after: GET /api/profiles 0 skill-tree stats/scandir on the request thread
(counts land from the background refresh by the next poll),
profiles.list 0, projects/tree 0; _count_skills (detail/control)
still reports 200 with 203 scandir; the vanished-dir case returns
199 and list_profiles() enumerates all 5 profiles.
Fixes#114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.
Docs: doctor reference and the network config section describe the check.
Follow-up to the salvaged hermes_constants expansion (#109212): the CLI has
~30 raw `os.environ["HERMES_HOME"]` readers (fast --version path, early
display.interface probe, profile re-home, dotenv loader, subprocess-home
helper) that never go through get_hermes_home(). A literal `~` (fish, or
any quoted value) left them resolving `~/.hermes` against cwd while the
resolver now expands it, so the process would disagree with itself.
Normalize the env var once, at the earliest point of hermes_cli/main.py
(stdlib-only, before the fast paths), and expand it in the one raw
hermes_constants reader (`_profile_home_path`). A relative value that is
not tilde/variable-shaped is deliberately left alone: rejecting it would
break `HERMES_HOME=./tmp-home` in tests and CI for no user-facing gain.
Tests: real-CLI subprocess with HERMES_HOME='~/.x' under a fake HOME asserts
`config path` lands in the fake home and that no literal `~` directory
appears under cwd (the reporter's acceptance criterion, #114353), plus a
unit test for the normalizer. Docs: environment-variables reference.
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.
agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.
Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.
Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.
Config-set atom of #114471.
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).
What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
(spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
`_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
`/proc/<pid>/environ` never reflects it. Without this, the common default-home install
(nothing exported) got a `--stop` that matched nothing — the regression flagged on
#113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
#113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
only a foreign backend prints "No hermes dashboard processes running for this profile."
instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
the (now filtering) finder. Docs: `--stop` row in the CLI reference.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).
resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.
Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).
- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
inherited UI transport label without an explicit --source) resolves to `oneshot`; the
platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
(HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
`sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
`--source` flag reference (explicit flag always stored as given).
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could
fail without the user ever seeing a question, and the agent then waited out the full
clarify_timeout and reported "[user did not respond within Nm]" (#112684):
- the platform rejects the card at once -> the wait aborted with a delivery sentinel but the
question was never re-asked;
- the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts,
a stale-thread retry) -> the runner never looked at the send future again and blocked for
the whole timeout for a card that never posted;
- no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported
timed_out=True with an empty notice.
gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers
move out of the gateway/run.py facade) now retries every definitive card failure once through the
adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress
DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a
possibly-delivered send so a late failure releases the waiter with
"[clarify prompt could not be delivered]". The no-surface case reports
"[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already
treats as a non-answer).
Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s):
card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and
the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt,
after the numbered prompt is sent and "2" resolves to the choice.
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.
_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.
Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.
The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).
Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
Closes the last atom of #112615. A Desktop app started from the Dock, Finder or a
desktop launcher inherits a minimal environment, so the GITHUB_TOKEN / GH_TOKEN rung
landed in #113190 never fires for exactly the users the issue is about (shared exit
IP, 60/hour anonymous budget exhausted by neighbours). The update check now walks the
same ladder as the Python GitHub client (tools/skills_hub_github.py::GitHubAuth):
1. GITHUB_TOKEN, then GH_TOKEN, from the launch env (unchanged).
2. `gh auth token` — execFile with an argv array (no shell), stdin closed, 3 s
timeout, windowsHide. `gh` is resolved on PATH plus the GUI-safe install
locations backend-env already appends for the backend (Homebrew, /usr/local,
~/.local/bin, nix profile, the Windows GitHub CLI installer dirs). The outcome
(token or none) is cached for the process lifetime so gh runs at most once.
3. Anonymous.
A rejected credential (401) still retries anonymously; the log line names the
source (env vars vs gh login) and never the token, once per source per process.
A rejected gh token also drops the cache so a re-login is picked up by the next
check. envTokenRejected is renamed githubTokenRejected since both rungs use it.
Live: with PATH=/usr/bin:/bin:/usr/sbin:/sbin and no env token, the module found
gh, resolved a credential in 135 ms and api.github.com reported a 5000/hour core
budget (anonymous: 60). A logged-out gh and a hanging gh both resolved to null
(anonymous) in 5 ms / 3001 ms.
Docs: the credential ladder is documented on the Desktop "Updating" page and in
the environment-variables reference.
The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.
Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.
The two one-shot entry points map the same outcome to different codes
(failed/partial/budget -> 2 on -z, 1 on -q/-Q; completed-with-no-text -> 1 vs 0)
and the reference documented both tables 60 lines apart without saying so.
Also `iteration_limit` is a metrics end_reason, not a turn_exit_reason; the
finalizer emits `max_iterations_reached(N/M)`.
`hermes -z` judged the run by whether stdout was non-empty: a turn that hit the
iteration budget, ended `partial`, failed on the provider, or was interrupted
still exited 0 as long as it printed an explanation, so scripts treated a
half-done job (or a 401 summary) as success. The `-q`/`-Q` paths already got
one result-derived contract in 23863ccbafe0d; this is the same class on the
scripted one-shot entry point (#111770).
- `hermes_cli/oneshot.py::_oneshot_exit_code`: 0 only when the turn completed;
130 interrupted; 2 failed / partial / completed:false; 1 a completed turn
with no text (unchanged message).
- `--usage-file` now also reports `partial`, `interrupted` and
`turn_exit_reason`, so a pipeline can tell a budget stop from a Ctrl-C.
- Docs: `hermes -z` exit-code table + usage-file fields in cli-commands.md.
Live probe (run_oneshot with a stubbed _run_agent): partial-with-text 0 -> 2,
budget-with-text 0 -> 2, interrupted-with-text 0 -> 130, failed-with-text 0 -> 2;
completed 0 -> 0 (control), partial-no-text 2 -> 2 (unchanged).
Replace the two salvaged per-command test files with a single suite that drives
the real /branch, /resume and /sessions <id> handlers against a real SessionDB
and asserts the invariant that matters: while a turn is running, the session id
(CLI and agent) is unchanged and the current row is not ended; when idle,
/branch still proceeds. Red on origin/main (3 busy cases), green with the guards.
The /branch MagicMock fixture gains an explicit `_agent_running = False`: a bare
MagicMock attribute is truthy and would trip the new guard.
Docs: /branch and /resume rows in the slash-command reference say they are
refused mid-turn in the classic CLI and why.
Co-authored-by: sonderhq <144971773+sonderhq@users.noreply.github.com>
Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
Follow-up to the #112725 salvage. When a gateway clarify batch stops at an
unanswered question, the batch payload only said `timed_out: true` with blank
answers — so a card the platform REJECTED ("[clarify prompt could not be
delivered]") read exactly like a user who walked away, which is the misreport
#112684 describes ("a timeout must not be presented as user inactivity when the
prompt was never delivered").
- gateway/run_turn_runner.py::_clarify_batch_sync: the unanswered question's
own response text rides along as `notice`.
- tools/clarify_tool.py::_run_batch/_batch_result: pass `notice` through into
the result JSON beside `timed_out` (only when present); schema description
mentions it. CLI/TUI batch callbacks send no notice -> byte-identical output.
- tests: fold the single-question re-arm pin into the existing bracket-answer
test (<=2 new tests), parametrize the end-to-end test over an undeliverable
Telegram-shaped adapter so the delivery sentinel is pinned as `notice`.
- docs: tools-reference.md describes `notice`.
Part of #112684
`_validate_config_key` is a DEFAULT_CONFIG walk, and several live keys are
deliberately unseeded so that a stored value counts as an explicit user pick
(browser.cloud_provider, stt.provider, cron.misfire_grace_minutes,
gateway.proxy_url / relay_url, dashboard.ssh_isolated_idle_grace_s,
computer_use.backend). The runtime reads all of them and Hermes' own wizards
write some of them, so `hermes config get browser.cloud_provider` asserting
"Hermes does not read it" was a false claim.
Match the set-path notice ("Hermes may not read it"): the key still gets
flagged so a leftover misspelling cannot pass unnoticed, but the CLI no longer
states something it cannot know. Docs updated to the same wording; one test
pins the hedge on an unseeded live key (red before this commit).
Slim follow-up to the salvaged commit from #112350 (@DavidMetcalfe): the two
helpers (`_is_unread_nested_key`, `_print_unread_key_notice`) collapse into
~10 lines at the single call site in `get_config_value`, reusing
`_validate_config_key` for both the verdict and the did-you-mean suggestion
instead of calling it twice. `is_env_key` is dropped: `.env` keys and
UPPER_SNAKE settings are never a known top-level section, so the guard
already excludes them.
Tests: keep the two invariants (phantom nested key -> notice on stderr while
stdout/--json stays parseable; known / custom top-level / open-subkey keys
are not flagged). Dropped `test_missing_key_exits_without_a_notice` — the
missing-key exit precedes the print and is pre-existing behaviour already
covered elsewhere.
Docs: one sentence each in cli-commands.md (`config get` row) and
configuration.md (the set-path did-you-mean paragraph).
Trim the salvaged ladder to its env rung and make the 403 copy actionable.
- github-api-auth.ts keeps githubTokenFromEnv / githubApiHeaders; the
`gh auth token` rung (resolveGithubToken, readGhCliToken, the per-process
cache and the execFile import in main.ts) is dropped. Whether a GUI app may
silently spend the user's gh CLI login on a passive background check is a
product call, not a bug fix; the pure seam makes re-adding it a ~15-line
follow-up.
- fetchGitHubApi reads GITHUB_TOKEN / GH_TOKEN from process.env per request
(nothing stored) and attaches x-ratelimit-remaining / x-ratelimit-reset plus
whether the call was authenticated to the rejected error.
- describeUpdateCheckFailure moves out of the unimportable main.ts into
update-api-check.ts. A 403/429 with x-ratelimit-remaining: 0 now says the
anonymous budget is 60/hour per network address shared with everyone behind
the same connection, gives the real reset time and names GITHUB_TOKEN as the
remedy; a 403 without rate-limit headers is reported as a plain HTTP 403
instead of a rate limit the user did not cause.
- Docs: desktop.md "Updating" and the GITHUB_TOKEN row in
environment-variables.md.
- Tests trimmed to two invariants per module.
Why: behind a shared exit IP "try again in an hour" is advice the user cannot
act on, and the old copy hid the one thing that does help.
Both readers now share utils.env_var_enabled(), so the reference table
spells out the falsy values instead of implying "any value enables".
Also registers the contributor email mapping for the salvaged commit.