Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.
Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.
An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
The switched-provider comment implied the guard catches the resolver raise;
it is suppress() leaving st.* at the session values that does. The
stay-on-provider comment now states the one semantic widening of moving to
the shared helper: an empty resolver result keeps the session endpoint on
any host, not only off openrouter.ai.
Gate finding: resolve_runtime_provider(requested="custom") never returns an
empty base_url — with no model.base_url but an OpenRouter key present it
lands on OpenRouter's default host, so the "keep the current endpoint"
fallback was dead and an Anthropic session adopting `provider: custom`
hopped to openrouter.ai. The guard _creds_for_current_provider already
carries for this (#74143) is now a shared helper used by both branches;
the branch assigns once instead of assign-then-undo; the import sits with
the other from-imports; the new test's write_text carries an encoding.
Second invariant test pins the no-endpoint case.
The bare-custom credential branch kept the session's current base_url and
key whenever the target was `custom`, regardless of where the session was.
That is right for a bare-custom session picking another model on its own
endpoint, but when the TUI's per-turn config sync adopts `provider: custom`
from an OpenRouter (or any other) session, the new model was paired with
OpenRouter's URL and key and every request 400/404'd (Bug 2 in issue 73680).
Arriving from another provider now resolves the configured custom endpoint
through resolve_runtime_provider; with none configured the current endpoint
is kept, so direct aliases still supply their own URL afterwards and the
bare-custom same-endpoint case is unchanged.
`display.show_reasoning` is a client-side display decision since 40f2368875,
and the Desktop transcript honors it through `$showReasoning` (#115335). The
composer path did not: `/reasoning` was marked `desktop="advanced"` in the
Python slash registry, so the Desktop refused it as "not available", while a
gateway-routed `/reasoning hide` would only have written config.yaml and left
the atom waiting for the next config refresh.
Route `/reasoning` to the gateway's `config.set key=reasoning` — the Ink TUI's
path (ui-tui/src/app/slash/commands/session.ts) — and mirror the answer:
`hide`/`show` flip `$showReasoning` on the spot, effort levels leave the gate
alone and get session scope (a throwaway slash-worker CLI could never pin the
live session's effort). No arg reports the current effort and display, like
Ink. Drops the registry's `advanced` marker and regenerates the desktop dump.
Live (real `hermes serve`, scratch HERMES_HOME, real WebSocket): config.set
key=reasoning value=hide -> {value: hide}, config.yaml show_reasoning=False;
show -> True; high -> effort only, display untouched.
Fixes#111761
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.
- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
current users). When false:
- `read_claude_code_credentials()` — the only reader of the borrowed Claude
Code login — returns None, so the resolver fallback, the expired-token
refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
file; the pool prunes a `claude_code` row an earlier adopting process
persisted.
- `_recover_codex_tokens_from_cli` returns None for both automatic recovery
paths (rejected refresh, half-empty singleton); the real AuthError is
surfaced instead. The interactive import offer in `hermes auth add
openai-codex` still asks first and is unaffected.
- One INFO line per process the first time adoption would have happened;
`hermes auth list` / `hermes auth status anthropic|openai-codex` print the
same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.
Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).
Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.
`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.
No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.
Part of #114587
Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
A dispatcher-spawned worker whose turn failed on a provider error that no
retry can heal (credential rejected: auth / auth_permanent, model_not_found,
ssl_cert_verification) now exits KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78,
BSD EX_CONFIG) instead of a plain 1, so the dispatcher can tell "the
provider will reject every further spawn" from "this attempt crashed".
The classification is the agent's own FailoverReason verdict already on
the turn result (failure_reason) — no error-text regex. billing stays in
the transient set: credit comes back, a revoked key does not.
Shared by both one-shot paths (-q and -Q) and gated on HERMES_KANBAN_TASK,
so a person's `hermes chat -q` keeps exiting 1 on the same error.
Part of #114587
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.
The venv-repair and runtime-repair paths reinstalled core, lazy and tool
deps but never refreshed memory-provider bridge packages, unlike the
git-pull and ZIP paths. Both repair paths now heal them last.
Fixes#113741.
The Windows CRT reports isatty()==True for every character device, the NUL
device included, so `hermes gateway start < NUL` (stdin=DEVNULL) still took
the TTY branch: prompt_yes_no() hit EOF, defaulted to Yes and registered a
Scheduled Task — the silent install #113977 describes. Confirm isatty with
GetConsoleMode(STD_INPUT_HANDLE); a handle no console accepts is treated
like a pipe. Pipes and HERMES_NONINTERACTIVE keep their behaviour.
`start()` asked "Install it now so the gateway starts on login?" with default Yes, then called
`install()` which asked two more default-Yes questions; on a redirected stdin (`< /dev/null`,
scripts, agent tool calls) all three answered themselves and a bare `start` wrote a Startup-folder
login item, spawned the gateway, and start() spawned it a second time (#113977).
- The persistence question is answered only explicitly: `HERMES_GATEWAY_INSTALL_START_ON_LOGIN`
(which start() never consulted before) or a prompt on a real TTY. Without a TTY, or under
`HERMES_NONINTERACTIVE=1`, the answer is No: the gateway starts, nothing persistent is written,
and the explicit `hermes gateway install` command is printed.
- Declining on a TTY also starts the gateway — the command is `start`; only persistence was declined.
- A Yes hands `start_now=True, start_on_login=True` to install(), which spawns and reports (including
the UAC hand-off), so start() no longer double-spawns and no longer prints the false
"Gateway install did not complete in this process" after an install that was skipped by intent.
- `install --no-start-on-login` without `--start-now` now hints `hermes gateway run` instead of the
`hermes gateway start` that used to offer the very auto-start just declined.
Supersedes #113981 (@KoNit-K), which flipped the first prompt's non-TTY default but left the
env override, the hint text and the false warning in place.
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.
`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.
Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
A provider whose catalog default has no path (api.anthropic.com) has no
sibling OpenAI-compatible path to infer an override's protocol from, so the
declared transport stays; the docstring says so and a test case pins it.
Round-2 gate folds: utils.base_url_path sits next to base_url_hostname on
the same scheme-tolerant parser, so a scheme-less override (api.minimax.io)
keeps host and path checks agreeing; a default with no path is an explicit
"nothing to compare" case; the dead `or "chat_completions"` after
determine_api_mode is gone; the test fixture seeds the catalog defaults
unconditionally so a warm models.dev cache cannot change what the contract
compares against, and the stale #53054 test comment states the new contract.
Gate findings folded: the accepted path is derived from the provider's
catalog default (/anthropic for MiniMax, /plan/anthropic for Tencent) instead
of a hardcoded /anthropic prefix, so /anthropic-compat no longer passes and a
non-/anthropic default provider's own endpoint does; an empty base_url keeps
the declared transport like an unknown default; the catalog read is
allow_network=False (determine_api_mode just warmed it); _base_url_path is
shared with URL detection. Tests pin the catalog defaults they compare
against so the contract holds offline.
33a2f29a6 made the runtime fallback honour the provider's declared transport
when URL detection has no opinion. For minimax/minimax-cn that is
anthropic_messages, which is right on api.minimax.io/anthropic but wrong for a
model.base_url override at an OpenAI-compatible relay (LiteLLM, token proxy,
corporate egress) or at the provider's own /v1 path: every request is built
Messages-shaped with x-api-key and 401s, or silently reroutes to a fallback
provider.
The fallback now keeps anthropic_messages only when base_url shares the
catalog default's host and is the bare host or a path under /anthropic;
everywhere else it lands on chat_completions. Other declared transports are
untouched (openai-api on a custom proxy still keeps codex_responses, as
33a2f29a6 intended), and URL detection still runs first.
Salvage of PR 76859 by webtecnica, rebuilt on the refactored module with the
path check teknium asked for in review.
Co-authored-by: webtecnica <webtecnica@gmail.com>
`hermes dashboard` from a named profile re-execs as `-p default dashboard --open-profile P`,
but `--open-profile` only reached the one URL `_maybe_open_browser()` opened. A `/chat?resume=<id>`
deep link without `?profile=` initialised `ProfileProvider` with the empty (launch) scope, so the
embedded chat ran in the default profile — no MCP servers, wrong model/skills — while the switcher
showed P. No error, no indication.
The server now records `initial_profile` on `app.state` and injects it into the SPA bootstrap as
`window.__HERMES_INITIAL_PROFILE__` (escaped for the script context); the Vite dev proxy forwards
it. `ProfileProvider` uses it only when the URL carries no `profile` param: an explicit
`?profile=` (including an explicit empty one) still wins, and the sticky-active-profile alignment
no longer replaces a launch-preselected scope.
Closes#73085. Salvage of #73260 (cherry-pick of 99e341cdef0 resolved onto the split
web_server_dashboard.py; start_server assertion dropped as a change-detector).
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).
- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
`served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
action env already funnel through (the dashboard site now calls it for its target).
Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
the profile is not the one the updater runs as; the module stays stdlib-only at
import time.
Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.
Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
Gate review: `degraded` became a whole-life serving state but four
sibling predicates only accepted `running`: derive_gateway_busy /
derive_gateway_drainable (the NAS drain gate reported a degraded gateway
with in-flight turns as idle and undrainable), the stale-heartbeat
detectors in `hermes gateway status` and the dashboard, and the Windows
doctor probe. All four accept `degraded`; a dead watchdog-stamped
`degraded` is still excluded by `gateway_running=False`.
The parked-platform ERROR pointed at `/platform resume`, which only
resumes platforms in the retry queue - a non-retryable failure never
enters it. The remedy is `hermes gateway restart`.
Test: a degraded gateway with active agents is busy; not live => not
drainable.
Gate review on the stack:
- `hermes gateway restart` under systemd only accepted `gateway_state == "running"`
as proof of the replacement; a boot with a parked platform stamps `degraded` for
its whole life, so every restart/update on such a host waited out the 60 s+
budget and reported a false failure while the gateway was serving. The verifier
now accepts `degraded` as restarted and prints one DEGRADED warning.
- Drain release and scale-to-zero wake re-stamped `running` unconditionally, wiping
the parked-platform signal after the first `.drain_request.json` cycle. Every
"we are serving" stamp now goes through `_serving_state()`; the mixed fatal +
retryable boot path sets the flag too (it fell through as a plain run before).
`_startup_parked_platforms` is a bool — the joined error text was only ever logged.
- `_wait_for_tcp_port_free`: an unresolvable or unreachable configured host raised
a non-refused OSError on every probe and burned the full 10 s wait; only a
connect timeout means "listener alive", anything else means nothing to wait for.
- Windows `restart()` replaces its blind `time.sleep(1.0)` "let Windows release the
port" with the same configured-address wait.
- The api_server bind retry rebuilds the AppRunner per EADDRINUSE attempt instead of
calling aiohttp's private `_unreg_site`; attempt count is a named constant.
Test: a `degraded` replacement is reported as restarted (red on the previous head).
E2E re-run: predecessor releases the port 0.5 s after the first bind → bound on the
third attempt (+0.61 s), connect() True.
Folds on the salvaged #92060 (@eliasburlison):
- `_wait_for_api_server_port_free` read only API_SERVER_HOST/PORT from the
environment, so an api_server port set in config.yaml (`platforms.api_server.port`)
was never waited on. The adapter's host/port resolution is now the shared
`gateway.platforms.api_server.listen_address()`, used by both the adapter and the
restart path, and the wait is skipped when api_server is disabled (a foreign
listener on the default port is nobody's race).
- Only ECONNREFUSED means the listener is gone. A timed-out connect (full accept
queue on a draining predecessor) was reported as "free", which would have let the
replacement start straight into EADDRINUSE again.
- The bind retry in `APIServerAdapter.connect` creates a fresh TCPSite per attempt
and unregisters the failed one: aiohttp registers the site before binding, so
re-starting the same object raises "already registered".
- The injected clock/sleeper/connect knobs are gone; the tests bind a real listener.
Tests: a real listener closed 300 ms in → wait returns True; a busy port taken from
config.yaml (env unset) → wait reports busy while enabled and is skipped when disabled.
E2E: with the predecessor releasing the port 0.5 s after the first bind, origin/main
logs `Errno 48 ... address already in use` and connect() returns False; this head
binds on the third attempt (+0.61 s) and connect() returns True.
On macOS api_server cannot SO_REUSEADDR, so replacing the process as
soon as the old PID exits still hits EADDRINUSE. The new process then
stays up with no API. Wait for the listen port to refuse connections,
and retry the bind a few times before treating the conflict as fatal.
Fixes#91547
(cherry picked from commit 5cb9545023646277b2fb59ebfb5e89b0fd161aba)
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
Third entry path with the same egress: `_resolve_startup_runtime` seeded the model
from HERMES_INFERENCE_MODEL and went straight to static provider detection, so the
qualified string reached the configured default. Route through
`resolve_startup_model_route` first, as HermesCLI and oneshot now do.
Also: an unconfigured `custom:<typo>:<model>` now logs a warning and leaves the
route undecided instead of steering a garbage model id at the bare custom endpoint,
and the redundant `qualified_model != raw` guard is gone.
`hermes -m custom:jetson-vllm:nemotron-nano-30b` (and `-z`) kept the configured default
provider and sent the UNSPLIT string as the model name, so the whole prompt reached
api.anthropic.com before a 404 — a local endpoint the user picked never saw the request.
`parse_model_input` already decodes the form; nothing on the startup path called it.
`resolve_startup_model_route`, the owner of startup provider/model decoding, now consumes
the colon-qualified form through `parse_model_input` ahead of the `provider/model` branch,
so HermesCLI picks it up unchanged. The oneshot resolver routes through the same owner
before provider auto-detection instead of only after a direct-alias miss.
Live, config `providers.jetson-vllm` at 127.0.0.1:8000 and `model.provider: anthropic`:
before — HTTP 401 from Anthropic, local sink idle; after — sink receives
POST /v1/chat/completions model=nemotron-nano-30b, Anthropic untouched.
Closes#73943. First reported fix: PR #74082 by @Drexuxux (HermesCLI.__init__ decode); the
seam moved to the startup route owner during the Sep-2026 cli.py decomposition.
Co-authored-by: Drexuxux <drexux0@gmail.com>
The recents window is a LIMIT-cap page by recency plus a back-fill of pins
the page missed. Both truncation flags — the batched /sidebar route and the
legacy per-slice derivation in the app — discounted every pinned row, on the
theory that pins only ever arrive past the LIMIT. Pins inside the window
occupy real slots, so with two pinned rows among the twenty newest the count
came to 18 < 20, `profiles_truncated` read false, and the sidebar never
mounted "Load more": every older session was unreachable (#81484).
Count every row. A list shorter than the cap has nothing past the page for
the back-fill to add, so pins cannot fake a full page there; at worst a list
of exactly cap rows costs one click that resolves to an empty page.
Salvaged from #81490.
A dashboard write scoped to another profile (POST /api/model/set ?profile=B
runs inside _profile_scope(B)) reached reconcile_record() unconditionally, so
_config_stamp()/_inventory_other_providers() resolved B's home and B's provider
flipped the LAUNCH profile's record to provider_configured=True with a
setup.ready broadcast. Bail out when get_hermes_home() is not the process
home; the negative case is added to the surface test, and the model.save_key
hook is now pinned by its own test.
The serve process's free-tier boot record (`free_tier_bootstrap._record`) is
built once; when the boot inventory found nothing (mint failed or the tier is
off) it says `provider_configured: false` for the process lifetime and
`setup.status` answers from it, so the Ink chat (dashboard /chat, `hermes
--tui`) parks every new session on "Setup Required" no matter what the user
configures afterwards — the Models page, a picker key save, or `hermes setup`
from a shell all land on disk and change nothing in the running process.
Reconcile the record on read instead of re-minting on write: `reconcile_record`
re-runs the cheap inventory (`_inventory_other_providers`, the resolver ladder
with the free-tier rung hidden) when a `False` record's config files
(`config.yaml` / `.env` / `auth.json`) moved since the inventory that built it,
replaces the record's inventory half and broadcasts `setup.ready`. The mint
verdict is kept as is: only the boot bootstrap and its retries mint.
`wait_for_record` (what `setup.status` reads) reconciles, which also covers
writes this process never saw (`/setup`'s `hermes setup` handoff, `hermes model`
in docker exec, a hand edit); the two in-process write paths from #114708
(`/api/model/set`, `model.save_key`) call it for the immediate broadcast.
Dropped from #114708: `retry_bootstrap_mint(force=True)` on the write paths — a
forced portal mint (network, cooldown bypassed) inside an HTTP write; the record
only needed its inventory refreshed. Not taken from #108773: OR-ing the record
with `_has_any_provider_configured()` — that first-run guard counts host-wide
credentials (other host-wide agent-CLI logins) and answers True on a blank machine (verified
live on this host), which would open the gate with nothing configured.
The 4001 hint for a sessionless `config.set model` named "Settings -> Models",
a Desktop-only surface; the Ink chat has no Settings and the dashboard has a
Models page. One string now names /setup, the Models page and Settings ->
Models. The TUI's "Setup Required" panel no longer advertises `/model` as the
in-place fix (sessionless picks are refused by design, 75b1b43ec1).
Co-authored-by: chelsealong <chelsealong@126.com>
A serve process whose boot-time free-tier mint failed keeps its bootstrap
record at provider_configured: false for the whole process lifetime, so the
web chat stays gated on "need setup" even after a provider lands on disk via
POST /api/model/set or a picker key save. Re-run the bootstrap inventory on
both write paths so setup.status (and the setup.ready broadcast) follow the
new configuration without a restart.
Fixes#114697
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).
Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
"slash command /<name> unavailable — name taken by built-in; use /skill <name>"
- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
`hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
carries the notice in its existing `warning` field (rendered by the Desktop
`/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
listing (Telegram/Discord/…) appends one ⚠ line per colliding skill
Fixes#113560
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.
Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):
- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
a base_url is: the target's registry/plugin host, a `providers:` /
`custom_providers:` entry resolving to the target, or bare custom/local
aliases (configured BY base_url) -> owned; another known provider's host
or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
alone, with no base_url, is old-route wire state and goes too); keeps an
owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
prints what was cleared and why, or warns that an unrecognised base_url
still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
(openai-codex + chatgpt.com), a custom entry with that URL, and bare
`custom` are untouched.
Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
When a user runs `hermes config set model.provider <new>` to switch
providers, the old provider's `model.base_url` is left behind in
config.yaml. This causes API calls to go to the wrong endpoint.
The wizard flow (`_update_config_for_provider`) already handles this
correctly by clearing stale base_url on provider switch, but the
`hermes config set` path did not.
Now `set_config_value` detects when `model.provider` is being changed
and removes the stale `model.base_url`, allowing runtime
auto-detection to resolve the correct endpoint for the new provider.
Fixes#40862
_has_sticky_block held on 'effective_limit in verdict and failures >= effective_limit',
which is true for every plain unified-budget trip (_record_task_failure only trips when
failures >= effective_limit). That made every breaker trip operator-only and regressed two
recovery paths main relies on: raising the dispatcher failure_limit past the counter
(#35072) and assign_task to a fresh profile (counter reset by design).
_record_task_failure now stamps 'sticky': true on the gave_up payload only for force_trip
(the clean-exit protocol-violation budget) and _account_crashes adds it for the systemic
same-error trip at limit 1; _has_sticky_block reads that marker and nothing else. A plain
gave_up carries no marker and is judged by the counter as before.
The fresh-process sweep test also pins that the decoded rc lands in the run row metadata
(exit_code), so quota (75) vs crash stays tellable after the fact even though the
worker_output tail is trimmed (#113611 related observation).
``_account_crashes`` trips the protocol-violation budget with
``force_trip`` after three consecutive clean exits, but the task's
``consecutive_failures`` is still 1 — so ``recompute_ready``, which
re-derives the threshold from its caller's ``failure_limit`` (default 2),
promoted the just-blocked card straight back to ``ready`` in the same
tick: ``protocol_violation -> gave_up -> promoted -> claimed`` forever
whenever ``failure_limit`` exceeds the violation count (the systemic
same-error trip at limit 1 had the same hole).
``_has_sticky_block`` now also honours the breaker's own verdict: the
newest ``gave_up`` since the last ``unblocked`` holds the card when it
recorded a violation-streak trip or ``failures >= effective_limit``.
``hermes kanban unblock`` remains the release path and still grants a
fresh budget. Tests cover the fresh-process classification (rc=0 and
rc=75), the held trip + operator unblock, and the trailer emission gate.
The dead-worker sweep learned a worker's exit status only from
``_recent_worker_exits``, which ``reap_worker_zombies`` fills via
``os.waitpid`` — so only the process that spawned the worker ever knows
how it exited. With ``kanban.dispatch_in_gateway: false`` every
``hermes kanban dispatch`` tick is a fresh process, the registry is
empty, and a worker that exited rc=0 without a terminal board call was
booked as a bare ``crashed`` / ``pid N not alive``: no
``protocol_violation`` marker, no corrective error text for the retry
worker, no violation streak — and a rc=75 quota wall was counted as a
failure instead of a neutral ``rate_limited`` requeue.
A Kanban worker (``HERMES_KANBAN_TASK`` set) now writes
``[kanban-worker-exit] rc=<code>`` as the last line of its own log on
every one-shot exit path; when the registry has no entry for a dead PID
the sweep reads that trailer and books the exit through the same
code -> kind mapping. A worker killed before its exit epilogue leaves no
trailer and stays a plain crash. ``_worker_final_output`` strips the
trailer so it never leaks into the board diagnostic.
Direction (a durable, process-independent witness in the worker log)
from PR #113638; its predicate keyed on the ``Resume this session with:``
summary, which the CLI prints before ``sys.exit`` for rc 0, 1, 75 and 130
alike and so would have booked quota walls and failed turns as protocol
violations.
Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
_probe_single_server's `connect_timeout` override (e.g. hermes mcp login's
315s, given so a user has time to finish a browser OAuth flow) only bounded
the outer asyncio.wait_for(_connect_server(...)). The deep transport code
(_negotiate_session) reads its own session.initialize() bound straight from
config["connect_timeout"], which was left at the unrelated 60s default. Any
OAuth login taking longer than a minute got its still-pending callback wait
cancelled mid-flow, then retried — reopening a second authorization against
the same still-tearing-down callback port ("port already in use").
Related #103633.
`_update_owes_fleet_restart` held a receipt whose restart phase completed to exactly
`post_update.sha`. A gateway restarted afterwards onto a checkout moved by hand runs code
NEWER than the update pulled, so the remedy the warning itself names (`hermes gateway
restart`) could not clear it until the next `hermes update` wrote a fresh receipt
(#113350 steps 3-4). A completed update is discharged when every owed gateway serves the
code it pulled OR today's checkout; the catch-up predicate `hermes update` runs is
unchanged.
`_live_fleet_covers_receipt` and `_marker_only_restart_obsolete` each looped the fleet
twice to fold `served_profiles` into the covered identities and re-validated a shape
`_fleet_row` already enforces. One `_fleet_covered_gateways(fleet)` answers the
`(kind, profile)` set (or None for an unidentified row) for both.
#114309 expects the dashboard, like `hermes cron status`, to say "scheduler last
ticked X hours ago" when a stopped ticker leaves next runs stranded in the past.
The Cron page only had the per-row `Overdue since` label and no way to date the
outage: GET /api/cron/jobs carried nothing about the ticker.
Every job the dashboard cron endpoints return now carries
`scheduler_heartbeat_age_s` — its own profile's ticker heartbeat age read inside
the same store scope as the job list (None when it cannot be dated) — and the Cron
page renders "Scheduler last ticked 7h ago — jobs that came due since then have not
fired" above the list when the oldest heartbeat among jobs expected to fire is
past the CLI's STALE_AFTER threshold (three missed 60s iterations plus slack).
Paused, disabled and completed jobs never raise the banner, matching the overdue
label's rule. The field is additive, so the response shape and every existing
consumer are unchanged.
Tests: one router test lists two profiles with one stale heartbeat file and checks
each job reports its own age (None for the profile without a heartbeat); one vitest
pins the stale-age selection and the relative label.
The slash handler behind the classic CLI's and the TUI's `/cron` overview and
`/cron list` still printed the stored `next_run_at` verbatim as `Next:` /
`Next run:`, so a stamp parked seven hours in the past read as an upcoming run —
the exact symptom of #114309 on the one CLI surface the branch had left out.
Both rows now go through `hermes_cli.cron._next_run_row`, the same decision
`hermes cron list` makes: past the doctor grace on an enabled, non-paused job the
row becomes `Overdue: <stamp> (7h ago — the job has not fired; is the scheduler
running?)`, while paused/disabled/completed jobs keep the plain label because they
are not expected to fire.
Test: one slash-handler test parks two jobs 7h in the past, pauses one, and checks
that `/cron` and `/cron list --all` flag exactly the enabled one.
`hermes cron list`, the web dashboard Cron page and the Desktop cron panel/sidebar all
rendered a `next_run_at` parked hours in the past as an ordinary upcoming "Next run" —
the only user-visible trace of a scheduler that stopped ticking (#114309). Every
surface now labels a slot past `cron doctor`'s 15-minute grace as overdue (CLI
`Overdue:` row with the lateness, web `Overdue since`, Desktop `Overdue since` label
on the detail panel and sidebar meta) while paused/disabled/completed jobs keep the
plain label because they are not expected to fire.
The CLI's overdue check now parses through `cron.jobs._parse_aware` and `hermes_time.now`
so status/list/doctor and the ticker agree on the instant (mixed offsets, DST folds,
legacy naive stamps read as system-local like the scheduler does), and both the
ordering and the subtraction normalise to UTC: Python compares same-tzinfo datetimes by
wall clock, which is wrong across a DST fold. Status still orders the soonest run by
instant (#113874) and prints the stored stamp.
Tests: the salvaged status tests are trimmed to two invariants (overdue + stale
heartbeat is loud on status AND list; within-grace stays plain on both), the DST-fold
ordering tests freeze the CLI clock as well so their 2026-11 fixtures never start
reading as overdue, and one vitest each pins the web and Desktop helpers.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Address review feedback: `cron doctor` tolerates a 15-minute grace
window (_OVERDUE_GRACE_SECONDS) before calling a next_run_at overdue,
but the new status OVERDUE line fired on any `scheduled < now`, so a
job a few minutes behind the ticker's own cadence could flash OVERDUE
while doctor still called the same job healthy.
- Gate the status OVERDUE line on the same _OVERDUE_GRACE_SECONDS so
status, list, and doctor tell one consistent story.
- Reuse _next_run_overdue_seconds inside _next_run_overdue_issue,
dropping the duplicated timestamp parsing.
- Add a regression test: a next_run_at 5 minutes in the past stays a
plain 'Next run' line.
When the scheduler host dies, a job's next_run_at is stranded in the
past and `hermes cron status` still printed it as an upcoming
"Next run", hiding the outage (#114309): the report's only signal was
the gateway line, while the schedule line kept presenting a 7h-old
timestamp as future.
- _print_active_jobs_summary: when the earliest next_run_at already
passed, print a loud OVERDUE line (age + scheduler question) instead
of a bare "Next run: <past ts>"; future timestamps render unchanged.
- cron_status gateway-down branch: when ticker_heartbeat is stale,
print when the scheduler last ticked so the frozen state is
first-class visible.
Fixes#114309
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.
The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").