Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.
Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.
An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
Closes#33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.
Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.
Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.
Reactive, destination-scoped recovery instead:
- error_classifier: new `FailoverReason.role_alternation` for the vendor
alternation wordings (checked before the request-validation table since
the body also carries `invalid_request_error`); same abort+fallback hints
as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
loop's `_merge_user_content`), `destination_key` (base_url|provider,
model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
adjacent user turns merged, remember the destination on the facade for
the session so later iterations pre-merge, never touch destinations that
accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.
Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.
Fixes#112358
When a Desktop tab references a session id that exists in no profile DB
(deleted elsewhere, or a stale id after a profile rename / wiped backend),
the resume path already drops the window to a fresh draft without toasting
or hot-looping (62af32efe7, bounded by goneSessionVerdict). But the text
the user had typed into that session stayed stashed under the dead key in
`hermes:composer-drafts:v3` — not lost, yet invisible, because nothing ever
opens that key again. From the user's side the message just vanished.
The gone verdict now announces the dead key (`announceGoneSessionDraft`),
and the composer's swap onto the fresh-draft scope consumes it once
(`adoptGoneSessionDraft`): the stash moves into the new-chat bucket, the
composer is seeded with it, and an inline "Restored your unsent message"
strip with Undo appears above the input. Undo returns the text to the dead
key while it is unchanged (still recoverable by the same path); once edited
it only dismisses. Keyed on the one-shot announcement rather than on the
composer observing an id→fresh transition, because the composer remounts
across the drop (a loading route mounts no composer). A fresh draft that
already has text is never clobbered; an empty stash shows nothing.
Offer, don't hijack (apps/desktop/AGENTS.md): no navigation beyond the drop
that already happened, no focus move, no toast.
Live (headless Electron over CDP, worktree backend): type into a live
session → delete its row + restart the backend → reload at the dead id.
Before: route drops to #/, composer empty, text only in localStorage under
the dead key. After: same drop, composer holds the text, strip + Undo shown;
Undo empties the composer and puts the text back under the key; a second
open of the dead id shows no strip; a dead id with no stash shows nothing.
Fixes#111868
`display.show_reasoning` is a client-side display decision since 40f2368875,
and the Desktop transcript honors it through `$showReasoning` (#115335). The
composer path did not: `/reasoning` was marked `desktop="advanced"` in the
Python slash registry, so the Desktop refused it as "not available", while a
gateway-routed `/reasoning hide` would only have written config.yaml and left
the atom waiting for the next config refresh.
Route `/reasoning` to the gateway's `config.set key=reasoning` — the Ink TUI's
path (ui-tui/src/app/slash/commands/session.ts) — and mirror the answer:
`hide`/`show` flip `$showReasoning` on the spot, effort levels leave the gate
alone and get session scope (a throwaway slash-worker CLI could never pin the
live session's effort). No arg reports the current effort and display, like
Ink. Drops the registry's `advanced` marker and regenerates the desktop dump.
Live (real `hermes serve`, scratch HERMES_HOME, real WebSocket): config.set
key=reasoning value=hide -> {value: hide}, config.yaml show_reasoning=False;
show -> True; high -> effort only, display untouched.
Fixes#111761
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.
- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
current users). When false:
- `read_claude_code_credentials()` — the only reader of the borrowed Claude
Code login — returns None, so the resolver fallback, the expired-token
refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
file; the pool prunes a `claude_code` row an earlier adopting process
persisted.
- `_recover_codex_tokens_from_cli` returns None for both automatic recovery
paths (rejected refresh, half-empty singleton); the real AuthError is
surfaced instead. The interactive import offer in `hermes auth add
openai-codex` still asks first and is unaffected.
- One INFO line per process the first time adoption would have happened;
`hermes auth list` / `hermes auth status anthropic|openai-codex` print the
same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.
Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).
Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.
`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.
No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.
Part of #114587
Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
Widen the salvaged Image API hunk (#114340) to the whole class:
- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
"image_generation", the billing provider and the priced model id. Dict and
SDK usage objects alike; a body without tokens is a no-op, as is a call
outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
token-billed) now records too; the Image API path uses the shared helper
(the contributor's local _record_image_api_usage is folded into it) and both
pass base_url so pricing resolves the route. Task renamed image_gen ->
image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
Images API usage block under the API model (gpt-image-2), not the Hermes
quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.
Fixes#114324
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.
Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).
Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
`start()` asked "Install it now so the gateway starts on login?" with default Yes, then called
`install()` which asked two more default-Yes questions; on a redirected stdin (`< /dev/null`,
scripts, agent tool calls) all three answered themselves and a bare `start` wrote a Startup-folder
login item, spawned the gateway, and start() spawned it a second time (#113977).
- The persistence question is answered only explicitly: `HERMES_GATEWAY_INSTALL_START_ON_LOGIN`
(which start() never consulted before) or a prompt on a real TTY. Without a TTY, or under
`HERMES_NONINTERACTIVE=1`, the answer is No: the gateway starts, nothing persistent is written,
and the explicit `hermes gateway install` command is printed.
- Declining on a TTY also starts the gateway — the command is `start`; only persistence was declined.
- A Yes hands `start_now=True, start_on_login=True` to install(), which spawns and reports (including
the UAC hand-off), so start() no longer double-spawns and no longer prints the false
"Gateway install did not complete in this process" after an install that was skipped by intent.
- `install --no-start-on-login` without `--start-now` now hints `hermes gateway run` instead of the
`hermes gateway start` that used to offer the very auto-start just declined.
Supersedes #113981 (@KoNit-K), which flipped the first prompt's non-TTY default but left the
env override, the hint text and the false warning in place.
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).
- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
`served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
action env already funnel through (the dashboard site now calls it for its target).
Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
the profile is not the one the updater runs as; the module stays stdlib-only at
import time.
Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.
Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
_send_file hit the identical stale -2 after its own tokenless re-send and still
raised the generic "sendmessage error"; it now raises the shared
_session_not_ready_error. The genuine rate-limit cooldown message carries
ret/errcode/errmsg so operators can tell a real -2 frequency limit from the
session case. One parametrised test (stored token / no token) asserts the send
fails once, classify_send_error does not call it rate_limited, and the breaker
stays closed; troubleshooting row in the Weixin docs.
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
The docs line implied a manual refresh step — open chats update as soon as
the setting saves. The ungrouped `Reasoning` gate reads as dead next to the
group's, so name the path that reaches it. The on/off assertions now fail
with the element, not `expected false to be true`.
The desktop guide listed every other Settings toggle but not this one, and
the key's type only needs the two shapes config.yaml can hold (boolean, or a
quoted string), matching `display.timestamps` beside it.
Gate review on the stack:
- A runtime claim released because the adapter was gone comes back as `failed`
with `send_path_degraded`, a fresh `updated_at` and its attempt refunded.
`pending_retries` listed it as due immediately, so the redelivery worker would
claim and release it every ~2 s for as long as the adapter stayed absent, never
bounded by the attempts cap. Such rows are reconnect-driven: `pending_retries`
skips them and only the reconnect sweep re-claims them.
- With the last attempt reserved for the boot sweep only two in-process retries can
happen, so the 600 s backoff tier was unreachable and the docs promised a
"30 s, 2 min, 10 min" schedule the code never ran. Two tiers, asserted against
MAX_ATTEMPTS, and the docs say what happens.
- `_finish_ledger_delivery` no longer wakes the timer for a dead chat (blocked bot,
deleted group): the ledger would never retry it, so the wake was a wasted task and
a full pending-rows scan per dead send.
- The runtime sweep classifies a row's error only after the ownership filter, so
rows this process can never claim skip the classifier under the ledger lock.
The messaging guide described automatic redelivery only for flood-control
refusals; since the ledger now retries any rejection with a backoff, one
bullet says so and names the exception (a permanently unreachable chat).
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.
The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).
Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.
Fixes#113667
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).
- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
ConnectionError; the no-fallback re-raise after a proven session, strict
redirect headers or a non-rejection) carry the detail. Debug-logs the
endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.
Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
``_account_crashes`` trips the protocol-violation budget with
``force_trip`` after three consecutive clean exits, but the task's
``consecutive_failures`` is still 1 — so ``recompute_ready``, which
re-derives the threshold from its caller's ``failure_limit`` (default 2),
promoted the just-blocked card straight back to ``ready`` in the same
tick: ``protocol_violation -> gave_up -> promoted -> claimed`` forever
whenever ``failure_limit`` exceeds the violation count (the systemic
same-error trip at limit 1 had the same hole).
``_has_sticky_block`` now also honours the breaker's own verdict: the
newest ``gave_up`` since the last ``unblocked`` holds the card when it
recorded a violation-streak trip or ``failures >= effective_limit``.
``hermes kanban unblock`` remains the release path and still grants a
fresh budget. Tests cover the fresh-process classification (rc=0 and
rc=75), the held trip + operator unblock, and the trailer emission gate.
The dead-worker sweep learned a worker's exit status only from
``_recent_worker_exits``, which ``reap_worker_zombies`` fills via
``os.waitpid`` — so only the process that spawned the worker ever knows
how it exited. With ``kanban.dispatch_in_gateway: false`` every
``hermes kanban dispatch`` tick is a fresh process, the registry is
empty, and a worker that exited rc=0 without a terminal board call was
booked as a bare ``crashed`` / ``pid N not alive``: no
``protocol_violation`` marker, no corrective error text for the retry
worker, no violation streak — and a rc=75 quota wall was counted as a
failure instead of a neutral ``rate_limited`` requeue.
A Kanban worker (``HERMES_KANBAN_TASK`` set) now writes
``[kanban-worker-exit] rc=<code>`` as the last line of its own log on
every one-shot exit path; when the registry has no entry for a dead PID
the sweep reads that trailer and books the exit through the same
code -> kind mapping. A worker killed before its exit epilogue leaves no
trailer and stays a plain crash. ``_worker_final_output`` strips the
trailer so it never leaks into the board diagnostic.
Direction (a durable, process-independent witness in the worker log)
from PR #113638; its predicate keyed on the ``Resume this session with:``
summary, which the CLI prints before ``sys.exit`` for rc 0, 1, 75 and 130
alike and so would have booked quota walls and failed turns as protocol
violations.
Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
#114309 expects the dashboard, like `hermes cron status`, to say "scheduler last
ticked X hours ago" when a stopped ticker leaves next runs stranded in the past.
The Cron page only had the per-row `Overdue since` label and no way to date the
outage: GET /api/cron/jobs carried nothing about the ticker.
Every job the dashboard cron endpoints return now carries
`scheduler_heartbeat_age_s` — its own profile's ticker heartbeat age read inside
the same store scope as the job list (None when it cannot be dated) — and the Cron
page renders "Scheduler last ticked 7h ago — jobs that came due since then have not
fired" above the list when the oldest heartbeat among jobs expected to fire is
past the CLI's STALE_AFTER threshold (three missed 60s iterations plus slack).
Paused, disabled and completed jobs never raise the banner, matching the overdue
label's rule. The field is additive, so the response shape and every existing
consumer are unchanged.
Tests: one router test lists two profiles with one stale heartbeat file and checks
each job reports its own age (None for the profile without a heartbeat); one vitest
pins the stale-age selection and the relative label.
The hermes-bots plugin's Routines card (and its inspector's `Next run` row)
still read `Next: 7 hr. ago` for a `next_run_at` parked past the scheduler
grace — the one Desktop surface the overdue labelling left out (#114309).
The card now switches to `t.cron.overdueSince` and the inspector row to
`Overdue since`, using the same `nextRunOverdueMs` decision as the core cron
panel and sidebar, so paused/completed/disabled jobs keep the plain label.
`nextRunOverdueMs` is exported through `@hermes/plugin-sdk` (the plugin only
imports from the SDK) and its parameter is widened to the optional-field shape
the plugin's `RoutineJob` carries; no plugin i18n key is needed because the
card already renders from core `t.cron`, which every locale resolves.
Test: one card test renders an overdue vs an upcoming job and checks both the
card label and the inspector rows; the fixture's fixed 2026-08-23 next_run_at
is made clock-relative so it never ages into "overdue" on its own.
`hermes cron list`, the web dashboard Cron page and the Desktop cron panel/sidebar all
rendered a `next_run_at` parked hours in the past as an ordinary upcoming "Next run" —
the only user-visible trace of a scheduler that stopped ticking (#114309). Every
surface now labels a slot past `cron doctor`'s 15-minute grace as overdue (CLI
`Overdue:` row with the lateness, web `Overdue since`, Desktop `Overdue since` label
on the detail panel and sidebar meta) while paused/disabled/completed jobs keep the
plain label because they are not expected to fire.
The CLI's overdue check now parses through `cron.jobs._parse_aware` and `hermes_time.now`
so status/list/doctor and the ticker agree on the instant (mixed offsets, DST folds,
legacy naive stamps read as system-local like the scheduler does), and both the
ordering and the subtraction normalise to UTC: Python compares same-tzinfo datetimes by
wall clock, which is wrong across a DST fold. Status still orders the soonest run by
instant (#113874) and prints the stored stamp.
Tests: the salvaged status tests are trimmed to two invariants (overdue + stale
heartbeat is loud on status AND list; within-grace stays plain on both), the DST-fold
ordering tests freeze the CLI clock as well so their 2026-11 fixtures never start
reading as overdue, and one vitest each pins the web and Desktop helpers.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
A worker process that carries HERMES_DELEGATED_CHILD_CONTEXT next to its
HERMES_KANBAN_TASK is fenced by kanban_path_is_fenced's env-marker branch: both
auto-heartbeat writes raised PermissionError at DEBUG only, so the board showed a
worker that never beats while the process was alive (the salvaged commit already
made the return value honest, but _touch_activity ignores it). Log the refusal once
per process at WARNING with the cause, and skip the bridge for an in-process
delegate child BEFORE stamping the rate-limit window so a chatty child cannot starve
the worker's own heartbeat.
No self-write grant is added: the marker + TASK combination means "descendant, not
worker" by design (b578261584, 8a8c3634e8), and the only in-tree grant path,
kanban_db_dispatch._default_spawn, already pops the marker. The real-spawn grant
test now proves that from a dispatcher that itself carries the marker: the worker's
auto-heartbeat lands and its handoff completes.
Docs: worker guide notes the automatic claim extension and what the refusal warning
means.
`hermes gateway status` (previous commit) only reported that a registered task predates
the current XML template; nothing rewrote it, so RestartOnFailure / the logon Delay only
ever reached fresh installs (#113670). Mirror gateway.py::refresh_systemd_unit_if_needed:
`reconcile_scheduled_task` runs the same allowlist compare and, on drift, delete+creates the
task from the current template. Called from the Windows `start()` path and from the
`hermes update` launcher refresh (`_refresh_windows_gateway_launchers`). A refused
re-register (Access Denied) prints the detail and points at the elevating
`hermes gateway install`.
RestartOnFailure remaining unreachable through the non-waiting wscript launcher is the
deliberate design of 433db17c0a (#45610) and stays a documented note.
A task registered before the hardened template (433db17c0a) keeps its old XML
forever: `_install_scheduled_task` has one prompted caller and the update flow
only rewrites the launchers. `status` now fetches the registered task's XML and
compares an allowlist of template leaves — Task `version`, `<RestartOnFailure>`,
LogonTrigger `<Delay>`, Exec `<Arguments>` — and prints
⚠ Scheduled Task registration predates the current template (missing: ...)
Repair: hermes gateway install
Report only: no automatic re-registration, no launcher change, silent when
schtasks cannot be queried. The allowlist deliberately excludes <UserId>
(exported as a SID, written as DOMAIN\user) so a healthy task is never flagged.
Docs: Windows service section stating that <RestartOnFailure> only covers the
launcher because the .vbs exits immediately by design, and that auto-restart
relies on the in-process restart path.
Fixes#113670
`_cron_job_action` printed `(^_^)b Triggered job … It will run on the next
scheduler tick.` unconditionally for `run`, so a run-now the tool refused
(`execution_skipped`: paused job, claim held by another run) was
indistinguishable from an accepted one. Print the refusal, and otherwise
reuse `hermes_cli.cron._run_outcome` — the same verdict line `hermes cron run`
already prints (ran now / background / skipped).
Tests: one invariant per atom — the real one-shot CLI shape
(`_SESSION_ASYNC_DELIVERY` scoped False) reaps a dead-owner 'running' ledger
row before the sync fallback; `/cron run` on a refused claim prints the reason
and never "Triggered". Both red on origin/main. Replaces the #113938 test hunk
(it relocated the prior test's assertion). Docs: execution-history section
notes the pre-manual-run reap.
Part of #113923.
The around route reads forward from a prompt, so when a single turn holds
more display rows than the page limit, the previous prompt's page ends before
the open window's first row. Merging the two would paint one continuous
transcript with a silent hole. Each page now carries `pagination.offset`
(display rows before its first row); an older page that does not reach the
current offset replaces the display page instead of being prepended — the
reader still moves backwards, nothing is omitted without notice — and the
prepend scroll anchor is not spent for a replacement.
Docs: the desktop guide now says Show earlier keeps paging back after a
timeline jump.
Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:
- `apply_instant_title` never received the preview, so the instant title of a
paste-only opener was the generated `@file:` path — and stayed that way,
because the upgrade thread's `derive_title` fallback writes `derived`
provenance, which never replaces the `derived` title already stored.
Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
but the generated attachment ref; the preview now leads for a ref-only
opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
manual attachments stay unread).
Bots filed before sectionName existed carry only sectionId in their profile
ui_meta. The creating desktop knows the section record but never wrote the
name back, so a second desktop on the same gateway had nothing to rebuild
the section from and still drew a flat list.
backfillBotSectionNames runs beside adoptBotSectionsFromMeta in the roster
pane effect: members whose sectionId matches a locally known section but
have no sectionName are re-stamped through moveBotsToSection (one write per
profile, sequential). Once written the name is present, so it is a one-time
pass per member. Sections nobody here knows are left alone.
Docs no longer tell users to re-file or rename to stamp the name.
Section RECORDS live in the creating desktop's plugin storage
(user-sections.ts, BOT_SECTIONS_KEY) while membership (`sectionId`) rides
each bot's profile ui_meta over the gateway. A second desktop on the same
backend therefore held sectionIds whose names it had never seen, and
renderUserSections fell back to the flat list — the "no section headings"
half of #114355.
- every filing writes `sectionName` beside `sectionId` (moveBotsToSection),
so the name travels with the membership
- adoptBotSectionsFromMeta rebuilds records this desktop never created from
its members' meta, and takes a rename once every member agrees on the new
name; the roster pane runs it on each roster/meta change
- renameBotSection re-stamps the members so the new name reaches other
desktops; order and empty sections remain per desktop
- two invariants in user-sections.test.ts (red on origin/main); docs updated
The other half of the report — no "New section" entry in the + menu — is a
stale Linux build: roster-pane-toolbar.tsx and bot-row.tsx render the
entries unconditionally and nothing under plugins/hermes-bots/ reads
process.platform / navigator.platform.
/kanban is a workspace page route: the titlebar slot is hidden there and the
switcher is projected into the Kanban page's header row (WORKSPACE_PAGE_HEADER_AREA
after #114960). Point readers, the i18n comment and the test comment at that row.
Follow-up to the salvaged #114647 hunk (visible "Board" label +
`Board: <name>` accessible name):
- replace the native `title` with the app's `Tip` ("Switch board"), the
same tooltip every other dropdown trigger uses, so hover names the
ACTION instead of repeating the board name
- lead the trigger with the plugin's `project` icon so the title-bar
text reads as the Kanban board control, not a static window title
- add `switchBoard` to the kanban plugin bundle (en/ja/zh/zh-hant)
- test registers the real locale bundle and asserts the rendered label,
accessible name and tooltip (the raw-key fallback matched
`/board:/i` by coincidence)
- docs: describe the Desktop switcher and its local selection
Fixes#114642
The lifecycleKeepAlive gate in watchPreviewTileMirror had no test: swapping
preview-tile.tsx back to main left the suite green, so a regression that
dropped the flag (and turned Hide back into an unmounting Minimize for the
in-app browser) would have shipped silently. preview-tile.test.ts now drives
the real mirror and asserts a url tile registers with lifecycleKeepAlive=true
while a text file peek stays evictable.
The Hide label and kept-mounted hidden body key on lifecycleKeepAlive for every
pane, and the Terminal pane sets it too; the user guide only mentioned the
browser, so the sentence now covers both.
The salvaged kernel kept EVERY minimized zone's panes mounted, contradicting
the deliberate park-on-hide budget documented in tree-group.tsx. Only panes
that declare `lifecycleKeepAlive` (embedded Browser / HTML preview, terminal)
stay mounted while their zone is hidden; everything else unmounts exactly as
before, and a hidden zone with no keep-alive tenant renders no body at all.
- keep the reconciler's lifecycle value for visible zones; a hidden zone
reports `hot-hidden` to its kept panes instead of a hard-coded rewrite
- PaneBody rides the app's ONE shared ResizeObserver
(hooks/use-resize-observer.ts) instead of a private observer per zone
- one invariant test: keep-alive pane survives hide/restore with state and
the "Hide" label, plain sibling still parks (red on origin/main)
- docs: Hide/Restore vs Close for the in-app browser
Follow-up to the salvaged commit: instead of a new resolver component in
fallback.tsx (a second copy of MarkdownImageContent's resolve/cancel
effect), render the activity image with the existing MarkdownImage. It
already owns the local/remote file-bridge resolution (readFileDataUrl
locally, profile-scoped /api/fs/read-data-url on a remote gateway), the
loading and "Couldn't load" states, and the shared lightbox.
Rendered regression: a completed vision_analyze row with a text-only
native-vision receipt shows its disclosure, resolves the image through
the right bridge for the connection mode, and opens the lightbox.
Docs: mention expanding an image-bearing activity row.
Review finding on #114941: the CLI-runner branch of `_spawn_delivery` still
answered `status: sent` with only a `process_id`, while the live-owner branch
answers `queued`/`claimed` + `delivery_id` and the relay branch `queued`. One
tool, three vocabularies; `sent` also reads as a delivery receipt, which is the
misreading this PR exists to remove.
`_spawn_delivery` now returns `status: queued`, a `delivery_id` and the
`process_id`. For CLI-runner deliveries the id is the same sha256 of the DM
file path the runner pins when it admits to a live owner (`_dm_delivery_id`,
replacing three inline copies), so the ack and any live receipt correlate; the
relay branch passes the envelope id (also on its notification_error ack).
`sent_at` becomes `queued_at`. Tool schema, Bot Mode user guide and the
assertions in tests/tools follow; two invariant tests pin the CLI-runner and
relay ids (red before, green after).
The CLI-runner fallback of `message_agent` returns `status: sent` synchronously
while delivery runs in a background process; when that process dies (e.g. the
`hermes` entrypoint missing from PATH in the Docker image, #114628) the failure
only arrives as the background-process completion notification. The old
`detail` ("Message dispatched ... when the delivery completes, its notification
carries the reply") and the schema description ("delivery acknowledgement")
read as a delivery receipt, which misled the calling agent into reporting a
successful hand-off and routing around the tool.
Reword `detail` and the tool description to say the ack is the hand-off to a
background delivery process, not a delivery receipt, and that the completion
notification carries the outcome — the reply or the delivery failure. Update
the Bot Mode user guide to match. `status: sent` is unchanged (renaming to
`queued` + `delivery_id` is a return-contract change left to the maintainer).
Fixes#114628
Supersedes #111562
The TUI/Desktop `/rollback diff <hash>` RPC still ran `mgr.diff(host_cwd, hash)`, which
stages the HOST working tree against a host checkpoint and presents it as the container
session's diff — the same operation `/rollback diff` and `/diff session` already refuse on
the CLI and messaging gateway. Refuse it with the same `unsupported_backend_reason`, via the
RPC error (what the TUI's `.catch(guardedErr)` already renders); `rollback.list` stays.
The backend-classification block moves into `_container_checkpoint_refusal` so restore and
diff share it. The docs sentence claiming `rollback.diff` remained available is corrected.
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.
Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.
Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).
Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
Reshape of the salvaged fix from #114513 (@whyyagswhy):
- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
or round-robin order. The contributor's version called the private
``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
The File attachments section still promised that removing an attachment always deletes the on-disk file. With reference-counted blobs the file is only unlinked once no other attachment row points at it, so the user guide now states that condition and names the operator CLI path (hermes kanban attach-rm ATTACHMENT_ID).