- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
(production entry) on a real AIAgent with a one-entry chain, so dropping the
reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
rate-limited primary's declared reset is sooner than N seconds,
try_activate_fallback returns False and no cooldown is armed; docs row added.
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.
The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.
Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
before req_reasoning: {}
after req_reasoning: {'reasoning_effort': 'medium'}
agent.reasoning_effort: low -> {'reasoning_effort': 'low'} (unchanged)
macOS smart-quote substitution turns the straight quotes a user types into “ ” in the composer,
so the quoted-span mask covers both; the user-visible rule (stop words inside code, quotes or
blockquotes never hold) now has its sentence in bot-mode.md alongside the behaviour.
A child that stalls under a configured delegation.child_timeout_seconds used to
learn about the budget only by dying, losing its whole context. The liveness
wait now queues a one-line "[delegation budget warning]" through the child's
steer channel once the idle window is 80% spent (delivered at the child's next
iteration boundary), so a slow-but-recoverable child can wrap up and return its
summary. The warning fires once per idle window and re-arms when progress
resets the window; a progressing child never sees it.
Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged.
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a
plain `settled.wait(timeout=child_timeout)`, so any child that outlived the
budget was abandoned even while the provider was actively serving it.
The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could
not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint
— but the mechanism it names is exactly this gate: a child waiting on an
in-flight LLM completion, killed with a nearly-finished context. A slow child is
already bounded elsewhere (the per-call stale watchdog, the heartbeat's
staleness verdict), so this cap could only ever kill children the runtime had
judged healthy.
`child_timeout_seconds` now measures time with NO progress: the wait runs in
slices and restarts the window on the same signals the heartbeat's stale verdict
reads (completed call, tool change, activity-clock tick). A frozen child is
still abandoned when the window elapses; a progressing one is never killed for
taking long.
Timeout entries also carry `last_event_age` (how long the child had been silent),
so operators can tell a slow provider from a runaway without transcript
forensics.
Fixes the mechanism reported in #116001. The budget warning and
continuation-respawn items in that issue are separate features and are not part
of this change.
`deliver: bot-chat:<other profile>` spawned the destination's agent turn with the sending
gateway's whole environment: its `.env` settings, bridged `TERMINAL_*` policy, platform
authorization gates and provider credentials. The lane called
`strip_launch_profile_env(env)` with no target, so the strip resolved against the ambient home
override — which is the SENDER's home, never the destination's. On an ordinary root-profile
gateway (`hermes gateway run`, no `-p`) `_is_routed_home` is then false and the strip is a
complete no-op, including the #113270 gate strip that lives after its early return.
This is the only cron child built for a profile other than the one whose tick spawned it; the
worker lane (`scheduler.py`) targets its own home, so its no-target call is correct. Build this
one through `served_profile_child_env(target_home=home, inherit_credentials=True)` — the helper
`kanban_db_dispatch` and `web_server_gateway` already use for cross-profile spawns: it strips the
launch residue against the real target, scrubs credentials the launch process was given by
systemd/Compose/the shell (which no name-based strip can see), points TMPDIR at the destination's
scratch, and overlays the destination's own secrets, as a standalone `hermes -p <profile>` has.
A failure to build that environment (an unreadable target home under per-user 0700, a broken
secret source) is reported as a refusal string like every other failure in this lane rather than
raised: `_deliver_result`'s fan-out does not catch, unlike the deferred drain.
Regressions drive the real `_deliver_to_bot_chat`; removing the fix fails the two leak witnesses
(`HERMES_MODEL leaked from the launch profile`, and the firing profile's gate reaching another
profile's turn under an active override) and leaves the four guard tests green.
Fixes#117220
(cherry picked from commit 6cc81d7ddad8c9793f21fdda7c2e260c3ee1ca44)
A hand-written Linux daemon unit (systemd user unit / XDG autostart entry
running `cua-driver serve`) can be dead for days — crash loop, stopped, never
started — while `hermes computer-use doctor` and `status` report a healthy
binary: the runtime contract only checks the binary (`manifest`), and nothing
ever connected to the daemon socket (#114748).
- tools/computer_use/cua_backend.py::cua_daemon_listening — socket-level
liveness via `cua-driver status [--socket PATH]` (rc 0 = a daemon answered,
"not running" = dead, anything else = unknown). Never raises.
- tools/computer_use/doctor.py::cua_daemon_units — one scan of the units that
run cua-driver (kind, unit, exec target, runs `serve`, `--socket` path with
`%h` expanded); the pruned-Exec guard now filters that list instead of
re-scanning.
- doctor.py::_apply_daemon_liveness_guard — per `serve` unit: `pass` when its
socket answers, `fail` (degrading `ok`) when not, with the hint that a driver
reinstall does not start the daemon. Unconfigured daemons are never probed:
on Linux the MCP runtime needs none, so a silent default socket is normal.
- `hermes computer-use status` prints the dead-daemon line and exits 1.
- docs: the Linux daemon-unit paragraph under the doctor section.
Not changed: _maybe_repair_runtime_contract. The repair is gated on the
binary-level contract only; a dead daemon never enters it, so a reinstall was
never triggered by the daemon (the `.release_installed/<version>` marker the
reporter saw is written by cua-driver itself on any first run of the binary —
observed via strace of `cua-driver status` under a fresh HOME).
Supersedes #114928 (@Finn763): same finding, ~600-line implementation with a
new module and a repair-gate rewrite; this is the ~90-line version on the
existing doctor seams.
The salvaged commit added https-proxy-agent, proxy-from-env and
@types/proxy-from-env as caret ranges; the repo pins every dependency to an
exact version so `npm ci` resolves the same tree everywhere. Pins are the
versions the lockfile already resolved (7.0.6 / 2.1.0 / 1.0.4), regenerated
with `npm install --package-lock-only`.
Documents that the Desktop update check now follows HTTPS_PROXY / HTTP_PROXY /
NO_PROXY in the environment-variables reference.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The statusbar workspace menu built its reveal item unconditionally while
the sidebar menus (file-actions.tsx, review/file-tree.tsx) already hid it
on a remote backend, and `hermes:fs:reveal` returned true after
`shell.showItemInFolder`, which silently no-ops on a missing item — so a
remote bot's workspace path gave a click that did nothing and reported
success.
- electron/fs-ipc.ts: `hermes:fs:reveal` answers false when nothing at the
(tilde-expanded) path exists on this computer.
- lib/desktop-fs.ts: revealDesktopPath surfaces that false as an error the
existing revealFile toast shows (new i18n key fileMenu.revealUnavailable;
locales fall back to English through defineLocale).
- store/file-actions.ts: shouldOfferLocalReveal() — the focused row's
Connections tag decides (a row tagged with another gateway is never
local, even under a local primary); an untagged row follows the window's
primary mode, the rule the sidebar already applies.
- use-statusbar-items.tsx: the reveal item is gated on it.
Slimmer redo of #115168 by @jonpol01 (19 files): same mechanism, without
the project-menu/workspace-header rewiring and the per-locale translations.
Fixes#115167
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
`_adapter_for_subscription` fail-closed on ANY connected secondary adapter of
the pinned profile: a profile that ran Signal/Home Assistant bots but held no
Telegram token (removed on purpose to avoid a duplicate-credential collision
with the shared bot) could never receive kanban notifications in a Telegram
group that `gateway.profile_routes` pins to it — the claim rewound every tick
and the docs' route-only promise could not be met because the profile is never
route-only.
Only an adapter for the subscription's OWN platform is a credential boundary
(`_authorization_adapter` already answered for it). Adapters on other
platforms no longer gate delivery; the exact-route match still authorizes the
primary solely for a chat `profile_routes` pins to that served profile — the
same authority the primary already exercises for that chat's inbound turns.
The "other-platform adapters but none for X" warning added for this branch is
superseded by delivery; the stamped-with-the-wrong-profile warning stays.
Fixes#115460
On a `custom` main route (llama.cpp, Ollama, vLLM, ...) whose
auxiliary.title_generation is not pinned elsewhere, the turn prologue fired the
`response_format: json_schema` title request on a daemon thread at the same
instant as the turn's own streaming request, against the same self-hosted
server. A single-slot server can decode the title grammar/completion into the
main reply: the user then receives `{"title": ...}` as the assistant turn, the
main loop persists it as a genuine assistant row, replays it, and the model
adopts the format (#117296). No Hermes writer routes the aux response into the
transcript; the leaked JSON is the main completion itself.
`maybe_auto_title` now returns the upgrade thread and leaves it UNSTARTED when
`title_upgrade_must_wait_for_turn(main_runtime)`; the prologue parks it on
`agent._deferred_title_upgrade` and `finalize_turn` starts it once the model
has answered. Hosted providers keep the turn-start timing. Usage accounting
(`task='title_generation'`) and `sessions.title` are unchanged.
Every catalog card gets an 'Open in Hermes Desktop' link:
hermes://plugin/install?catalog=<name>. The app resolves the reviewed pin
itself, so the page hands it a catalog name and never a repo URL; the CLI
install command stays in the expanded card for people without the app.
In the in-app picker embed the existing '+ Add to this Agent' button is
unchanged.
A `catalog=<name>` deep link now resolves the name against the live plugin
catalog feed (the same `/docs/api/plugins.json` the Capabilities → Plugins
picker renders) and opens the Install Plugin dialog in its reviewed/pinned
catalog mode — identical to an in-app catalog pick, via one shared
`openCatalogPluginInstall` helper the Plugins tab now uses too.
The `catalog` param claims the link outright: an unknown, invalid, or
unresolvable name is a clear error toast and nothing else. It is never
reinterpreted as a git identifier, so a link cannot smuggle an unreviewed
repo behind a familiar-looking name (a `repo=` riding along is ignored).
5141b312 moved the godmode skill's self-referencing script paths from
skills/red-teaming/ to skills/security/ without re-running
website/scripts/generate-skill-docs.py, so the Docs Site check fails on
every PR that touches website/. Generator output, no hand edits.
Review follow-ups on the external-write mirror:
- Idle trigger: `openGroupChat` now runs `sweepExternalGroupWrites`, one
`session.resume` per stored member session whose thread the room still
shows, then the existing mirror. A Bot posting into its own room session
between rounds (the reporter's scenario) is posted the moment the user
opens the room, not only once the room next drives that member and it
happens to be a responder. No polling; a room mid-round is left to the
round, which sweeps its responders itself.
- Classifier: the single `[System:` skip becomes the canonical synthetic
user-row set (mirrors `agent/context_compressor.py::
_SYNTHETIC_USER_ROW_PREFIXES`, comment links both) plus the gateway's
`display_kind` on typed scaffolding rows. A compaction handoff, cron
delivery, delegation result or steer marker is never mirrored as member
speech, and the assistant row reacting to it closes the exchange instead
of inheriting the previous writer's origin.
- Stranded harvest picks the FIRST substantive assistant row after the
header-prefixed prompt and stops at the next outside user row, so a CLI
answer written after the late reply is no longer posted as the turn reply
and then mirrored again.
- Cursor edges (documented in the module header): first sight of a session
seeds the cursor at the transcript's current length — a room hydrated from
the gateway mirror (which carries no cursors) on a second Desktop does not
re-post its history; "late, never lost" holds from that moment on, at the
cost of no history replay. A cursor past the end after compaction is
reset to the end. A missing snapshot leaves the cursor alone.
Tests stay at 2 in group-external-writes.test.ts: the negative control now
also opens the room without a drive and asserts the peer exchange arrives
while the compaction/cron/auto-continue rows and their answers do not. Red
with the `openGroupChat` call removed and with the prefix set reverted.
A member's hidden per-group session is an ordinary Hermes session, so the
CLI (`hermes -p <bot> chat --resume "Group: <room> · <thread>"`), cron and
the agent's own tools append to it too. Those rows reached the transcript
but never the room log, so the room silently diverged from what the member
actually said.
What: new sibling group-external-writes.ts sweeps each member session's
unseen tail on the two paths that already read it — the pre-resume
snapshot in runGroupChatMemberTurnLeased and the stranded-marker harvest —
and appends the rows the room engine did not write itself, authored by
that member, in the session key's thread. Cursor keyed by the SESSION key
(`thread:<t>::<memberKey>`) in room.externalCursors, persisted through all
three room projections (updateGroupChat, durableGroupChatRooms, plugin.tsx
hydrate) so a window restart never re-mirrors a row.
Why the classifier is header-based: every room-fed prompt opens with the
header buildGroupChatTurnPrompt writes (now the exported
GROUP_PROMPT_HEADER_PREFIX), agent-injected `[System:` rows continue the
open exchange, and an assistant row answers whichever user row preceded it.
Why the round watermark walk: mirrored rows land after the round's submit
anchor, so `anchorIdx + 1` re-fed the member its own CLI conversation as
room news and drove an extra turn. A member's own entries are never news
to their author; the walk generalises the existing tail bump for replies.
Why the `at` stamps: the gateway mirror merge orders same-millisecond
entries by id, so a burst of mirrored rows appended in one tick came back
shuffled.
Ports the types.ts/group-chat.ts externalCursors persistence hunks of
PR #94340; its memberKey cursor and 'legacy' thread predate per-thread
member sessions and are replaced by the session-key cursor.
Co-authored-by: YusukeOshima-5564 <yusuke_oshima@capsor.co.jp>
On Discord, Telegram, Slack and Matrix a plain `/branch` used to rebind the
CURRENT chat/thread's session key to the clone, ending the original session
on that surface. The user could not keep the original path live while
exploring an alternate one — the opposite of what a branch is for.
Now the handler opens a sibling thread through the adapter's existing
`create_handoff_thread` BEFORE cloning (a failed create never orphans a
branch row), binds the thread's own session key to the clone with the
thread's routing columns written at create time, and leaves the origin key
untouched. `/branch --here` keeps the legacy in-place switch; platforms
without threads, DMs, unknown Discord parents and adapters that cannot open
a thread fall back to in-place with a one-line note. The CLI strips the
flag through the same parser so `--here` never becomes a session title.
Destination source shapes mirror each adapter's inbound key (Discord keys
threads on their own id; Telegram/Slack/Matrix on the parent chat), the
same rules the CLI->platform handoff uses.
Live repro (real gateway + real Slack adapter against a stand-in Slack
Socket Mode/Web API): base ends the origin session and rebinds its key;
fixed posts the thread seed, replies "this chat stays on it", the origin
thread keeps its session and the follow-up typed in the new thread lands on
the branch (parent_session_id = origin).
Design and first implementation by Angello Picasso (#66014, #66024);
this is a slim port onto the split slash_commands_* layout.
Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
A 42-comment r/hermesagent thread ("Frustrated.") shows the recurring
shape: the model answers "Done, I will remember that", never calls the
memory tool, and the next session knows nothing. The page had no place
that tells a user to open MEMORY.md and check, or lists the other
reasons a write can be invisible (staged approval, another profile,
memory disabled, frozen snapshot). Add an ordered checklist and say
plainly that .env variables are not memory.
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
The shipped tool-surface references still document the pre-consolidation
surface: tools-reference.md lists cronjob/todo/process/project_create/
project_list/project_switch/open_preview/close_preview/read_preview/tour/tip
(6 uncallable, 5 hidden dispatch-only aliases), and toolsets-reference.md
still claims web_search is a member of the browser toolset — membership
decacbac3 deliberately removed (#64503) with a regression test. Both rename
commits (e16ad33a9, 217ab2f8d) left website/ untouched.
Pin the pages to the live registry with a contract test (real
discover_builtin_tools()/resolve_toolset queries over the shipped .md data);
rename the rows to the registered surface (cronjob_manage, todo_list,
process_manage, desktop_project enum, desktop_preview, gui_tour, show_tip)
and add the browser row's actual members (browser_vault_*, browser_exec,
apply_layout) that the docs never mentioned.
`generate-skill-docs.py` now deletes every page under `user-guide/skills/{bundled,optional}`
it did not write this run, together with the zh-Hans mirror twin, so a skill that moves,
merges or leaves the shipped set takes its page with it instead of lingering as an orphan
that cross-links still reach (24 such pages after the shipped-set slim, plus 21 zh-Hans
copies whose English page was already gone).
The Docs Site Checks workflow regenerated the docs and never compared the result with the
committed copies, which is what GitHub renders; it now fails with a pointer to the generator
when they differ. Also regenerates the two pages that drifted since the salvaged commits.
generate-skill-docs.py writes one page per discovered skill but never prunes
pages for skills that were moved or merged. #98539 ("shipped-set slim") moved 15
skills to optional-skills/, merged the six github-* skills into one and let pdf
absorb ocr-and-documents; 25 bundled pages survived it.
They are invisible to the catalogs (regenerated from the tree) and to
check_doc_links.py, but they still render on GitHub and in the docs site, still
claim "Source | Bundled (installed by default)" for a skill that is no longer
bundled, and are still reachable through the cross-links the six github-* pages
maintain between each other.
Each page states the skill path it documents in its | Path | row, so the page can
be checked against the tree: this deletes every page whose row points at a
directory with no SKILL.md, plus the one catalog row and the one sidebar entry
that referenced the deleted merge-reconciler page.
tests/skills/test_skill_pages_match_shipped_skills.py is the guard — the next
shipped-set change that forgets its pages fails there instead of leaving them to
rot.
website/scripts/generate-skill-docs.py is the documented source of both catalogs
and of the per-skill pages, but nothing compares its output with what is
committed, so the committed copies drifted:
- optional-skills-catalog.md was missing agent-merge-conflict-arbiter and listed
pr-lens under blockchain (it lives in software-development);
- skills-catalog.md still listed merge-reconciler, which #98539 moved out of the
bundled set;
- 196 pages and both catalogs carried Windows path separators — in the Path
column and inside GitHub blob links, where a backslash is a broken URL.
This commit is the generator's output (re-running it on this branch is a no-op),
plus the orphan page #98539 left behind for merge-reconciler: no skill backs it,
its catalog row is gone, and no page links to it.
tests/skills/test_skill_docs_contract.py is the guard: the next skill that ships
without a catalog row, or a page regenerated on Windows, fails there instead of
on the published page.
A session far above the model window (~356k tokens on a 131k window in
#116472) re-ran context compression on every turn: a preflight pass that
reclaimed nothing still let the request go to the provider (400 -> overflow
handler -> another pass), and a summary stream that kept emitting tokens
while never committing held the pre-commit wait to the full 600s ceiling.
On the Desktop that blocked the gateway event loop for 10-20 minutes per
turn and the renderer was eventually killed.
- agent/turn_context.py::_fail_closed_on_insufficient_progress: when a
preflight pass makes no (or sub-5%) progress and the request provably
exceeds the model window, raise PreflightCompressionTimedOut with
"start a new session (/new)" guidance so no provider call is sent. An
unknown window or a fitting request keeps the send-as-is behaviour; a
pass that no-op'd on a transient guard (summary-failure cooldown) keeps
its typed cooldown result. Called from both insufficient-progress
branches of turn_context_compaction._run_preflight_passes.
- agent/conversation_compression.py::run_compress_context_with_progress_timeout:
an over-window request's pre-commit wait is bounded by one inactivity
budget (compression.context_timeout_seconds) instead of
context_total_ceiling_seconds; the existing first-stall deterministic
fallback then carries the compaction. Config-derived, no new knob.
Slim slice of #116592's Python half.
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
`hermes mcp login`, the dashboard re-auth (web_server_mcp.py) and the Desktop
re-auth (tui_gateway/mcp_oauth_sessions.py) each bounded the login probe at
`max(connect_timeout, 315)`. The 315 was the default 300 s callback window
plus headroom, frozen: a user who set `oauth: {timeout: 3600}` still had the
probe cancelled at 315 s. That expiry was a bare `asyncio.TimeoutError`, whose
`str()` is '', so the CLI printed a blank `✗ Authentication failed:` line
(#116278, reporter's steps 5).
`tools/mcp_oauth.py::login_connect_timeout(config)` computes the bound once —
`max(connect_timeout, oauth.timeout + 15)` — and all three call sites use it.
`_probe_single_server` re-raises its wait_for expiry as a TimeoutError naming
the server, the elapsed bound and both governing knobs, so every caller's
`humanized or exc` renders a reason.
Slimmer redo of the blank-line half of #114527 (@liuhao1024): that PR races
the connect against `asyncio.wait` so an inner exception can win; with the
window now sized from oauth.timeout the callback waiter's own
`OAuth callback timed out` message wins on its own, and the described
TimeoutError covers the remaining case in one place.
Completes the timeout atom of #116278 (closed by #116658); refs #103633
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
bot_relay.deliver has two live-owner branches: the in-process prompt.submit
handoff (#100523) and the sibling-process mailbox handoff (#113753). Both
answered the sender with a receipt sentence — "Delivered into / Queued for
@x's open Bot Chat; the reply will appear there" — so a bot on another machine
never received the target's answer whenever the target's Bot Chat happened to
be open. The owner's poller already settles a receipt carrying the reply, and
local DMs wait on it (bot_mode_dm._wait_live_dm); the relay just never read it.
Owner-first now: a live owner that advertises a mailbox — this process or a
sibling — gets the DM through it, and the handler waits on the receipt on the
local lane's budget: the reply comes back (a bare silence marker as ""), a
failed or cancelled turn as the typed 5092 refusal, and a DM still unanswered
at the budget is reported queued there with "do not resend". prompt.submit
stays as the fallback for a live session with no mailbox.
Fixes#115316
(cherry picked from commit 1f76e8f39be55ec63d7833da765ad5427a01e090)
Rebuilt onto main as one commit after #116903/#116968/#117055 landed: relay tests guard both child seams; the cross-connection docs bullets are deduplicated.
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.
Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.
Fixes#115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
`MCP: registered N tool(s) from M server(s) (2 failed)` left the failing
identity diagnosable only by elimination from the per-server `registered`
lines. The per-server WARNING fires only on the immediate-failure path; a
candidate skipped for its retry cooldown (a failure from an earlier pass) is
counted as failed with no line of its own at all.
`_connected_summary` now returns `(name, reason)` pairs and `_log_summary`
prints them inline: `(2 failed: github (Connection closed); notion (HTTP 401
...))`. The reason is the recorded `_server_connect_errors` entry (already
credential-scrubbed by `_format_connect_error`); a candidate never attempted
this pass reads `not attempted (in retry cooldown)`. One line, no second
WARNING per server on the path that already warns.
Slimmer redo of #114794 (@liuhao1024, earliest) and #114872 (@Finn763): both
name the failures via an extra WARNING per server; inline on the summary keeps
the immediate-failure path at its current two WARNINGs and still covers the
cooldown case. #114872's extra `_sanitize_error` pass is redundant with the
recorder.
Fixes#114746
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Both routes serve the slug (tools supported, 1,048,576 context, $0.37/$1.25 per M).
The Nous list is derived from OPENROUTER_MODELS so one tuple edit covers both; the docs
manifest is regenerated in the same commit.
DEFAULT_CONTEXT_LENGTHS gets its own key: substring matching would otherwise land the
slug on the glm-5.3-flash entry (1,310,720) and overstate the window by 25%.
`hermes sessions archive --older-than N` (SessionDB.archive_sessions) selected
every ENDED row matching the filters, which includes the compression ancestors
of a long conversation: they are ended (`end_reason='compression'`) and old by
construction. `set_session_archived` then flipped `archived` across the whole
lineage — including the OPEN, actively written, lease-holding live tip — and
the default `archived=exclude` listing (Desktop sidebar, `hermes sessions
list`, /resume) lost the chat while messages kept flowing.
A lineage is now matched through its tip only: `_prune_filter_where` gains
`lineage_tips_only`, which `archive_sessions` always sets and the CLI sets in
archive mode so the dry-run preview and the confirmation count show exactly
the rows the archive will touch. An idle, ended tip still archives its whole
chain (the lineage stays one unit — the listed row is the root, projected to
the tip, so sparing only the tip would leave the chat hidden anyway).
Prune is unchanged. Docs: the bulk-archive section says how compacted
conversations are matched.
Direction from #115500 by @whyyagswhy (automatic archives must not hide the
open live tip); the mechanism differs because the listing keys on the root.
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
`hermes doctor` reported "GitHub token configured (authenticated API access)" for any
value in .env, including an expired classic PAT — the very token that shadowed a working
`gh` login for every git-auth clone (#115257). The resolver half (fall through to the gh
CLI when GitHub refuses the .env token) landed separately; this adds the diagnostic half:
a "GitHub token" row under API Connectivity that sends the configured token to
api.github.com/user and, on 401, names the variable and the .env file so the user can
remove or replace it. No token configured: the row is skipped (the Skills Hub section
already reports gh-CLI / no-token state).
- request_restart() opens the same drain window as stop() (new turns refused,
in-flight work awaited), so pollers of GET /v1/runs/{id} need the
shutdown_requested_at marker from that moment as well; the marker is idempotent
so stop() re-marking after the restart wait is a no-op.
- Collapse the three adapter-level tests into one invariant that drives the real
GET /v1/runs/{id} route: live run keeps status=running but carries the marker
(durably), a status set after the boundary inherits it, a terminal run never does.
- Shutdown-path fakes gain _mark_api_runs_shutdown_requested (the real mixin method
where the fake already borrows _api_server_hook, a 0-stub where it has no adapter).
- Document the field on the runs API page.
Refs #115133.
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.
Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
The 20-minute hard cap on a visibly busy member turn truncated real work:
in a five-bot local-model room 44% of turns (mean 33 min, longest 154 min)
crossed it, the member was recorded timed-out and stranded, the next
members were handed a turn on work that did not exist yet, and the room
settled while the worker was mid-deploy (#100274).
The cap is a runaway guard, not a budget: a member that stops reporting
work already expires on the 3-minute idle timeout, so the cap only ever
cut off a member that was demonstrably producing. Raise it to 180 minutes
(no measured turn exceeded it) instead of removing it, so a session stuck
reporting `running` still cannot hold a room forever. The quiet-room
harvest window follows the constant.
Slim redo of #100288 (@astraltrekkin), which dropped the clamp entirely
behind a per-room `turnHardCapMs` knob no editor sets.
Co-authored-by: astraltrekkin <astraltrekkin@users.noreply.github.com>
A member's room turn keeps running on its own gateway when the Desktop that
submitted it quits or crashes: the gateway keeps a client-absent turn that is
still producing (session_lifecycle._ws_orphan_turn_activity_is_fresh) and only
pops its runtime once it has finished, leaving the stored row resumable. A
remote member's turn therefore always outlives the Desktop; a local member's
backend dies with it. The stranded marker that lets the next boundary harvest
a late reply was written only on the timeout path, so a turn abandoned any
other way left nothing behind: the finished reply became silent prior context
in the member's room session, and the next drive re-submitted into that
still-running session, interrupting exactly the work it should have waited for.
The marker is now written at submit, token-stamped, and cleared by the poll
that owns it on a reply, a death or an explicit stop; a timeout leaves it as
before. While that poll runs, its marker is live: a harvest that races it
leaves the turn to it. A marker with no live owner — a previous process's, or
a timeout's — is stranded exactly as before. A harvest that finds the session
genuinely gone (4007) drops the marker instead of keeping it forever, which
would have silenced the member in every later round; unreachability still
keeps it.
Fixes#115431
Disbanding a room pushed a deletion tombstone to every connected gateway
through the debounced sync job, and nothing else remembered the disband.
Two gaps let a stale mirror resurrect the room forever:
- the pending job's deletedRooms is the ONLY memory of the disband; once
the retry ladder gives up (gateway offline, push rejected) or the window
closes, every later pullGroupChatServerState() from that gateway merges
the room back because "missing remote rooms are not deletions";
- an ordinary room write inside the 350 ms debounce replaced the timer, and
the widened secondary-gateway jobs were queued from the LAST call's
(empty) deletedRooms, so the tombstone landed on the active gateway only.
Fix: keep a durable disband memory in the mirror's own tombstone shape
(room key -> revision, plugin storage 'group-chat-tombstones', hydrated
before the first pull). It rides every local publish snapshot — so any
later write re-tombstones a mirror that still projects the room — and is
applied by every pull/read-back merge. Secondary targets now inherit the
active job's coalesced changedRooms/deletedRooms. Id-keyed tombstones are
final (ids are never reused); name-keyed ones keep the revision ordering,
so a same-name recreate with a fresh roomId is never blocked.
Slim redo of #105303 by @liuhao1024 (same idea: durable disband memory
applied on read/publish/read-back), without the per-pull sweep over every
connection.
Fixes#105275
Salvages #105303
The roster the Desktop pushes to every gateway carries `connection_label`, and that
label is what a bot reads when it picks a teammate (`tools/bot_mode_probe.py` renders
"@handle on <label or id>"), what `message_agent` echoes back on a send, and how
`tools/bot_relay.py` names the machine when it refuses a target as offline.
`relayAgentsOn` read that label off the route, which carries identity only —
connectionId, mode, profile, targetProfile — so every push fell through to the raw
connection id, as the TODO this replaces said. On a live two-machine fleet the peer's
roster reads "@chii on 127-0-0-1-9119" instead of the machine's name.
Take the label from the connection registry (`host.connections()`), the only place that
has one. A Desktop build without a registry rejects that call and the ids stay, exactly
as today — the second row of the test pins that.
Follow-up to the cherry-picked #116014:
- Pin per the dependency policy (pre-1.0: `<0.(minor+2)`): httptools
`>=0.6.3,<0.9` (floor = uvicorn[standard]'s own floor), uvloop
`>=0.15.1,<0.24`. `watchfiles>=0.20,<2` already complied.
- Copy uvicorn's own uvloop marker (win32, cygwin, PyPy) plus
`sys_platform != 'android'` so `pip install '.[all]'` on those
hosts does not fail on the extra either.
- `tools/lazy_deps.py` mirrors the `web` extra for the lazy dashboard
install: it also requested `uvicorn[standard]`, so a Termux user
opening the dashboard would have hit the same uvloop build at first
use. The web_server install hint follows.
- `uv lock` regenerated; the lock delta is exactly the pyproject delta.
- Two invariant tests: no Termux-reachable extra (or core, or the lazy
dashboard feature) requests uvloop; `[all]` still does, off Android.
- Docs: troubleshooting entry in the Termux guide.
Two composer gaps, one roster path:
- #103731: `host.agents()` rows carry `profileMetadata` (title/display_name/
ui_meta) since 2ed39365d6, but mergeMultiSourceRoster dropped it, so a
remote `default` titled "CoS Bot" could only ever tag as `@hermes(-device)`.
Carry the metadata onto the remote row: the picker now offers `@cos-bot`
and the middleware resolves it to `default@<connection>`. When two rows tag
alike (two remotes both titled "CoS Bot") the bare slug names neither, so
the picker inserts `@cos-bot@<connection>` and resolveRosterMentions
accepts that form, pinning the row to one connection. botHandle() is
untouched: the local default stays the only `@hermes` in either roster
order (Map last-wins ruling).
- #94018 (renderer atom only): useRoster was the sole `host.agents()` caller
and the Bots pane its only mount, so a launch that never opened the pane
left the composer blind to other connections and the middleware's cold
fallback asked the ACTIVE gateway for profiles.list, which cannot
enumerate them. Extract fetchRosterSnapshot, add primeRoster() (one
fetchQuery into the pane's own cache key), prime on the first gateway
open and on a cold middleware submit. The message_agent grant is not
touched.
Part of #103731 (Python half: #116983)
Part of #94018
Salvages #103767 (Desktop half, slim redo) and the data.ts/plugin.tsx half of #102925.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Zeus-Deus <github.commits@widow.cc>
The relay's delivery child shared one 600s deadline between the target's turn
and the one-shot exit linger, whose own budget is the same 600s — so a turn that
answered in seconds and then handed off to a teammate was killed mid-linger,
reported to the sender as delivery_timeout (auto-retried: the turn ran twice),
its reply lost, and the handoff delivery the linger protected destroyed.
The -Q child's turn report (#113608) now carries the answer the run will print,
rewritten when a follow-up turn displaces it, and the poll loop that books a
child from that report moves next to the contract as
quiet_single_query.run_reported_turn. The cron lane keeps its policy (book after
a 2s exit grace); the relay waits for exit under the cap as before — a teammate's
reply during the linger may still become the printed answer — and at the cap
books a reported child from its latest report and leaves it to finish. Only a
turn that never ends is a timeout. The report is 0600 from creation now that it
carries the answer.
Fixes#114980
The re-offer window (2 x 600 s = 1200 s) was shorter than a legitimate in-flight
bot_relay.deliver hold (lock wait + two attempts = 1320 s) and than the
Desktop's own deliver deadline (1500 s), so a slow-but-live delivery with no
reply on disk yet was handed out AGAIN by the next drain: a double turn on the
target (or target_busy), whose fast error then won first-settled-wins over the
real answer. And every re-offer bumped the mtime and opened a new window, so an
envelope nobody ever answered was re-offered every 1200 s forever - the 6 h sweep
is mtime-based too - long after the sender's waiter had given up.
- REOFFER_AFTER_SECONDS = DESKTOP_DELIVER_TIMEOUT_SECONDS + 60: past the point
where the Desktop has provably posted its own delivery_timeout reply (the
gateway-side hold ends before it by construction), silence means a dead
Desktop. The false "never in flight that long" comment is gone.
- REPLY_WAIT_SECONDS = REOFFER_AFTER_SECONDS + DESKTOP_DELIVER_TIMEOUT_SECONDS
+ 60, so the waiter is still listening when the one re-offered delivery hits
its own deadline.
- One re-offer per envelope (`reoffered_at` stamped on the claimed file); once
created_at + REPLY_WAIT_SECONDS passes unanswered the drain writes a
delivery_timeout reply instead, so the sender learns and no turn loop runs
against a target nobody is waiting for. Age is created_at, not bumped mtime.
- bot_mode.envelope_ttl_seconds applies to the re-offer leg exactly as to the
outbox: the message is back in the queue from claim + REOFFER_AFTER_SECONDS,
and a drain that comes a whole TTL later refuses it with queued_expired.
- write_reply's first-settled-wins is now documented as safe BECAUSE two
deliveries of one envelope can no longer overlap.
The existing constants test pins REOFFER > Desktop deadline > live hold and
REPLY_WAIT > REOFFER + Desktop deadline; the drain-side test drives the real
outbox.drain handler through re-offer, no second re-offer, and the timeout
reply. Docs bullet reworded to the new window.
After outbox.drain moved an envelope to claimed/, a Desktop that disconnected
before bot_relay.deliver left it there with no reply: the sender's waiter
learned nothing until its deadline, every later drain saw an empty outbox, and
the 6h sweep deleted the message. claim_pending_envelopes now re-offers claimed
envelopes unanswered for REOFFER_AFTER_SECONDS (two turn attempts — a live
delivery never runs that long without the Desktop posting its own timeout
reply), bumping the mtime so each re-offer opens a new window; the claim itself
now stamps the mtime so the window counts from the claim, not the enqueue.
write_reply is idempotent by envelope id: the first settled reply is kept, so a
re-offered delivery's second outcome never displaces the answer the waiter read.
Slim redo of #111207 (@JoaoMarcos44): the Desktop already drains on every
reconnect (b469be8cc3, 1eb771e2ff), so no Desktop change and no per-envelope
receipt store are needed for the loss the PR reproduced. Closes the residual of
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
POST /api/sessions/{id}/chat/stream is the SSE sibling of the chat route and
went _prepare_session_chat -> _run_agent with no owner check, so it stayed a
second writer into a live-owned canonical Bot Chat (#114959, "Remaining
siblings"). It now admits through the same _admit_to_live_bot_chat door; the
owner's settled receipt is streamed as the run's single assistant.completed
event, a receipt still open at the budget as run.queued (the 202 shape), a
failed one as an error event carrying the reason. The receipt wait is shared
with the JSON route (_await_live_bot_chat_receipt) and sends SSE keepalives
while it waits.
Docs: the peer dm paragraph now carries both the read-timeout wording
(#116885) and the open-chat wording in one paragraph so either landing order
resolves to this text.
`hermes peer run` posts POST /v1/runs with the peer's canonical Bot Chat as
session_id. Like the /chat transport before #114959, the run executed here
while a Desktop session held that chat's lease — a second writer the open chat
never showed, with the two transcripts interleaved in state.db.
The admission that /api/sessions/{id}/chat now performs moves onto the adapter
as one helper both peer transports call, so the two lanes cannot drift. When
the selected session is the live-held canonical Bot Chat, /v1/runs admits the
message to the owner's mailbox and drives the run from the owner's receipt
instead of an executor: `settled` completes it with the reply, a failed
receipt fails it with the owner's classified reason, and the run retires the
way an executor-backed one does. `peer run` keeps its run_id and `peer status`
keeps working; the status carries the delivery_id.
/stop cannot reach the owner's turn — the mailbox has no recall once a record
is claimed — so a stop ends this run as cancelled while the chat finishes on
its own; the stop handler already reports a run without an in-process agent as
not interruptible here.