Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:
- `apply_instant_title` never received the preview, so the instant title of a
paste-only opener was the generated `@file:` path — and stayed that way,
because the upgrade thread's `derive_title` fallback writes `derived`
provenance, which never replaces the `derived` title already stored.
Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
but the generated attachment ref; the preview now leads for a ref-only
opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
manual attachments stay unread).
Bots filed before sectionName existed carry only sectionId in their profile
ui_meta. The creating desktop knows the section record but never wrote the
name back, so a second desktop on the same gateway had nothing to rebuild
the section from and still drew a flat list.
backfillBotSectionNames runs beside adoptBotSectionsFromMeta in the roster
pane effect: members whose sectionId matches a locally known section but
have no sectionName are re-stamped through moveBotsToSection (one write per
profile, sequential). Once written the name is present, so it is a one-time
pass per member. Sections nobody here knows are left alone.
Docs no longer tell users to re-file or rename to stamp the name.
Section RECORDS live in the creating desktop's plugin storage
(user-sections.ts, BOT_SECTIONS_KEY) while membership (`sectionId`) rides
each bot's profile ui_meta over the gateway. A second desktop on the same
backend therefore held sectionIds whose names it had never seen, and
renderUserSections fell back to the flat list — the "no section headings"
half of #114355.
- every filing writes `sectionName` beside `sectionId` (moveBotsToSection),
so the name travels with the membership
- adoptBotSectionsFromMeta rebuilds records this desktop never created from
its members' meta, and takes a rename once every member agrees on the new
name; the roster pane runs it on each roster/meta change
- renameBotSection re-stamps the members so the new name reaches other
desktops; order and empty sections remain per desktop
- two invariants in user-sections.test.ts (red on origin/main); docs updated
The other half of the report — no "New section" entry in the + menu — is a
stale Linux build: roster-pane-toolbar.tsx and bot-row.tsx render the
entries unconditionally and nothing under plugins/hermes-bots/ reads
process.platform / navigator.platform.
The honest "unavailable (remote backend)" state (previous commit) tells the
user the copy will never happen; the tooltip now also says what does work
against a remote backend — Install from Git with the Desktop target checked
clones the desktop half onto this machine (the install modal already takes
that branch for connection.mode === 'remote'). Docs note the new state next to
the existing remote-backend paragraph.
/kanban is a workspace page route: the titlebar slot is hidden there and the
switcher is projected into the Kanban page's header row (WORKSPACE_PAGE_HEADER_AREA
after #114960). Point readers, the i18n comment and the test comment at that row.
Follow-up to the salvaged #114647 hunk (visible "Board" label +
`Board: <name>` accessible name):
- replace the native `title` with the app's `Tip` ("Switch board"), the
same tooltip every other dropdown trigger uses, so hover names the
ACTION instead of repeating the board name
- lead the trigger with the plugin's `project` icon so the title-bar
text reads as the Kanban board control, not a static window title
- add `switchBoard` to the kanban plugin bundle (en/ja/zh/zh-hant)
- test registers the real locale bundle and asserts the rendered label,
accessible name and tooltip (the raw-key fallback matched
`/board:/i` by coincidence)
- docs: describe the Desktop switcher and its local selection
Fixes#114642
The contributor test covers the queue fall-through; the Desktop's normal case is
a redirect-capable AIAgent, where busy_input_mode=interrupt turned the edit into
a mid-turn redirect and left the un-edited transcript in place. Pin that branch
too, and document in the rewind section that a truncating submit refuses with
4009 while a turn runs so hosts interrupt + retry (the Desktop already does).
The lifecycleKeepAlive gate in watchPreviewTileMirror had no test: swapping
preview-tile.tsx back to main left the suite green, so a regression that
dropped the flag (and turned Hide back into an unmounting Minimize for the
in-app browser) would have shipped silently. preview-tile.test.ts now drives
the real mirror and asserts a url tile registers with lifecycleKeepAlive=true
while a text file peek stays evictable.
The Hide label and kept-mounted hidden body key on lifecycleKeepAlive for every
pane, and the Terminal pane sets it too; the user guide only mentioned the
browser, so the sentence now covers both.
The salvaged kernel kept EVERY minimized zone's panes mounted, contradicting
the deliberate park-on-hide budget documented in tree-group.tsx. Only panes
that declare `lifecycleKeepAlive` (embedded Browser / HTML preview, terminal)
stay mounted while their zone is hidden; everything else unmounts exactly as
before, and a hidden zone with no keep-alive tenant renders no body at all.
- keep the reconciler's lifecycle value for visible zones; a hidden zone
reports `hot-hidden` to its kept panes instead of a hard-coded rewrite
- PaneBody rides the app's ONE shared ResizeObserver
(hooks/use-resize-observer.ts) instead of a private observer per zone
- one invariant test: keep-alive pane survives hide/restore with state and
the "Hide" label, plain sibling still parks (red on origin/main)
- docs: Hide/Restore vs Close for the in-app browser
Follow-up to the salvaged commit: instead of a new resolver component in
fallback.tsx (a second copy of MarkdownImageContent's resolve/cancel
effect), render the activity image with the existing MarkdownImage. It
already owns the local/remote file-bridge resolution (readFileDataUrl
locally, profile-scoped /api/fs/read-data-url on a remote gateway), the
loading and "Couldn't load" states, and the shared lightbox.
Rendered regression: a completed vision_analyze row with a text-only
native-vision receipt shows its disclosure, resolves the image through
the right bridge for the connection mode, and opens the lightbox.
Docs: mention expanding an image-bearing activity row.
`444c75c10a` rendered `titleBar.center` in the workspace panel header on
full pages so the kanban board switcher would sit beside the page title
instead of colliding with the sidebar tab strip. That made the area migrate
between two React subtrees on chat <-> page navigation: the old instance's
effect cleanup ran after the new instance's setup and wiped every global
side effect a third-party plugin had just re-created (#114290).
Keep both intents: `titleBar.center` is a permanent titlebar slot again (the
salvaged commit), and page-owned controls move to a dedicated
`WORKSPACE_PAGE_HEADER_AREA` (`workspace.pageHeader`) that the workspace
pane projects into its vetoed tab row while `$workspaceIsPage` holds —
exactly the placement `444c75c10a` introduced, just not through the
plugin-facing titlebar area. Kanban's board switcher contributes there.
- SDK exports `WORKSPACE_PAGE_HEADER_AREA`; `TITLEBAR_AREAS` documents the
permanent-mount contract.
- Test: a `titleBar.center` component's effect runs setup once and cleanup
never across a chat -> /skills -> chat round trip (red on base).
- Docs: plugin SDK guide covers the lifecycle contract and the page-header
area.
Review finding on #114941: the CLI-runner branch of `_spawn_delivery` still
answered `status: sent` with only a `process_id`, while the live-owner branch
answers `queued`/`claimed` + `delivery_id` and the relay branch `queued`. One
tool, three vocabularies; `sent` also reads as a delivery receipt, which is the
misreading this PR exists to remove.
`_spawn_delivery` now returns `status: queued`, a `delivery_id` and the
`process_id`. For CLI-runner deliveries the id is the same sha256 of the DM
file path the runner pins when it admits to a live owner (`_dm_delivery_id`,
replacing three inline copies), so the ack and any live receipt correlate; the
relay branch passes the envelope id (also on its notification_error ack).
`sent_at` becomes `queued_at`. Tool schema, Bot Mode user guide and the
assertions in tests/tools follow; two invariant tests pin the CLI-runner and
relay ids (red before, green after).
The CLI-runner fallback of `message_agent` returns `status: sent` synchronously
while delivery runs in a background process; when that process dies (e.g. the
`hermes` entrypoint missing from PATH in the Docker image, #114628) the failure
only arrives as the background-process completion notification. The old
`detail` ("Message dispatched ... when the delivery completes, its notification
carries the reply") and the schema description ("delivery acknowledgement")
read as a delivery receipt, which misled the calling agent into reporting a
successful hand-off and routing around the tool.
Reword `detail` and the tool description to say the ack is the hand-off to a
background delivery process, not a delivery receipt, and that the completion
notification carries the outcome — the reply or the delivery failure. Update
the Bot Mode user guide to match. `status: sent` is unchanged (renaming to
`queued` + `delivery_id` is a return-contract change left to the maintainer).
Fixes#114628
Supersedes #111562
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.
Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.
Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.
Part of #114271
The TUI/Desktop `/rollback diff <hash>` RPC still ran `mgr.diff(host_cwd, hash)`, which
stages the HOST working tree against a host checkpoint and presents it as the container
session's diff — the same operation `/rollback diff` and `/diff session` already refuse on
the CLI and messaging gateway. Refuse it with the same `unsupported_backend_reason`, via the
RPC error (what the TUI's `.catch(guardedErr)` already renders); `rollback.list` stays.
The backend-classification block moves into `_container_checkpoint_refusal` so restore and
diff share it. The docs sentence claiming `rollback.diff` remained available is corrected.
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.
Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.
Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).
Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
Reshape of the salvaged fix from #114513 (@whyyagswhy):
- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
or round-robin order. The contributor's version called the private
``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
The File attachments section still promised that removing an attachment always deletes the on-disk file. With reference-counted blobs the file is only unlinked once no other attachment row points at it, so the user guide now states that condition and names the operator CLI path (hermes kanban attach-rm ATTACHMENT_ID).
#113683 changed user-visible behaviour: an owner whose liveness cannot be proved
is now counted as live (MAX_CONCURRENT_SESSIONS) instead of refusing every claim
(SESSION_COORDINATION_UNAVAILABLE), and it only fences its own session id. Note
that in the existing max_concurrent_sessions section.
sessions.md lists the state-owned sources the auto-prune sweep closes; add
`recovered`. The second test drives `_reconstruct_missing_sessions` itself so
the placeholder shape the sweep must age out is the one recovery writes, and
keeps a fresh placeholder open as the control.
Fixes#114730
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).
Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:
- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
{pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
Report present -> the delivery is booked from it (real exit code and
stream tails when the child exits within a short grace) and the still-
lingering child is left running, drained and reaped by a daemon thread.
No report by the cap -> the turn never ended: killed and booked as a
timeout, exactly as before. The linger itself is untouched.
Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.
Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.
Fixes#113608
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:
- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
which logs the single final warning AND drops the supervisor from
``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
(``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
attached" instead of a dead entry, and the next browser call for the task
starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
(two invariants: bounded + unregistered after attach; first-dial failure
still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
main — that file is the opt-in real-Chrome E2E suite and its module-level
skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
reconnect.
Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.
Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.
Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).
guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
The Portal "Profile setup" section promised that every profile picks up a
Portal login automatically. The shared token store only merges fresher
tokens into a profile that already has a Nous state; a profile with no
providers.nous entry of its own fails closed at boot and is asked to run
the setup flow (hermes_cli/auth.py:1545, auth_nous.py:1056), because
profiles are independent islands (#111724). Correct the promise to
"sign in once per profile" and describe the fail-closed behavior, in both
the English and zh-Hans copies.
Fixes#114259
`_runtime_credentials_ready()` collapsed "nothing is configured" and "the
configured credential is unusable right now" into one False, so a profile
whose only Nous credential was benched by a failed refresh (or quarantined
DEAD) was told "No inference provider is configured yet" and offered the
provider picker — re-running setup on an OAuth provider with single-use
refresh tokens can rotate the grant away from the session that was working
(#113720).
The readiness probe now returns the resolver's exception too
(`_probe_runtime_credentials`). At CLI startup only the resolver's
`no_provider_configured` verdict reaches the wizard; any other failure is
explained by `_explain_unusable_credentials`: `format_auth_error(exc)`
plus, from the provider's pool, "cooling down ... re-enters rotation in
about Nm" or "sign-in was lost (<reason>); run `hermes auth add <provider>`".
An actually-empty profile still gets the wizard.
Fixes#113720
Supersedes #113732 (@whyyagswhy): its `_inventory_other_providers()` gate
returns False for a Nous-configured profile, so the reported Nous case would
still have reached the wizard, and it printed no reason.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.
save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.
security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
tools.skills_hub already owns _guarded_http_get (SSRF pre-check, guarded
client, bounded manual redirects that fail closed on a missing Location,
website-policy check). Route both the skills.sh sitemap index and each
<loc> sitemap through it instead of carrying a second inline guarded
client in the skills.sh adapter. The explicit Accept-Encoding: gzip
header goes away — httpx sends it by default.
Docs: the SSRF section now names the remote-party fetches the guard
covers and the LAN-provider opt-in.
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).
#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.
- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
(RuntimeError subclass) so the dispatcher can tell a host refusal from a
card failure.
- _record_task_failure(infrastructure=True): run + event carry
`infrastructure: true`, consecutive_failures is left alone, the breaker
never trips; the card stays `ready` with the linger remedy in
last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
is such a refusal inside the rate-limit cooldown window, so a dead bus is
retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
managed gateway.
#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.
- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
is scope-wrapped when the user bus is reachable. Without a bus it degrades
with a loud once-per-process warning naming the consequence and both
remedies (enable-linger, KillMode=process) rather than refusing — the
unit's KillMode/lifetime is unknowable and a long-lived sequencer without
linger must keep working.
Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.
Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:
- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
between auxiliary_client._is_unsupported_parameter_error and the new
classifier helper); is_reasoning_field_rejection() replaces
is_reasoning_disable_rejected() with the same token gate and a
symmetric "unsupported" window so both word orders match
("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
model "requires reasoning" (a chat-only relay does not); a second
reasoning-field rejection in the same turn is treated as spent and
takes the fallback chain instead of replaying the identical request
max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
docs mention the reversed wording and the main-loop recovery.
Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
Follow-up on the picked #113768 commit so it clears the constraint that got
e9a54c48f2 reverted (a9fabe43c4): `/model <direct-alias>` resolves the alias
LABEL first and applied the alias endpoint only afterwards, so any guard that
fires on the label alone turns a working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery went red again
with the PR as-is).
- hermes_cli/runtime_provider.py::_raise_if_local_alias_missing_endpoint keys
on the class (anything auth.resolve_provider maps to `custom` without a rung
of its own; llamacpp keeps its managed-server fail-fast) instead of a second
hardcoded alias set, requires the providers.<alias> block to actually carry a
base_url (an entry without one falls through to OpenRouter too), and names
the alias plus where to set its endpoint. OPENROUTER_BASE_URL never counts as
the alias endpoint; an explicit api_key does not lift the guard.
- hermes_cli/model_switch.py::_creds_for_switched_provider hands a URL-bearing
direct alias's base_url to the resolver as explicit_base_url, so the alias
has an endpoint at the moment the guard runs (the same URL
_apply_direct_alias_endpoint installs later).
- Tests trimmed to two invariants (raise-and-name incl. the OPENROUTER_BASE_URL
/ explicit-api_key non-lifts + bare-custom control; every endpoint source
resolves to its URL). Docs: troubleshooting entry in the local Ollama guide.
Cron replay snapshots that store the resolved `custom` instead of the alias
are #109765's atom and untouched here.
With a bound project the settings dialog omits a blank directory from the PATCH and rename_board mirrors the project's primary folder, so clearing the field cannot fall back to scratch until the project is unbound; the Settings bullet now says so. The bundle test drops five wording-level substring asserts that pinned the payload shape already covered behaviourally by test_kanban_board_project_api.py, keeping only the switcher's unbind seam.
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.
Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.
Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.
- list_profiles(lazy_skill_count=True): skill_count is the last known value;
a missing/aged entry schedules ONE background _count_skills per profile per
60 s recheck window, so the request thread does zero skill-tree I/O and the
refresh cadence is decoupled from the poll rate. The two polled callers and
the REST fallback entry use it; the sync default (CLI, detail views) is
unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
os.walk with excluded/support dirs pruned) instead of rglob — half the
fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
skill add/remove immediately on the next refresh).
Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
profiles.list 2189 + 2015, projects/tree 2177 + 2015;
a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
after: GET /api/profiles 0 skill-tree stats/scandir on the request thread
(counts land from the background refresh by the next poll),
profiles.list 0, projects/tree 0; _count_skills (detail/control)
still reports 200 with 203 scandir; the vanished-dir case returns
199 and list_profiles() enumerates all 5 profiles.
Fixes#114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
A Nous pool entry whose refresh reached the resolver's state-shape raises
("Hermes is not logged into Nous Portal.", "No access token found ...",
"No refresh token is available ...") carried relogin_required=True but no
OAuth code, so `OAuthProviderFlow.is_terminal_refresh_error` classified it
transient: `_recover_failed_refresh` benched the only credential for an
hour with null last_error_* fields and nothing at WARNING (#113718).
Give those raises codes (`nous_auth_missing`, `nous_auth_missing_access_token`,
`nous_auth_missing_refresh_token`) and list them in the Nous flow's
terminal_refresh_codes, mirroring the Codex/xAI `*_auth_missing_refresh_token`
convention. The pool then quarantines, marks the row DEAD with the reason,
and logs the existing WARNING naming `hermes auth add nous`; every consumer
of the classifier (pool, singleton `_refresh_nous_or_quarantine`) agrees.
A failure that says nothing about the login (network error, non-JSON 5xx
body) still benches as before.
Fixes#113718
Salvages #113726 (@whyyagswhy) — tests kept, consumer-side hunk replaced
by the classifier fix.
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.
- send() now reports a mid-split refusal through the existing partial_overflow
contract (the key _edit_overflow_split already sets and the stream consumer
reads): delivered_chunks / total_chunks / last_message_id, plus
undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
(flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
_resume_partial_send hook (default None = keep the partial failure, never
re-send the head; the plain-text fallback is skipped for partials too). The
Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
only around the API calls — never across the reconnect wait — on send() and the
media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
capped 300s); sends inside the window fail closed locally with the same
flood_control:<s> result and no API call, so ledger recognition and redelivery
timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
"waiting Ns" followed by no wait.
Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.
Fixes#114396
Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
- hermes_cli/config_defaults.py: the agent.clarify_timeout comment claimed
"CLI clarify blocks indefinitely and ignores this"; the CLI reads it (via
resolve_clarify_timeout) — describe the real contract and the legacy
clarify.timeout precedence once.
- agent/tool_executor.py: the sequential-middleware docstring named clarify
as exempt from the generic deadline while _SEQUENTIAL_DEADLINE_EXEMPT_TOOLS
does not list it; the exemption is the _NEVER_PARALLEL_TOOLS inline path
(61645cde82). Point at it so nobody re-adds a dead set entry.
- website/docs: configuration.md Clarify section covers all surfaces and the
deadline exemption; discord.md (+ zh-Hans discord/telegram) still said the
default was 600 s.
Two follow-ups on the salvaged commits:
- `_ensure_default_soul_md`: widen the cyclic-only (ELOOP) branch to any symlink the seed
cannot write through — a link dangling into a missing directory raises ENOENT on the same
write and bricked home init the same way (HomeInitializationError -> exit 75 -> supervisor
relaunch storm). Shape: write through the link first (a working link stays operator
wiring), and only on OSError seed the default in place of the link via mkstemp + replace.
Tests trimmed to two invariants (cyclic/dangling replaced; resolving link preserved).
- `_run_pre_update_backup`: the quick snapshot stays best-effort by design (8ed599dc05:
"a broken backup never blocks the update"), but a failure was swallowed at DEBUG level, so
the user only learned from the receipt afterwards that no recovery point existed. Print a
stdout warning with the reason and continue; together with the receipt skip/fail split the
receipt now says "failed" only when a snapshot was requested and not produced.
- docs: updating.md describes the warning and the receipt split.
User-visible behaviour changed by the post `files` fix: every top-level attachment is
collected (folders skipped) with an [Attachment: name] marker, and small .txt/.md content
is appended after the caption instead of replacing it.
`get_codex_auth_status` / `get_xai_oauth_auth_status` back every credential-gated
listing (`/model` picker rows, `hermes doctor`, dashboard auth cards). The shared
`_pool_first_oauth_status` read the pool with `select()`, which is a runtime lease:
it refreshes an expiring single-use token and, when that speculative POST fails
transiently (503, endpoint unreachable), benches the entry with a persisted
exhaustion cooldown. The picker then rendered the provider as an unconfigured
skeleton ("needs setup" / "0 models") while the runtime resolver kept serving
the same credential. On a round-robin pool the read also rotated and persisted
the priority order.
Read the pool with `peek()`: no refresh, no accounting, no rotation, no persist.
Refreshing stays with the runtime resolver reached through the `resolve`
fallback, whose failures persist nothing (same design as the Nous status).
Slimmer redo of #114379 by @Finn763 (observe flag threaded through four pool
methods plus non-refreshing resolvers); fixtures adapted from that PR.
Live: pool-only openai-codex entry, expired access token, stubbed token endpoint
returning 503 — before: 1 refresh POST, entry persisted `exhausted`,
`has_available()` False, picker row gone; after: 0 POSTs, entry untouched,
`has_available()` True, picker row present with the catalog. Runtime `select()`
still refreshes (control).
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Two invariant tests: the handler returns 0/1/3 for connected / connection
failed / not in config, and the `hermes mcp` CLI dispatcher forwards that code
to `main()` (which already exits on an int return). The MCP guide documents the
codes so probes and watchdogs can stop parsing the output.
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.