Commit Graph

3065 Commits

Author SHA1 Message Date
teknium1
34ba61bf67 fix(agent): paste title hint reaches the instant title and the prompt.submit contract
Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:

- `apply_instant_title` never received the preview, so the instant title of a
  paste-only opener was the generated `@file:` path — and stayed that way,
  because the upgrade thread's `derive_title` fallback writes `derived`
  provenance, which never replaces the `derived` title already stored.
  Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
  but the generated attachment ref; the preview now leads for a ref-only
  opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
  TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
  manual attachments stay unread).
2026-09-18 10:56:00 -07:00
teknium1
996c18274b fix(desktop): the creating desktop backfills section names onto bots filed before names rode along
Bots filed before sectionName existed carry only sectionId in their profile
ui_meta. The creating desktop knows the section record but never wrote the
name back, so a second desktop on the same gateway had nothing to rebuild
the section from and still drew a flat list.

backfillBotSectionNames runs beside adoptBotSectionsFromMeta in the roster
pane effect: members whose sectionId matches a locally known section but
have no sectionName are re-stamped through moveBotsToSection (one write per
profile, sequential). Once written the name is present, so it is a one-time
pass per member. Sections nobody here knows are left alone.

Docs no longer tell users to re-file or rename to stamp the name.
2026-09-18 10:55:24 -07:00
teknium1
1197afc2b9 fix(desktop): bot roster sections follow the roster to every desktop
Section RECORDS live in the creating desktop's plugin storage
(user-sections.ts, BOT_SECTIONS_KEY) while membership (`sectionId`) rides
each bot's profile ui_meta over the gateway. A second desktop on the same
backend therefore held sectionIds whose names it had never seen, and
renderUserSections fell back to the flat list — the "no section headings"
half of #114355.

- every filing writes `sectionName` beside `sectionId` (moveBotsToSection),
  so the name travels with the membership
- adoptBotSectionsFromMeta rebuilds records this desktop never created from
  its members' meta, and takes a rename once every member agrees on the new
  name; the roster pane runs it on each roster/meta change
- renameBotSection re-stamps the members so the new name reaches other
  desktops; order and empty sections remain per desktop
- two invariants in user-sections.test.ts (red on origin/main); docs updated

The other half of the report — no "New section" entry in the + menu — is a
stale Linux build: roster-pane-toolbar.tsx and bot-row.tsx render the
entries unconditionally and nothing under plugins/hermes-bots/ reads
process.platform / navigator.platform.
2026-09-18 10:55:24 -07:00
teknium1
aee7de4db5 fix(desktop): point the remote-backend desktop-half tooltip at the install path
The honest "unavailable (remote backend)" state (previous commit) tells the
user the copy will never happen; the tooltip now also says what does work
against a remote backend — Install from Git with the Desktop target checked
clones the desktop half onto this machine (the install modal already takes
that branch for connection.mode === 'remote'). Docs note the new state next to
the existing remote-backend paragraph.
2026-09-18 10:53:41 -07:00
teknium1
e052500722 docs(desktop): the Kanban board switcher sits in the page header, not the window title bar
/kanban is a workspace page route: the titlebar slot is hidden there and the
switcher is projected into the Kanban page's header row (WORKSPACE_PAGE_HEADER_AREA
after #114960). Point readers, the i18n comment and the test comment at that row.
2026-09-18 10:53:00 -07:00
teknium1
d235ae0991 fix(desktop): Kanban board switcher reads as a control — icon, tooltip, i18n
Follow-up to the salvaged #114647 hunk (visible "Board" label +
`Board: <name>` accessible name):

- replace the native `title` with the app's `Tip` ("Switch board"), the
  same tooltip every other dropdown trigger uses, so hover names the
  ACTION instead of repeating the board name
- lead the trigger with the plugin's `project` icon so the title-bar
  text reads as the Kanban board control, not a static window title
- add `switchBoard` to the kanban plugin bundle (en/ja/zh/zh-hant)
- test registers the real locale bundle and asserts the rendered label,
  accessible name and tooltip (the raw-key fallback matched
  `/board:/i` by coincidence)
- docs: describe the Desktop switcher and its local selection

Fixes #114642
2026-09-18 10:53:00 -07:00
teknium1
66ba4e5114 test(tui_gateway): pin that a live-turn rewind is not redirected; document the 4009 contract
The contributor test covers the queue fall-through; the Desktop's normal case is
a redirect-capable AIAgent, where busy_input_mode=interrupt turned the edit into
a mid-turn redirect and left the un-edited transcript in place. Pin that branch
too, and document in the rewind section that a truncating submit refuses with
4009 while a turn runs so hosts interrupt + retry (the Desktop already does).
2026-09-18 10:52:27 -07:00
teknium1
26a7a95ceb fix: pin Browser tile keep-alive through the mirror; document Hide for Terminal
The lifecycleKeepAlive gate in watchPreviewTileMirror had no test: swapping
preview-tile.tsx back to main left the suite green, so a regression that
dropped the flag (and turned Hide back into an unmounting Minimize for the
in-app browser) would have shipped silently. preview-tile.test.ts now drives
the real mirror and asserts a url tile registers with lifecycleKeepAlive=true
while a text file peek stays evictable.

The Hide label and kept-mounted hidden body key on lifecycleKeepAlive for every
pane, and the Terminal pane sets it too; the user guide only mentioned the
browser, so the sentence now covers both.
2026-09-18 10:47:00 -07:00
teknium1
a0c61500bc fix(desktop): gate the kept-mounted hidden body on lifecycleKeepAlive
The salvaged kernel kept EVERY minimized zone's panes mounted, contradicting
the deliberate park-on-hide budget documented in tree-group.tsx. Only panes
that declare `lifecycleKeepAlive` (embedded Browser / HTML preview, terminal)
stay mounted while their zone is hidden; everything else unmounts exactly as
before, and a hidden zone with no keep-alive tenant renders no body at all.

- keep the reconciler's lifecycle value for visible zones; a hidden zone
  reports `hot-hidden` to its kept panes instead of a hard-coded rewrite
- PaneBody rides the app's ONE shared ResizeObserver
  (hooks/use-resize-observer.ts) instead of a private observer per zone
- one invariant test: keep-alive pane survives hide/restore with state and
  the "Hide" label, plain sibling still parks (red on origin/main)
- docs: Hide/Restore vs Close for the in-app browser
2026-09-18 10:47:00 -07:00
teknium1
9e4b10da41 fix(desktop): resolve tool activity images through the media pipeline
Follow-up to the salvaged commit: instead of a new resolver component in
fallback.tsx (a second copy of MarkdownImageContent's resolve/cancel
effect), render the activity image with the existing MarkdownImage. It
already owns the local/remote file-bridge resolution (readFileDataUrl
locally, profile-scoped /api/fs/read-data-url on a remote gateway), the
loading and "Couldn't load" states, and the shared lightbox.

Rendered regression: a completed vision_analyze row with a text-only
native-vision receipt shows its disclosure, resolves the image through
the right bridge for the connection mode, and opens the lightbox.

Docs: mention expanding an image-bearing activity row.
2026-09-18 10:46:22 -07:00
teknium1
14d862384b fix(desktop): page-owned header controls get their own area; titleBar.center stays permanent
`444c75c10a` rendered `titleBar.center` in the workspace panel header on
full pages so the kanban board switcher would sit beside the page title
instead of colliding with the sidebar tab strip. That made the area migrate
between two React subtrees on chat <-> page navigation: the old instance's
effect cleanup ran after the new instance's setup and wiped every global
side effect a third-party plugin had just re-created (#114290).

Keep both intents: `titleBar.center` is a permanent titlebar slot again (the
salvaged commit), and page-owned controls move to a dedicated
`WORKSPACE_PAGE_HEADER_AREA` (`workspace.pageHeader`) that the workspace
pane projects into its vetoed tab row while `$workspaceIsPage` holds —
exactly the placement `444c75c10a` introduced, just not through the
plugin-facing titlebar area. Kanban's board switcher contributes there.

- SDK exports `WORKSPACE_PAGE_HEADER_AREA`; `TITLEBAR_AREAS` documents the
  permanent-mount contract.
- Test: a `titleBar.center` component's effect runs setup once and cleanup
  never across a chat -> /skills -> chat round trip (red on base).
- Docs: plugin SDK guide covers the lifecycle contract and the page-header
  area.
2026-09-18 10:45:44 -07:00
teknium1
250bdcc311 fix(bot-mode): message_agent CLI-runner ack is queued + delivery_id like the other branches
Review finding on #114941: the CLI-runner branch of `_spawn_delivery` still
answered `status: sent` with only a `process_id`, while the live-owner branch
answers `queued`/`claimed` + `delivery_id` and the relay branch `queued`. One
tool, three vocabularies; `sent` also reads as a delivery receipt, which is the
misreading this PR exists to remove.

`_spawn_delivery` now returns `status: queued`, a `delivery_id` and the
`process_id`. For CLI-runner deliveries the id is the same sha256 of the DM
file path the runner pins when it admits to a live owner (`_dm_delivery_id`,
replacing three inline copies), so the ack and any live receipt correlate; the
relay branch passes the envelope id (also on its notification_error ack).
`sent_at` becomes `queued_at`. Tool schema, Bot Mode user guide and the
assertions in tests/tools follow; two invariant tests pin the CLI-runner and
relay ids (red before, green after).
2026-09-18 10:39:50 -07:00
teknium1
ef49567128 fix(bot-mode): message_agent ack says it is a dispatch hand-off, not a delivery receipt
The CLI-runner fallback of `message_agent` returns `status: sent` synchronously
while delivery runs in a background process; when that process dies (e.g. the
`hermes` entrypoint missing from PATH in the Docker image, #114628) the failure
only arrives as the background-process completion notification. The old
`detail` ("Message dispatched ... when the delivery completes, its notification
carries the reply") and the schema description ("delivery acknowledgement")
read as a delivery receipt, which misled the calling agent into reporting a
successful hand-off and routing around the tool.

Reword `detail` and the tool description to say the ack is the hand-off to a
background delivery process, not a delivery receipt, and that the completion
notification carries the outcome — the reply or the delivery failure. Update
the Bot Mode user guide to match. `status: sent` is unchanged (renaming to
`queued` + `delivery_id` is a return-contract change left to the maintainer).

Fixes #114628
Supersedes #111562
2026-09-18 10:39:50 -07:00
teknium1
09cf1b926f fix(state): reset forks leave the Python lineage walk too; /resume ranks lineages by activity
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.

Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.

Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.

Part of #114271
2026-09-18 10:39:13 -07:00
teknium1
1efc02716a fix(checkpoints): refuse rollback.diff from container-backed sessions on the TUI/Desktop RPC
The TUI/Desktop `/rollback diff <hash>` RPC still ran `mgr.diff(host_cwd, hash)`, which
stages the HOST working tree against a host checkpoint and presents it as the container
session's diff — the same operation `/rollback diff` and `/diff session` already refuse on
the CLI and messaging gateway. Refuse it with the same `unsupported_backend_reason`, via the
RPC error (what the TUI's `.catch(guardedErr)` already renders); `rollback.list` stays.
The backend-classification block moves into `_container_checkpoint_refusal` so restore and
diff share it. The docs sentence claiming `rollback.diff` remained available is corrected.
2026-09-18 10:38:08 -07:00
teknium1
dac4c60fbb fix(checkpoints): refuse host rollback and session diff from container sessions on the gateway too
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.

Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-18 10:38:08 -07:00
Sora-bluesky
3220b9ed2f fix(checkpoints): classify the backend on every /rollback instead of remembering it
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
2026-09-18 10:38:08 -07:00
Sora-bluesky
10ee9819f9 fix(checkpoints): do not feed container paths into host checkpoint storage
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.

Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).

Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
2026-09-18 10:38:08 -07:00
teknium1
303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00
teknium1
92bb5b92b8 fix: revert a quota-benched credential through the pool, not a private probe; cover /model and chained rotations
Reshape of the salvaged fix from #114513 (@whyyagswhy):

- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
  benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
  refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
  or round-robin order. The contributor's version called the private
  ``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
  and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
  429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
  the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
  the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
  by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
  hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
  ``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
  fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
  and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
2026-09-18 10:35:37 -07:00
teknium1
9cc7886afe fix(docs): kanban attachment removal keeps a blob other rows still reference
The File attachments section still promised that removing an attachment always deletes the on-disk file. With reference-counted blobs the file is only unlinked once no other attachment row points at it, so the user guide now states that condition and names the operator CLI path (hermes kanban attach-rm ATTACHMENT_ID).
2026-09-18 10:34:31 -07:00
teknium1
b941f41ead docs(acp): every tool call reaches a terminal status (event bridge) 2026-09-18 10:33:54 -07:00
teknium1
d9ca304047 docs(acp): note the failed-turn transcript boundary in the ACP lifecycle 2026-09-18 10:33:21 -07:00
teknium1
d60280fde8 fix: document how unprovable lease owners count toward max_concurrent_sessions
#113683 changed user-visible behaviour: an owner whose liveness cannot be proved
is now counted as live (MAX_CONCURRENT_SESSIONS) instead of refusing every claim
(SESSION_COORDINATION_UNAVAILABLE), and it only fences its own session id. Note
that in the existing max_concurrent_sessions section.
2026-09-18 10:32:48 -07:00
teknium1
a834988165 docs(state): name recovered placeholders in the stale-open sweep; exercise the real recovery path
sessions.md lists the state-owned sources the auto-prune sweep closes; add
`recovered`. The second test drives `_reconstruct_missing_sessions` itself so
the placeholder shape the sweep must age out is the one recovery writes, and
keeps a fresh placeholder open as the control.

Fixes #114730
2026-09-18 10:30:20 -07:00
teknium1
4ce832eeb9 fix(cron): bot-chat delivery cap bounds the bot's turn, not its exit linger
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).

Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:

- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
  report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
  HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
  {pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
  Report present -> the delivery is booked from it (real exit code and
  stream tails when the child exits within a short grace) and the still-
  lingering child is left running, drained and reaped by a daemon thread.
  No report by the cap -> the turn never ended: killed and booked as a
  timeout, exactly as before. The linger itself is untouched.

Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.

Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.

Fixes #113608

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:29:13 -07:00
teknium1
aefa503479 fix(browser): gave-up supervisor unregisters itself; tests trimmed to two invariants
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:

- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
  which logs the single final warning AND drops the supervisor from
  ``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
  (``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
  attached" instead of a dead entry, and the next browser call for the task
  starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
  (two invariants: bounded + unregistered after attach; first-dial failure
  still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
  main — that file is the opt-in real-Chrome E2E suite and its module-level
  skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
  reconnect.

Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.

Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
2026-09-18 10:28:40 -07:00
teknium1
b9d3c70303 docs(portal): name the exact per-profile sign-in path and fix the guide page too
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.

Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).

guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
2026-09-18 10:28:04 -07:00
PRATHAMESH75
0269ddb7ca docs(portal): the shared token store refreshes a login, it does not seed one
The Portal "Profile setup" section promised that every profile picks up a
Portal login automatically. The shared token store only merges fresher
tokens into a profile that already has a Nous state; a profile with no
providers.nous entry of its own fails closed at boot and is asked to run
the setup flow (hermes_cli/auth.py:1545, auth_nous.py:1056), because
profiles are independent islands (#111724). Correct the promise to
"sign in once per profile" and describe the fail-closed behavior, in both
the English and zh-Hans copies.

Fixes #114259
2026-09-18 10:28:04 -07:00
teknium1
d98b1480b9 fix(cli): a benched or signed-out credential prints its reason instead of the first-run provider wizard
`_runtime_credentials_ready()` collapsed "nothing is configured" and "the
configured credential is unusable right now" into one False, so a profile
whose only Nous credential was benched by a failed refresh (or quarantined
DEAD) was told "No inference provider is configured yet" and offered the
provider picker — re-running setup on an OAuth provider with single-use
refresh tokens can rotate the grant away from the session that was working
(#113720).

The readiness probe now returns the resolver's exception too
(`_probe_runtime_credentials`). At CLI startup only the resolver's
`no_provider_configured` verdict reaches the wizard; any other failure is
explained by `_explain_unusable_credentials`: `format_auth_error(exc)`
plus, from the provider's pool, "cooling down ... re-enters rotation in
about Nm" or "sign-in was lost (<reason>); run `hermes auth add <provider>`".
An actually-empty profile still gets the wizard.

Fixes #113720
Supersedes #113732 (@whyyagswhy): its `_inventory_other_providers()` gate
returns False for a Nous-configured profile, so the reported Nous case would
still have reached the wizard, and it printed no reason.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
2026-09-18 10:24:15 -07:00
teknium1
ab59c81f32 docs(skills): project skills follow the TUI/Desktop session workspace
User-visible note for #114359: a placeholder terminal.cwd no longer hides a trusted repo's skills in the TUI/Desktop.
2026-09-18 10:22:47 -07:00
teknium1
b44a481334 fix(video-gen): trust the operator-configured origin on the first download hop
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.

save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.

security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
2026-09-18 10:22:14 -07:00
teknium1
5236a58005 refactor(skills-hub): sitemap fetches reuse the hub's guarded GET
tools.skills_hub already owns _guarded_http_get (SSRF pre-check, guarded
client, bounded manual redirects that fail closed on a missing Location,
website-policy check). Route both the skills.sh sitemap index and each
<loc> sitemap through it instead of carrying a second inline guarded
client in the skills.sh adapter. The explicit Accept-Encoding: gzip
header goes away — httpx sends it by default.

Docs: the SSRF section now names the remote-party fetches the guard
covers and the LAN-provider opt-in.
2026-09-18 10:22:14 -07:00
teknium1
eb28bc1bf7 fix(kanban): host spawn refusals stop charging cards; oneshot-unit dispatch scope-wraps workers
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).

#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.

- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
  (RuntimeError subclass) so the dispatcher can tell a host refusal from a
  card failure.
- _record_task_failure(infrastructure=True): run + event carry
  `infrastructure: true`, consecutive_failures is left alone, the breaker
  never trips; the card stays `ready` with the linger remedy in
  last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
  is such a refusal inside the rate-limit cooldown window, so a dead bus is
  retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
  managed gateway.

#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.

- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
  old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
  is scope-wrapped when the user bus is reachable. Without a bus it degrades
  with a loud once-per-process warning naming the consequence and both
  remedies (enable-linger, KillMode=process) rather than refusing — the
  unit's KillMode/lifetime is unknowable and a long-lived sequencer without
  linger must keep working.

Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.

Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 10:21:42 -07:00
teknium1
a830061c9d fix(agent): recognise reversed "reasoning_effort 'none' unsupported" on every surface (#114460)
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:

- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
  between auxiliary_client._is_unsupported_parameter_error and the new
  classifier helper); is_reasoning_field_rejection() replaces
  is_reasoning_disable_rejected() with the same token gate and a
  symmetric "unsupported" window so both word orders match
  ("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
  model "requires reasoning" (a chat-only relay does not); a second
  reasoning-field rejection in the same turn is treated as spent and
  takes the fallback chain instead of replaying the identical request
  max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
  docs mention the reversed wording and the main-loop recovery.

Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
2026-09-18 10:21:05 -07:00
teknium1
fea2396981 fix(runtime): local alias guard keys on a missing endpoint, /model passes the alias URL
Follow-up on the picked #113768 commit so it clears the constraint that got
e9a54c48f2 reverted (a9fabe43c4): `/model <direct-alias>` resolves the alias
LABEL first and applied the alias endpoint only afterwards, so any guard that
fires on the label alone turns a working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery went red again
with the PR as-is).

- hermes_cli/runtime_provider.py::_raise_if_local_alias_missing_endpoint keys
  on the class (anything auth.resolve_provider maps to `custom` without a rung
  of its own; llamacpp keeps its managed-server fail-fast) instead of a second
  hardcoded alias set, requires the providers.<alias> block to actually carry a
  base_url (an entry without one falls through to OpenRouter too), and names
  the alias plus where to set its endpoint. OPENROUTER_BASE_URL never counts as
  the alias endpoint; an explicit api_key does not lift the guard.
- hermes_cli/model_switch.py::_creds_for_switched_provider hands a URL-bearing
  direct alias's base_url to the resolver as explicit_base_url, so the alias
  has an endpoint at the moment the guard runs (the same URL
  _apply_direct_alias_endpoint installs later).
- Tests trimmed to two invariants (raise-and-name incl. the OPENROUTER_BASE_URL
  / explicit-api_key non-lifts + bare-custom control; every endpoint source
  resolves to its URL). Docs: troubleshooting entry in the local Ollama guide.

Cron replay snapshots that store the resolved `custom` instead of the alias
are #109765's atom and untouched here.
2026-09-18 10:20:32 -07:00
teknium1
c2bf40d1ce fix(kanban): docs say a bound Project keeps the directory when the field is cleared; UI test pins only the unbind seam
With a bound project the settings dialog omits a blank directory from the PATCH and rename_board mirrors the project's primary folder, so clearing the field cannot fall back to scratch until the project is unbound; the Settings bullet now says so. The bundle test drops five wording-level substring asserts that pinned the payload shape already covered behaviourally by test_kanban_board_project_api.py, keeping only the switcher's unbind seam.
2026-09-18 10:19:54 -07:00
teknium1
0985aeb732 fix(kanban): board switcher shows the bound project with an unbind action
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.

Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.

Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
2026-09-18 10:19:54 -07:00
teknium1
4590ef8b58 fix(profiles): polled profile lists never walk skill trees; vanished skill dirs no longer abort enumeration
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.

- list_profiles(lazy_skill_count=True): skill_count is the last known value;
  a missing/aged entry schedules ONE background _count_skills per profile per
  60 s recheck window, so the request thread does zero skill-tree I/O and the
  refresh cadence is decoupled from the poll rate. The two polled callers and
  the REST fallback entry use it; the sync default (CLI, detail views) is
  unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
  os.walk with excluded/support dirs pruned) instead of rglob — half the
  fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
  uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
  projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
  skill add/remove immediately on the next refresh).

Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
  before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
          profiles.list 2189 + 2015, projects/tree 2177 + 2015;
          a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
  after:  GET /api/profiles 0 skill-tree stats/scandir on the request thread
          (counts land from the background refresh by the next poll),
          profiles.list 0, projects/tree 0; _count_skills (detail/control)
          still reports 200 with 203 scandir; the vanished-dir case returns
          199 and list_profiles() enumerates all 5 profiles.

Fixes #114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:18:40 -07:00
teknium1
89c6a8a145 fix(auth): Nous "not logged in" refresh failures leave rotation instead of a silent hour-long bench
A Nous pool entry whose refresh reached the resolver's state-shape raises
("Hermes is not logged into Nous Portal.", "No access token found ...",
"No refresh token is available ...") carried relogin_required=True but no
OAuth code, so `OAuthProviderFlow.is_terminal_refresh_error` classified it
transient: `_recover_failed_refresh` benched the only credential for an
hour with null last_error_* fields and nothing at WARNING (#113718).

Give those raises codes (`nous_auth_missing`, `nous_auth_missing_access_token`,
`nous_auth_missing_refresh_token`) and list them in the Nous flow's
terminal_refresh_codes, mirroring the Codex/xAI `*_auth_missing_refresh_token`
convention. The pool then quarantines, marks the row DEAD with the reason,
and logs the existing WARNING naming `hermes auth add nous`; every consumer
of the classifier (pool, singleton `_refresh_nous_or_quarantine`) agrees.
A failure that says nothing about the login (network error, non-JSON 5xx
body) still benches as before.

Fixes #113718
Salvages #113726 (@whyyagswhy) — tests kept, consumer-side hunk replaced
by the classifier fix.
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
2026-09-18 10:17:59 -07:00
teknium1
595f3a289c fix(telegram): split replies resume from the refused chunk, never re-send the head; per-chat send order + flood cooldown
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.

- send() now reports a mid-split refusal through the existing partial_overflow
  contract (the key _edit_overflow_split already sets and the stream consumer
  reads): delivered_chunks / total_chunks / last_message_id, plus
  undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
  (flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
  have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
  _resume_partial_send hook (default None = keep the partial failure, never
  re-send the head; the plain-text fallback is skipped for partials too). The
  Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
  only around the API calls — never across the reconnect wait — on send() and the
  media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
  capped 300s); sends inside the window fail closed locally with the same
  flood_control:<s> result and no API call, so ledger recognition and redelivery
  timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
  "waiting Ns" followed by no wait.

Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.

Fixes #114396

Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
2026-09-18 10:17:23 -07:00
teknium1
f39ae9000c docs(sessions): say what user_id holds for desktop and dashboard sessions 2026-09-18 10:16:46 -07:00
teknium1
1b50849d27 docs(clarify): one clarify timeout for every surface; say where the deadline exemption lives
- hermes_cli/config_defaults.py: the agent.clarify_timeout comment claimed
  "CLI clarify blocks indefinitely and ignores this"; the CLI reads it (via
  resolve_clarify_timeout) — describe the real contract and the legacy
  clarify.timeout precedence once.
- agent/tool_executor.py: the sequential-middleware docstring named clarify
  as exempt from the generic deadline while _SEQUENTIAL_DEADLINE_EXEMPT_TOOLS
  does not list it; the exemption is the _NEVER_PARALLEL_TOOLS inline path
  (61645cde82). Point at it so nobody re-adds a dead set entry.
- website/docs: configuration.md Clarify section covers all surfaces and the
  deadline exemption; discord.md (+ zh-Hans discord/telegram) still said the
  default was 600 s.
2026-09-18 10:16:14 -07:00
teknium1
69b63e7407 fix(config,update): replace an unwritable SOUL.md symlink; say when the pre-update snapshot failed
Two follow-ups on the salvaged commits:

- `_ensure_default_soul_md`: widen the cyclic-only (ELOOP) branch to any symlink the seed
  cannot write through — a link dangling into a missing directory raises ENOENT on the same
  write and bricked home init the same way (HomeInitializationError -> exit 75 -> supervisor
  relaunch storm). Shape: write through the link first (a working link stays operator
  wiring), and only on OSError seed the default in place of the link via mkstemp + replace.
  Tests trimmed to two invariants (cyclic/dangling replaced; resolving link preserved).
- `_run_pre_update_backup`: the quick snapshot stays best-effort by design (8ed599dc05:
  "a broken backup never blocks the update"), but a failure was swallowed at DEBUG level, so
  the user only learned from the receipt afterwards that no recovery point existed. Print a
  stdout warning with the reason and continue; together with the receipt skip/fail split the
  receipt now says "failed" only when a snapshot was requested and not produced.
- docs: updating.md describes the warning and the receipt split.
2026-09-18 10:15:11 -07:00
teknium1
6a6d37075a docs(feishu): describe post-message file attachments and caption-preserving inlining
User-visible behaviour changed by the post `files` fix: every top-level attachment is
collected (folders skipped) with an [Attachment: name] marker, and small .txt/.md content
is appended after the caption instead of replacing it.
2026-09-18 10:14:38 -07:00
teknium1
3ab6f36613 docs: caller thinking-off beats auxiliary.<task>.reasoning_effort 2026-09-18 10:14:06 -07:00
teknium1
fe6b5cfd15 docs: explicit auxiliary providers walk their own fallback_chain on auth errors 2026-09-18 10:13:32 -07:00
teknium1
964fbaae2c fix(auth): status snapshot peeks the credential pool instead of leasing it
`get_codex_auth_status` / `get_xai_oauth_auth_status` back every credential-gated
listing (`/model` picker rows, `hermes doctor`, dashboard auth cards). The shared
`_pool_first_oauth_status` read the pool with `select()`, which is a runtime lease:
it refreshes an expiring single-use token and, when that speculative POST fails
transiently (503, endpoint unreachable), benches the entry with a persisted
exhaustion cooldown. The picker then rendered the provider as an unconfigured
skeleton ("needs setup" / "0 models") while the runtime resolver kept serving
the same credential. On a round-robin pool the read also rotated and persisted
the priority order.

Read the pool with `peek()`: no refresh, no accounting, no rotation, no persist.
Refreshing stays with the runtime resolver reached through the `resolve`
fallback, whose failures persist nothing (same design as the Nous status).

Slimmer redo of #114379 by @Finn763 (observe flag threaded through four pool
methods plus non-refreshing resolvers); fixtures adapted from that PR.

Live: pool-only openai-codex entry, expired access token, stubbed token endpoint
returning 503 — before: 1 refresh POST, entry persisted `exhausted`,
`has_available()` False, picker row gone; after: 0 POSTs, entry untouched,
`has_available()` True, picker row present with the catalog. Runtime `select()`
still refreshes (control).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
2026-09-18 10:12:25 -07:00
teknium1
23f4db708f test(mcp): pin hermes mcp test exit codes; document them
Two invariant tests: the handler returns 0/1/3 for connected / connection
failed / not in config, and the `hermes mcp` CLI dispatcher forwards that code
to `main()` (which already exits on an int return). The MCP guide documents the
codes so probes and watchdogs can stop parsing the output.
2026-09-18 10:11:46 -07:00
teknium1
2164221844 fix(send): name the default-root gateway and external secret sources in the 'not configured' error
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.
2026-09-18 10:11:14 -07:00