Commit Graph

36614 Commits

Author SHA1 Message Date
liuhao1024
b08ec00a70 fix(desktop): hand preview guest links to the audited opener via a guest preload
Preview-pane guest pages (Streamlit's traceback "Ask Google" / "Ask …"
buttons, plain `<a target="_blank">` anchors) could not open anything: the
`<webview>` has no `allowpopups`, so Chromium drops the popup before any
handler runs. Opening from `setWindowOpenHandler` is banned
(GHSA-9f4c-93c8-jc8g, window-open-policy.ts), so this adds an explicit
click bridge instead:

- main.ts installs a guest preload through `will-attach-webview`, keyed on
  the `persist:hermes-preview` partition only.
- The preload forwards a clicked `_blank` anchor's resolved href to the
  host renderer via `ipcRenderer.sendToHost`; it opens nothing itself.
- PreviewPane admits the URL and routes it through the existing audited
  `hermes:openExternal` IPC.

Salvaged from PR #112959 (squash of its two commits).

Fixes #112941
2026-09-17 09:05:24 -07:00
teknium1
091b8c4f53 docs: call out the client.capabilities fail-closed gate as breaking for third-party WS clients
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
2026-09-17 09:04:38 -07:00
teknium1
2afb405337 fix(approval): withdraw the queue entry when no client can answer; commit choices under the lock
Why: for an old WebSocket client that never advertised server→client requests,
send_async already failed fast (on_result(None)) but _emit_approval_request ignored
it, so the approval wait — owned by tools.approval's queue, not server_requests —
still idled for the whole approvals.timeout (300s) with no prompt anywhere. The
same held for a -32601 error frame. on_result(None) now withdraws the queue entry
(withdraw_gateway_approval: cancelled cause, never a user deny) so
_await_gateway_decision returns at once; the agent sees a withdrawn prompt.

resolve_gateway_approval committed entry.result/reason/event.set() AFTER
releasing _lock, so _drop_entry (which reads result and leaves the queue under
the lock) could still pop-and-lose a choice the client was acked for. Every
committer (resolve, clear_session, unregister_gateway_notify) now commits inside
the same critical section that pops the entry, which is what _drop_entry's
docstring promised.

_drop_entry mapped a withdrawn wait to settle("set") — the raw poll-state token
went out as the request.cancel reason. It is now session_closed, and a choice
from another surface is resolved (both RequestCancelReason values).

Tests: restore test_server_request_error_response_fails_fast (an error frame
settles send() to None promptly); the approval fail-fast probe from the review
(silent WS peer, timeout 3 -> was 3.2s, now immediate); a lock-instrumented
resolve test; a settle-reason test. The test_protocol `server` fixture imports
server_requests before its sys.modules patch window so the module server.py binds
its sinks on is the one tests import — the PR's fail-fast test only passed in a
full-file run before (order dependency).

Part of #112548
2026-09-17 09:04:38 -07:00
teknium1
6c057dbaa5 test(desktop): the no-heartbeat check expects the capability advertisement
Same stale invariant the ui-tui test had: since the shared channel sends one
client.capabilities frame on every gateway.ready, "no frames" is no longer
what this case guards. What it guards is that an older backend without
heartbeat gets no gateway.ping and stays open, so assert the wire holds
exactly the advertisement.
2026-09-17 09:04:38 -07:00
teknium1
f7e70e954e test(ui-tui): the no-heartbeat check expects the capability advertisement
Every gateway.ready now puts one client.capabilities frame on the wire, so
"no frames" is no longer the invariant; "no gateway.ping" is.
2026-09-17 09:04:38 -07:00
teknium1
f9d178f78e fix(tui_gateway): old app builds no longer stall the agent on clarify/approval; late approval choices count
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.

Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
  the approval lock while leaving the queue, so an /approve that lands after
  the deadline check but before the entry is dropped is an answer, not a
  timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
  restated pre-existing behaviour is replaced by the two capability
  invariants (never advertised → fails fast; advertised → frame written,
  waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
  (4371ed34a9); nothing to change.

Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
2026-09-17 09:04:38 -07:00
teknium1
1038e0fc77 fix: descendant sweep spares the stopper, reports leftovers, and skips subtrees of unkillable roots
Review follow-ups on the #112631 descendant sweep:

- `_posix_descendants` excluded nothing of the caller. `hermes dashboard --stop` /
  `hermes update` run from a shell escape inside the hosted Chat TUI are same-session
  descendants of the backend, so the sweep SIGTERMed its own process mid-run (probe:
  caller in snapshot, backend SIGKILLed, "survived" line never reached) — the POSIX
  twin of the Windows #98814 hazard. The caller, its subtree and its ancestor chain are
  now pruned from the same `ps` snapshot; an ancestor's *other* children (the wedged
  TUI) stay in the sweep.
- Descendants still alive after their SIGKILL grace were silently dropped and the stop
  declared complete; they now land in `failed` so `_kill_stale_dashboard_processes`
  prints them — a wedged ui-tui still holding the deleted WAL is the incident itself.
- The Desktop boot reaper SIGKILLed descendants of roots whose own kill raised
  (EPERM: not ours), orphaning a foreign tree half-way. Descendants now carry their
  root and are skipped when the root is in `failed`. First test of that hunk.
- `_is_detached_session_leader` documents that compute_host / slash_worker children
  are exempt by the same rule and rely on their own ppid watchdogs.
2026-09-17 09:04:11 -07:00
teknium1
f3bdcd0877 fix: dashboard stop sweeps the wedged descendants that outlive the SIGKILLed backend
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.

_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.

Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.

The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).

Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
2026-09-17 09:04:11 -07:00
teknium1
3f80dc9a5a fix(plugins): stream observer hooks and event subscribers fail-report once, not per event
Two per-call WARNING surfaces were still outside the warn-once reporter:

- agent/plugin_stream_hooks.py::_worker — the on_stream_start/on_stream_delta/
  on_stream_end consumer, which fires once per streaming delta (far more often than
  per tool call). A mis-declared callback (signature naming tool_data) logged
  "Hook ... raised" at WARNING on every delta: 20 deltas -> 20 WARNING lines.
- hermes_cli/plugins_dispatch.py::_deliver_event — plugin event subscribers that
  raise identically were warned on every emit.

Both now go through PluginManager._report_hook_failure (keyed by module/qualname,
cleared on unload): the first failure warns and names the fields the hook/event
provides, identical repeats are DEBUG. Skip-and-continue semantics are unchanged.

Part of #111922
2026-09-17 09:03:29 -07:00
teknium1
0f98d5a1d9 fix(plugins): execution-chain middleware failures are reported once, not per call
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.

Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.

Part of #111922
2026-09-17 09:03:29 -07:00
teknium1
dea7c21bc7 docs(i18n): zh-Hans mirrors of the vision.embed_target_bytes / max_calls_per_image sections 2026-09-17 09:02:55 -07:00
teknium1
87f29fb1b9 fix(vision): reserve the per-image embed slot atomically
repeat_refusal checked the counter and record_embed incremented it in separate
lock sections, so a concurrent tool batch on the same image (the incident issued
4 at once) all passed a check taken before any of them recorded and the cap of 3
let 6 embeds through. Reserve the slot inside the check's lock section and
release it when the embed then fails.
2026-09-17 09:02:55 -07:00
teknium1
f4452169d1 fix(gateway): busy-mode steer also reaches the parent's active subagents
A parent blocked inside delegate_task only drains its steer queue after the tool
returns, i.e. after the child finishes. On Telegram (busy_input_mode=steer, the
/steer command, and the PRIORITY path) the gateway queued the text on that parent
alone and acked "Steered into current run", so a looping delegated child never saw
it (#112095). Fan the text out to running_agent._active_children — the same
identity-scoped snapshot interrupt() uses — and make the ack say where it went.
2026-09-17 09:02:55 -07:00
teknium1
ed7eb5685a docs(vision): document vision.embed_target_bytes and vision.max_calls_per_image
User-visible knobs need a home: the Vision feature page explains why native embeds ride
the session and what each key does (subagent-only default cap, clamp range), the
configuration reference points at it next to auxiliary.vision so the two sections are
not confused, cli-config.yaml.example carries the commented block, and tools/AGENTS.md
names vision_tools_history_budget.py as the single owner of embed-cost policy.
2026-09-17 09:02:55 -07:00
teknium1
abc56351aa fix(vision): delegated subagents stop re-loading the same image on the native fast path
A delegate_task child asked to transcribe five screenshots called vision_analyze 158
times on those files (full loads alternating with region crops) until its provider
quota ran out; every native load bakes the image into history and is re-sent on each
later API call, and nothing refused the repeat (#112095).

tools/vision_tools_history_budget.py now keeps a per-(session id, resolved source)
embed counter. Region crops share their file's key. When the cap is reached
_vision_analyze_native returns a tool error that says the image has already been loaded
N times and to answer from what is visible or ask the user, instead of another embed.
Only successful embeds count, so a failed read never eats the budget.

Default: `vision.max_calls_per_image` unset caps only delegated subagents
(agent.delegation_context.is_delegated_child_process_context) at 3 — they run
unattended and CLI /steer targets the parent — while the main agent stays unlimited;
an explicit value applies everywhere and 0 means unlimited.

Co-authored-by: Evi Nova <66773372+Tranquil-Flow@users.noreply.github.com>
Co-authored-by: gumclaw <gumclaw@gumroad.com>
2026-09-17 09:02:55 -07:00
Mohamad Kanso
f37336522b feat(vision): vision.embed_target_bytes replaces the hardcoded 256 KB native embed budget
The native vision_analyze fast path (and the browser screenshot twins) downscaled every
embed to a fixed _EMBED_TARGET_BYTES = 256 KB. That is fine for photos but turns a
1080x2340 phone screenshot of a table into 540x1170, which the model then reads as
"unreadable" and re-requests (#112095).

The budget is now `vision.embed_target_bytes` in config.yaml (default unchanged: 256 KB,
clamped to 64 KiB..4 MiB so one setting cannot make every later request a multi-megabyte
resend), resolved in the new topical sibling tools/vision_tools_history_budget.py and read
by vision_analyze, browser_vision and browser_exec screenshots alike.

Ported from #112947 by @MohamadKanso (resolver + clamp), relocated out of the facade.
2026-09-17 09:02:55 -07:00
teknium1
428056e7b9 fix(auth): Claude Code dead refresh token is reported once; CLAUDE_CONFIG_DIR honoured; no 'claude setup-token' hints
Review follow-up on the sibling refresher path (_refresh_oauth_token), which
the PR title already claimed but only the pool path delivered:

- A refresh token the endpoint rejected terminally (invalid_grant etc.) is
  fingerprinted in a process-local set; later attempts skip the POST and log
  at DEBUG instead of replaying the dead token and re-firing the WARNING on
  every resolve. A rotated token (user re-logs into Claude Code) has a new
  fingerprint and is tried normally.
- claude_code_credentials_path() honours CLAUDE_CONFIG_DIR (blank = unset,
  same rule as hermes_cli.foreign_sessions), so the opt-out the PR body
  documents actually exists and matches the Claude CLI's own relocation.
- The three remaining 're-run claude setup-token' hints inside the refresher
  now say 'hermes auth add anthropic' like the rest of the PR; setup-token
  does not repair the Claude Code file anyway.
2026-09-17 09:02:26 -07:00
teknium1
a7b13e68a9 test(auth): terminal 401 on a manual Codex grant is 'dead', not 'exhausted'
The parent commit deliberately marks a pool row whose refresh token is
terminally rejected (invalid_grant) STATUS_DEAD via _mark_dead_refresh_grant,
so a manual `hermes auth add` login leaves rotation until re-auth instead of
being benched for a TTL and replaying the dead token every hour (#113023).
STATUS_DEAD already was the pool's "permanent OAuth failure" status, and the
new tests/agent/test_credential_pool_terminal_refresh_visibility.py asserts
exactly this for a manual:device_code openai-codex row.

tests/hermes_cli/test_auth_pool_operations.py[401] still encoded the old
contract and went red. Only the 401 leg changes; the 503 leg (transient
failure) keeps asserting 'exhausted', which is the control for the new
terminal/transient split.
2026-09-17 09:02:26 -07:00
teknium1
d6add16059 fix(auth): dead OAuth logins are reported once and leave rotation; hints point at hermes auth
A terminally rejected refresh token (invalid_grant / invalid_token /
refresh_token_reused) is the moment a login is lost. Three gaps remained
after #113197 (which added the WARNING for Codex/xAI/Nous):

- Anthropic: the token endpoint's HTTPError carried no classifiable code,
  so a dead Anthropic / Claude Code grant fell through to a transient
  'exhausted' bench at DEBUG and was replayed every hour. The endpoint
  error is now a structured AnthropicOAuthError (status + OAuth error
  code); _recover_failed_refresh logs one WARNING with the repair command
  and marks the row DEAD. A dead grant is not replayed at the fallback
  endpoint. The auxiliary Claude Code refresher (sibling path) warns the
  same way. Claude Code's own credentials file is never touched.
- Codex/xAI/Nous: _quarantine_sources drops only singleton-seeded rows, so
  an independent `hermes auth add` (manual:*) login survived unmarked and
  re-fired the WARNING on every later refresh attempt. The survivor is now
  marked DEAD (leaves rotation until a write-side re-auth clears it).
- Hints on live Hermes paths (anthropic 401 troubleshooting block, the
  no-credentials error) recommended an external CLI's login command,
  which cannot repair Hermes' own login; they now say
  `hermes auth add anthropic` / `hermes auth list anthropic`.

Docs: credential-pools.md documents the dead-login behaviour.
Part of #113023.
2026-09-17 09:02:26 -07:00
teknium1
ec14e8dfa2 test(desktop): session reads no longer expect the foreground dial tag
sessionScoped() strips the tag since the review fix; the assertions added by the
first commit of this branch still expected it on session reads and failed CI.
Profile writes keep the tag.
2026-09-17 09:01:59 -07:00
teknium1
6ae79db3eb fix(desktop): explicitly scoped session reads keep the ambient dial priority
Moving the foreground tag into capabilityScoped() made every helper that
spreads it dial foreground, including sessionScoped() and therefore
getSession/getSessionMessages. resolveStoredSession probes every OTHER
profile sequentially with getSession(id, profile) while resolving a
remembered or 404'd session id — a boot-time path, not a scope-selector
click — so it would have cold-started each profile's backend on the pool's
reserved foreground slot. sessionScoped() now strips the tag: the tag is
for the user-facing scope selectors (Settings, Capabilities, Messaging),
session lookups stay on main's background default.

Part of #111651
2026-09-17 09:01:59 -07:00
teknium1
53517082c8 fix(desktop): every explicitly scoped api/ read dials its profile backend foreground
Fold the scoped-dial priority rule into profileScoped()/capabilityScoped()
themselves: an explicit scope (string or object) now carries
priority 'foreground' from the scope helper, so every api/*.ts request
that targets a selected profile — skills hub/official lists, learning
nodes, toolset config/models, MCP, messaging/pairing, memory providers,
sessions, profiles — inherits it. The per-call `...scopedDialPriority()`
spreads (config/models/mcp/skills/toolsets) and the helper are removed;
the ambient path (`undefined`) and the deliberate `null` → primary path
stay untagged so hydration cannot take the reserved foreground slot.

WHY: the first pass tagged the six Settings pages, then models.ts, then
Capabilities' three page-open reads, and each round left siblings behind
(getOfficialSkills / getSkillHubSources / getLearningNode /
getToolsetConfig / getToolsetModels / the Messaging page) still queueing
behind roster hydration when the selected profile's backend was cold —
an infinite spinner once the pool's background slots were saturated. A
rule that must be re-spread at every call site is the failure mode; one
that rides on the scope helper cannot be forgotten by the next helper.

messaging.ts spread profileScoped() into request BODIES for the
body-scoped pairing/onboarding endpoints; those now bind the scope once
and forward only `profile`, so the dial tag never reaches the backend.

Tests: hermes-capability-scope.test.ts gains two invariants (red on
base): every scoped helper outside models.ts carries the tag and `null`
does not; the tag never leaks into a body. Exact-shape request
assertions in hermes.test.ts, session-transcript-paging.test.ts and
use-desktop-integrations.test.tsx gain the tag for explicit scopes.

Fixes #111651
2026-09-17 09:01:59 -07:00
teknium1
07c92d675a fix(compression): repeated summary stall escalates to the deterministic fallback summary
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).

WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.

Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.

Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
2026-09-17 08:59:50 -07:00
teknium1
ba2bcfa6e7 fix(gateway): auth-fallback test stub accepts the target_model kwarg the fallback resolver now passes
`_try_resolve_fallback_provider()` now resolves each fallback entry with
`target_model=entry["model"]` (intended: the model-keyed ladder rungs must see the
model the fallback will send, not config's primary default). The auth-fallback test
in test_session_model_override_routing.py monkeypatches resolve_runtime_provider with
a stub whose signature lacked `target_model`; the resulting TypeError was swallowed
by the per-entry `except Exception: continue`, the fallback chain returned None, and
the gateway re-raised the primary AuthError ('No Codex credentials stored').

Source behavior is correct (real resolve_runtime_provider accepts target_model; the
other stubs in test_api_server/test_session_api were already updated), so this only
brings the stub in line and additionally asserts the fallback rung is resolved
against the fallback model, locking in the PR's intent.
2026-09-17 08:59:22 -07:00
teknium1
931b5ff9e7 fix(opencode): every credential-resolution surface keys off the model it will send; family heal only for built-in providers
Closes the remaining atoms of #112600.

A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
   its siblings still resolved credentials against config's `default`: the CLI
   auth-fallback rung, `--resume` credential re-resolution, the gateway
   provider-override helper (channel overrides, persisted /model switches,
   API-server provider refresh), the gateway fallback chain, the TUI /model
   switch-from runtime and ACP agent construction. With a `*-free` default the
   OpenCode free-tier rung fired first and a Go-only model was built against
   the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
   the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
   optional `target_model` and the two test stubs of it accept the kwarg.

B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
   provider matched by opencode_provider_family, including custom providers
   merely named after a family (`opencode-go-bridge`, #85589) whose relay the
   user declared explicitly in `providers:`. The family heal now applies to the
   built-in canonical providers only; custom prefix-named providers keep their
   per-model api_mode routing and /v1 handling. Documented in the providers
   guide.

C) Same function: the official-host check uses parsed.hostname (a port no
   longer defeats the heal) and only the path is edited, so query/fragment
   round-trip instead of being dropped.

Fixes #112600
2026-09-17 08:59:22 -07:00
teknium1
5b80838ea2 fix(aux): local-server aliases never borrow OPENAI_API_KEY for the lane's base_url
A keyless `provider: ollama|vllm|llamacpp` aux lane ran through the anonymous
`custom` key chain, so with OPENAI_API_KEY in the environment (or a main key on
the same host) the lane sent that secret as `Authorization: Bearer` to whatever
host `base_url` named. The alias means "this is my own server": only an explicit
`api_key` or the `no-key-required` placeholder is ever sent. `provider: custom`
keeps its documented OPENAI_BASE_URL + OPENAI_API_KEY contract.

Review follow-up on #113961; extends the existing parametrized invariant with the
OPENAI_API_KEY-set case (red on the previous head).
2026-09-17 08:56:27 -07:00
teknium1
200f8a6b4e fix(aux): keep the /v1 tail on the config path and trim the salvage to two invariants
Follow-up to the cherry-picked #106018 (@liuhao1024). The alias-table fix resolved the
keyless lane, but the reporter's second variant (provider: ollama + bare base_url +
explicit api_key → 404) still failed through the real config path:
_resolve_task_provider_model collapses "base_url + api_key" lanes to "custom" before
resolve_provider_client runs, so the custom branch never saw the ollama alias and
skipped the /v1 tail. Keep the local-server alias identity on both base_url paths
(config lane and explicit kwargs via _preserve_provider_with_base_url) so the tail
applies whenever the user wrote provider: ollama/vllm/llamacpp.

Shape: inline the 5-line _bare_host_base_url helper (urlparse never raises on str;
no new facade helper), trim the five contributor tests to two invariants that drive
the reporter's lane through _resolve_task_provider_model, and document the alias
group under the auxiliary provider list.

Live: stand-in mirroring Ollama's /v1 OpenAI-compatible surface — before: RuntimeError
"no API key was found" (keyless) / 404 on /chat/completions (explicit key); after: both
POST /v1/chat/completions and return OK, sync and async; provider: custom + /v1 unchanged.
2026-09-17 08:56:27 -07:00
liuhao1024
b7e0d715be fix(agent): route explicit local-server provider aliases (ollama/vllm) through the aux custom branch
An auxiliary lane with an explicit `provider: ollama` + local base_url + empty
api_key kept the ollama identity, matched no PROVIDER_REGISTRY entry (the aux-side
alias table lacks the local-server alias group hermes_cli.auth has), and raised
"Provider 'ollama' is set in config.yaml but no API key was found" on every call
(#106010) — while `provider: custom` against the same server worked.

- Mirror the hermes_cli.auth local-server alias group (ollama/vllm/llamacpp/
  llama.cpp → custom) in the aux-side _PROVIDER_ALIASES.
- Append the /v1 tail to bare host base_urls coming through those aliases: the
  OpenAI-compatible surface of these servers lives under /v1, so without it the
  first call 404s on the native API.

Fixes #106010
2026-09-17 08:56:27 -07:00
teknium1
f7631c0ee7 fix: fold a duplicate custom_providers row at the match level so its own models still route
Dropping an identity-equal legacy `custom_providers` entry from the candidate list also dropped the
models only that entry declared: `/model gpt-5.4-mini` (declared by the legacy row but not by
providers.relay) fell through to the current provider (openrouter) instead of routing to the shared
endpoint — a fail-open flip versus the pre-#112788 behaviour, which routed it to custom:relay.

Keep every entry as a candidate and credit a duplicate's hits to the providers row it duplicates, so
the collapse loses no declaration. `_duplicates_configured_row` now returns the owning slug.

Also require credential equality (not just endpoint) in the raw-list `provider_key` shortcut: a
hand-written entry stamped `provider_key: relay` on the same URL but a different key_env is a
distinct route and stays ambiguous, matching the issue's "preserve ambiguity when credential differs"
criterion.
2026-09-17 08:55:04 -07:00
teknium1
01d2e6c3d1 fix: /model collapses compat views of one configured provider by identity, keeps custom:<slug> sessions
`_configured_provider_matches` now drops a `custom_providers` entry when it is a second view of a
`providers.<slug>` row: either that row's compat projection (provider_key names the slug AND the
endpoint matches) or a legacy duplicate with the same provider identity (normalized name, endpoint,
credential identity, api_mode — read through the same normalizer that builds the compat view).
Any difference in endpoint, credential or wire protocol keeps two rows distinct.

Why: the first pass (#113103) only filtered entries whose provider_key equalled a providers slug.
A hand-migrated config that kept the same endpoint in both sections was still reported as
"declared by multiple configured providers (custom:relay, relay)", and on the raw-list fallback
(gateway/agent_init pass cfg['custom_providers'] verbatim when the compat view fails) a hand-written
provider_key pointing at a different endpoint was silently hidden instead of ambiguous.

`_route_configured_provider` treats `custom:<name>` of a projected row as an alias of its slug when
checking whether the session is already on the matched provider, so a session running on
`custom:relay` that types a model providers.relay declares stays on `custom:relay` instead of
flipping target_provider/explicit_provider to `relay`.

Closes the remaining atoms of #112788. Identity-tuple approach ported from the
`_configured_provider_identity` hunk of #112794.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-17 08:55:04 -07:00
teknium1
3187d68b49 fix(kanban): goal_mode workers keep the live tool feed in their worker log
The dispatcher forced `-Q` on goal_mode cards because cli.py only ran the
kanban judge loop in the fully-quiet one-shot branch. `-Q` strips every tool
callback, so the dashboard's Worker log (and `hermes kanban log`) stayed
blank for the whole run while non-goal cards logged normally; users read
that as "the worker is doing nothing" (Discord report, Sep 2026).

Run the judge loop on the `-q` path too, driving follow-up turns through
cli.chat so each turn's tool activity lands on stdout (= the worker log),
and print each judge verdict there as well. The dispatcher spawns goal_mode
and one-shot workers with the identical argv; the mode travels only in
HERMES_KANBAN_GOAL_MODE. The `-Q` hook stays for manual quiet runs.

Live A/B (real `hermes kanban dispatch` + spawned worker against a scripted
loopback provider that answers a terminal tool call then text, judge always
"continue", --goal-max-turns 2): base log = 94 bytes, 0 tool-feed lines,
0 verdict lines; fix log = 1910 bytes, 4 tool-feed lines, 3 verdict lines;
both arms end blocked "exhausted 2/2 turns" (loop behaviour unchanged).
2026-09-17 08:55:01 -07:00
teknium1
d4dfbba485 fix(cron): pin the worker's PYTHONPATH from the sanitized env in a sibling helper; skip under a wheel install
The pin landed inline in the cron/scheduler.py facade and rebuilt PYTHONPATH from
raw os.environ. That resurrected the Hermes-owned entries build_subprocess_env
had just stripped (runtime site-packages, launcher spellings of the repo root)
instead of extending what the sanitizer kept. It also ran unconditionally: under
a wheel / pipx / uv-tool install `Path(__file__).parent.parent` IS purelib, so the
pin hoisted site-packages above the stdlib on the worker's sys.path instead of
being a no-op.

cron/scheduler_worker_env.py::pin_hermes_tree_on_pythonpath prepends repo_root
to worker_env's own PYTHONPATH and returns the env untouched when repo_root is
sysconfig purelib (cron/ is already importable there). Test: A/B with the
previous inline pin -> the raw-environ-only entry leaked into the spawn env.

Part of #112729: this hardens the PYTHONSAFEPATH / cwd-not-checkout case; the
reporter's failure did not reproduce on main from cwd=checkout, so the real
worker stderr tail (now captured by e094e25b26 / f857a4ed99) is still needed.
2026-09-17 08:54:39 -07:00
teknium1
9269b19e4f fix(cron): restart-safe worker imports the gateway's own checkout instead of relying on cwd
The external cron worker is spawned as `sys.executable -m cron.scheduler`. Its
entry module is `cron.scheduler`, not `hermes_cli.main`, so it never runs the
bootstrap that puts the checkout on the gateway's sys.path; it only imported
`cron` at all through the implicit `-m` cwd entry (cwd was already the repo
root). That implicit path breaks on real hosts: a venv whose editable install
maps a moved or deleted checkout (the finder in this repo's own venv points at a
worktree that no longer exists), or a host that sets PYTHONSAFEPATH so `-m`
ignores cwd. The worker then dies with "No module named 'cron'" before its
ownership ack and every fire records
"cron external worker exited before ownership acknowledgement (exit 1)" (#112729).

The shared subprocess sanitizer strips Hermes-owned PYTHONPATH entries because
user children must not see our tree; this child IS Hermes, so after the env is
built the worker gets an explicit PYTHONPATH: the checkout `cron/scheduler.py`
lives in first, then whatever PYTHONPATH the gateway itself was started with.
The cwd stays the same checkout. Kanban workers spawn `-m hermes_cli.main` and
already get the bootstrap, so this is the only affected entry point.

Live probe (worktree python from cwd=/tmp with PYTHONSAFEPATH=1, real
_launch_external_cron_worker + real Popen): before, the worker stderr tail is
"Error while finding module specification for 'cron.scheduler'
(ModuleNotFoundError: No module named 'cron')"; after, the worker imports
cron.scheduler, loads the payload and reaches the durable-ownership check.

The two review minors on the first pass (worker unlinks its own stderr capture;
comments no longer claim DEVNULL) already landed in f857a4ed99 and are covered
by that PR's tests.
2026-09-17 08:54:39 -07:00
teknium1
7b085a734b fix(desktop): a logged-out gh is asked again on the next update check
githubTokenFromGhCli cached the outcome of `gh auth token` for the whole
process, including "none": a user who ran `gh auth login` after launching
the Desktop stayed anonymous (and rate-limited) until an app restart, since
only a 401 on a real token cleared the cache. Only a token is cached now; a
null answer (gh missing, logged out, hung) is re-asked by the next hourly
check, one bounded spawn per check at most.

Part of #112615
2026-09-17 08:54:13 -07:00
teknium1
c61a11474b fix(desktop): update check falls back to the gh CLI login before going anonymous
Closes the last atom of #112615. A Desktop app started from the Dock, Finder or a
desktop launcher inherits a minimal environment, so the GITHUB_TOKEN / GH_TOKEN rung
landed in #113190 never fires for exactly the users the issue is about (shared exit
IP, 60/hour anonymous budget exhausted by neighbours). The update check now walks the
same ladder as the Python GitHub client (tools/skills_hub_github.py::GitHubAuth):

1. GITHUB_TOKEN, then GH_TOKEN, from the launch env (unchanged).
2. `gh auth token` — execFile with an argv array (no shell), stdin closed, 3 s
   timeout, windowsHide. `gh` is resolved on PATH plus the GUI-safe install
   locations backend-env already appends for the backend (Homebrew, /usr/local,
   ~/.local/bin, nix profile, the Windows GitHub CLI installer dirs). The outcome
   (token or none) is cached for the process lifetime so gh runs at most once.
3. Anonymous.

A rejected credential (401) still retries anonymously; the log line names the
source (env vars vs gh login) and never the token, once per source per process.
A rejected gh token also drops the cache so a re-login is picked up by the next
check. envTokenRejected is renamed githubTokenRejected since both rungs use it.

Live: with PATH=/usr/bin:/bin:/usr/sbin:/sbin and no env token, the module found
gh, resolved a credential in 135 ms and api.github.com reported a 5000/hour core
budget (anonymous: 60). A logged-out gh and a hanging gh both resolved to null
(anonymous) in 5 ms / 3001 ms.

Docs: the credential ladder is documented on the Desktop "Updating" page and in
the environment-variables reference.
2026-09-17 08:54:13 -07:00
teknium1
a1238adfae fix: log the provider in the vision-gate skip and prove capacity 429s do not rotate the pool
Review follow-up. The vision capability gate in _try_main_agent_model_fallback
said "main agent model %s" but formatted main_provider, so operators reading
the INFO line would take a provider name for a model name. Reword to "main
agent provider" and keep logging only the provider, matching the sibling
_vision_main_provider_client (model names can trip the clear-text-logging
scanner).

test_upstream_capacity_429_is_not_a_credential_to_bench only checked the
_is_overloaded_error predicate, so a regression that re-added the pool
rotation for a capacity 429 would still have passed. Parametrize it over a
capacity 429 (must not call pool.mark_exhausted_and_rotate) and a plain
per-key "rate limit exceeded" 429 (must rotate) against a stub pool routed
through _recover_provider_pool. Verified red by temporarily dropping the
`not _is_overloaded_error` guard: the capacity case fails.
2026-09-17 08:53:54 -07:00
teknium1
6368ec6883 fix(aux): aux billing keywords share the main classifier table; quarantine logs name the real reason; vision fallback skips text-only models
Three auxiliary error-classification defects with one shared cause: the aux
ladder kept its own copies of the classifier's phrase tables and its own
fixed log wording, and they drifted.

- `_PAYMENT_KEYWORDS` is now `error_classifier._BILLING_PATTERNS` plus the
  aux-only phrasings, and `_BILLING_PATTERNS` gains OpenRouter's org cap
  "budget limit exceeded". A 403 "Budget limit exceeded (monthly limit)" (or
  "Key limit exceeded") is billing for the main loop AND a payment error for
  the aux ladder, so compression/title/vision fall through to the configured
  fallback instead of raising while the main model is served elsewhere.
- `_mark_provider_unhealthy` takes the caller's reason and level; the skip
  line echoes it. Absent credentials (no OPENROUTER_API_KEY, no Nous auth)
  are DEBUG with the real reason; the confirmed-402 rungs keep the WARNING
  payment wording. Local-only users no longer read "payment / credit error"
  for providers they never configured.
- `_try_main_agent_model_fallback` applies the auto-route's vision capability
  gate, so a transient 429 on the pinned vision lane no longer hands an image
  to a text-only main model (guaranteed 400 "content.type is invalid").
  `_recover_provider_pool` no longer benches a pool entry on an
  upstream-capacity 429 ("at capacity", same `_OVERLOADED_PATTERNS` table).

Fixes #107166, #64144, #108349.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: awizemann <319078+awizemann@users.noreply.github.com>
Co-authored-by: kokhlo <47825603+kokhlo@users.noreply.github.com>
2026-09-17 08:53:54 -07:00
teknium1
0a3b74379f fix(tools): core-tool drop warns only after the probe had admitted it
The live pass on a stock home counted 13 new WARNINGs per process: every
unconfigured core-gated tool (browser, image_gen, HA) warned on its first
probe. The actionable case is a core tool that was available earlier in
this process and then got dropped; never-configured stays INFO.
2026-09-17 08:53:48 -07:00
teknium1
4e41ab0979 fix(tools): warn once when a check_fn drops a core tool; reject a mistyped catalog_provider
Issue #112649 atom 4: a core (`_HERMES_CORE_TOOLS`) tool whose check_fn returns
False leaves neither the schema nor the tool_search catalog (it is not
deferrable), so the model's "no such tool" is accurate and nothing in the log
points at the probe. ae5666f7fc made the False verdict INFO for optional,
unconfigured toolsets — that stays; the core-tool drop is now a WARNING that
names the dropped tool(s), once per probe per process (the 30 s TTL re-probe
would otherwise re-warn every turn), reset by invalidate_check_fn_cache().

Review minor: `catalog_provider: deepsek` was accepted silently and the typo
leaked into ModelInfo.provider_id. An alias that is neither a Hermes provider
id nor a models.dev id (loaded catalog, no network) now warns once, mirroring
the unknown-key warning, and the row keeps its own slug.
2026-09-17 08:53:48 -07:00
teknium1
d1c6786d1e fix(models): custom providers can inherit a vendor catalog; unknown models get no synthesized output cap
A metadata-only model_overrides entry on a custom provider (only context_window set) took the
unknown-model branch of get_model_capabilities and reported max_output_tokens=8192 from the
_UNKNOWN_MODEL_BASE template; /model info showed "Max output: 8,192 tokens" for a model nobody
knows. The output limit is now left unknown (ModelCapabilities.max_output_tokens=None,
ModelInfo.max_output=0), matching the fail-open treatment #113146 gave supports_vision and
supports_reasoning. Consumers already handle a missing value: the dashboard only renders a
positive limit and the /model card only prints a truthy max_output.

There was also no way for a custom provider that resells a catalogued vendor's models (a
gateway/proxy) to reach _BUILTIN_MODEL_METADATA or the models.dev catalog, which is what forced
the metadata-only override in the first place. providers.<name>.catalog_provider (also honoured
on legacy custom_providers[] rows) names that vendor; models_dev resolves the catalog id through
it for capabilities, context, model info and provider info. It affects metadata lookups only,
never routing or credentials, and an explicit model_overrides entry still wins.

Closes the remaining atoms of #112649 (docs: providers page).
2026-09-17 08:53:48 -07:00
teknium1
03e17ba7e2 fix(agent): route a planning-monologue reasoning-only stop through the stall guard, not a fake completion
A reasoning-only clean stop with tools offered and zero tool calls whose
reasoning ENDS on a first-person plan ("Let me batch the terminal calls and
run them in parallel.", "I need to check the log.") was promoted to the
final answer: the turn reported "complete" while the model had done nothing,
masking a stalled model and aborting multi-step tool loops (#111761,
KushBeaverTTV / Iman-Sharif). The existing stall-guard tail detector only
matched "let me now" / "I'll now" / "next, I" and caps at 400 chars, so these
verbatim tails sailed through.

Add promoted_reasoning_announces_action(): a broader, tail-only first-person
plan detector applied ONLY on the promoted-reasoning path (content empty,
tools offered, no tool call). It feeds the SAME stall-guard continuation
(interim row with api_content sidecar + nudge) under the SAME 2-continuation
cap, so a model that never acts still ends after two nudges. Visible-content
replies keep the narrow trailing_continue_intent() shape, and reasoning that
states an answer ("The answer is 42.", plan-then-answer) still promotes on
the first call.

Part of #111761
2026-09-17 08:53:20 -07:00
teknium1
194695b58a test(agent): lock the reasoning-only stop row shape on the stall-guard and prefill paths
#113188 (merge 578bff8) keeps a reasoning-only clean stop out of the assistant
row's ``content`` (promoted text lives in the ``api_content`` sidecar) and the
merge commit stamped the same sidecar on the stall-guard / codex-ack interim row
and made the gateway history rebuild replay sidecar-only rows. The interim-row
half landed without a test, and the thinking-prefill stub named in #111761 had
no guard that its ``content`` stays empty.

- stall-guard interim row: promoted reasoning that trips
  ``trailing_continue_intent`` must persist ``content=""`` + ``api_content`` and
  the continuation request must carry the text as the assistant turn (red on
  578bff8^, where the interim row had no sidecar).
- thinking-prefill stub: the row appended by ``recover_empty_response`` keeps
  the reasoning in its reasoning fields only, and no transcript row ever stores
  the chain-of-thought as ordinary content (red when the stub copies reasoning
  into ``content``).

Part of #111761
2026-09-17 08:53:20 -07:00
teknium1
d41aec0beb fix(aux): named custom anthropic_messages entries send their extra_headers too
`_resolve_named_custom_branch` lifted a `providers:` entry's `extra_headers`
onto both OpenAI-wire clients but built the anthropic_messages client without
them, so a gateway that needs a header on an Anthropic-style relay still
dropped it on aux calls. `with_options(default_headers=...)` merges the entry
headers onto the beta / credential-Omit headers the builder installs, so the
one-credential invariant is unchanged.

Review follow-up on #113969; the existing anthropic_messages named-custom test
now asserts the header (red on the previous head).
2026-09-17 08:53:15 -07:00
teknium1
b1db026f3a fix(aux): keep key_cmd bearer and extra_headers on async aux calls to named custom providers
An `auxiliary.<task>.provider` pinned to a bare-named `providers:` entry with
`key_cmd` sent every async aux request (async_call_llm, async vision, the
async fallback ladder) with NO Authorization header: the OpenAI SDK parks a
callable api_key in `_api_key_provider` and leaves `.api_key` == "", and
`_to_async_client` rebuilt AsyncOpenAI from that empty snapshot alone, so the
SDK's auth_headers came back empty. The entry's `extra_headers` were never
lifted onto the aux client at all (sync or async), unlike the main runtime.

- agent/auxiliary_async_rebuild.py: carry the sync client's token provider
  (wrapped for await, run off-loop) and its configured default_headers onto
  the async twin.
- _to_async_client uses both; covers the primary async route and every
  fallback-candidate rebuild.
- _resolve_named_custom_branch lifts the entry's extra_headers onto the
  OpenAI-wire client (both call sites).

key_env (static string) was not affected on main; key_cmd was.

Fixes #109595
Salvages #109626

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-17 08:53:15 -07:00
teknium1
cdc4fc4c8f fix(kanban): only a handoff to a different profile lifts the active_pr guard
The handoff exemption treated ANY `assigned` event after the PR comment as a
handoff. A no-op same-profile assign (`hermes kanban assign <id> dev`,
dashboard PATCH /tasks/{id}, `reassign --reclaim`), an unassign
(`assignee: null`), and the dispatcher's own `kanban.default_assignee`
fill-in all record an `assigned` event without changing the owner — and each
lifted `active_pr` for the very implementer that opened the PR, re-spawning
it against its own PR. That is the duplicate-work protection #111910 says
must be preserved.

`assign_task` now records `from` in the payload and the guard counts an
`assigned` event as a handoff only when it moves the card to a different,
non-null profile and was not written by the default-assignee pass. Events
without `from` (pre-existing rows) are not trusted — fail closed.

Docs sentence qualified accordingly.
2026-09-17 08:52:53 -07:00
teknium1
2e805005ea fix(kanban): active_pr respawn guard no longer holds a card handed to a closer or back to its implementer
A ready card whose recent comment links a GitHub PR is held by the
``active_pr`` respawn guard for every assignee. That stopped exactly the
spawns the PR evidence calls for: an operator reassigning the card to a
closer/recovery profile, and a reviewer's changes-requested verdict
routing it back to the implementer to fix that PR — the card sat in
``ready`` as ``respawn_guarded=active_pr`` with zero workers (#111910).

``check_respawn_guard`` now lifts ``active_pr`` when a handoff event
(``assigned``, ``changes_requested``, ``review_reopened``) was recorded
strictly after the newest PR comment: the handoff names the profile that
must work on that PR, so it is not a duplicate implementation. Duplicate
protection is kept: no handoff (a crash/reclaim is not one) leaves the
opener guarded, a newer PR comment after the handoff guards again,
same-second ties stay guarded, and ``recent_success`` is untouched. The
review-lane exemption from a235d1917e stands unchanged.

Docs: the kanban respawn-guard section names both supported routes
(``request-review`` and ``assign``).
2026-09-17 08:52:53 -07:00
teknium1
3ea0089efb fix(process): let the yield-to-background kill test signal the reparented orphan
tests/tools/test_terminal_yield_to_background.py went red after the parent-first
teardown (f11499fdb3a): kill_process() returned {'status': 'error'} instead of
'killed' for a live backgrounded `bash -c 'echo started; sleep 60; echo done'`.

Root cause is the test harness, not the registry. _terminate_host_pid() now
SIGTERMs the shell before its descendants; a non-interactive bash dies at once
and `sleep 60` is reparented to init before the registry reaps it from the
descendant snapshot. tests/conftest.py's live-system guard allowlists a PID by
walking its parent chain up to the test process, so the reparented orphan looks
foreign and the guard raises RuntimeError inside psutil.Process.terminate().
That is not a psutil/OSError, so `suppress(gone)` lets it escape and
kill_process() wraps it as 'error'. Outside pytest the same path returns
'killed' and the orphan is gone (verified with a standalone probe).

Opt the test into real signal delivery with the established
@pytest.mark.live_system_guard_bypass (it already spawns and kills a real
process) and tighten it: capture the shell's descendants before the kill and
assert none survives, so the anti-orphan snapshot cleanup is covered end to end
instead of being hidden behind the harness. Assertions are unchanged otherwise.
2026-09-17 08:52:27 -07:00
teknium1
019a34346c test: prove a self-reaping supervisor's child is never signalled by the registry
Adds the real-process invariant behind the parent-first teardown salvaged from
#111603 (#111598): a bash supervisor that reaps its own child on SIGTERM must
receive the only SIGTERM and exit 0, with the child never signalled by the
registry. Red on origin/main (the child logged a registry TERM before the parent
and the parent's own kill found no such process); green on the salvaged head.

Documents the teardown order in tools/AGENTS.md, including why scope teardown
(_stop_systemd_unit) never precedes the PID signal for a live parent — the
issue's cgroup-split concern (Electron migrating only the browser PID into its
own scope) cannot make the registry stop the scope first.
2026-09-17 08:52:27 -07:00
KoNit-K
039165a3d4 fix(process): grace host PID tree teardown 2026-09-17 08:52:27 -07:00
teknium1
7d57fbfca9 fix: digest only the peeked entry's key in the aux client cache hint
Review follow-up on #113790. _pool_credential_digest hashed every pool
entry, so rotating a sibling account's token churned the cached client of
the selected, unchanged entry (multi-account pools rebuilt every aux client
on every account's refresh). Digest only the peeked entry's key when peek()
names one; keep the all-entries digest solely for the '<provider>::'
cooldown case so a rotation during a per-model cooldown still changes the
key (#113022 point 5).

Also migrate the six remaining literal 7-tuple cache keys in
test_async_httpx_del_neuter.py to _client_cache_key() so no test encodes
the positional key layout, as the PR body claims.
2026-09-17 08:51:58 -07:00