Commit Graph

41168 Commits

Author SHA1 Message Date
teknium1
17a137d8a3 fix(tui-gateway): a SIGTERM-ignoring command no longer survives the gateway's SIGTERM exit
The SIGTERM handler arms a 1s os._exit timer, then runs _shutdown_sessions: a flush of up to
5s, then _stop_turns_before_exit, whose kill was the graceful TERM, wait 1s, KILL. A command
that ignores SIGTERM was still alive when the timer fired, and os._exit left it reparented to
init (live: `trap '' TERM; sleep 3600` survived a SIGTERM to `python -m tui_gateway.entry`).

- kill_live_foreground_processes(now=True): SIGKILL each in-flight foreground tree at once,
  no TERM grace, no wait (BaseEnvironment._force_kill_process; LocalEnvironment kills the
  recorded process group, never our own).
- The grace timer's exit (entry._hard_exit) runs it before os._exit.
- _stop_turns_before_exit SIGKILLs whatever is still alive halfway through its settle budget
  (it ignored the interrupt's TERM), so the tool call still ends with a result the teardown
  persists instead of a dangling tool_call in state.db.
- The other hard exits that skip cleanup do the same before os._exit: the serve parent-death
  watchdog, the CLI exit watchdog, the kanban worker's SIGTERM path, and the messaging
  gateway's shutdown and loop-liveness watchdogs.
- Deflake test_shutdown_mid_tool_kills_the_command_and_keeps_its_result: the 0.5s settle
  budget was too tight under -n 40 (1 red in 9 runs); the join returns when the turn ends.
2026-09-23 17:09:54 -07:00
teknium1
537ce77f52 fix(tui-gateway): exiting mid-tool no longer orphans the foreground command's process tree
Foreground terminal commands run in their own session (start_new_session) so an
interrupt can kill the whole tree, which also puts them outside the host's process
group. When the tui_gateway left mid-command (client closed stdin, or SIGTERM) nothing
killed them: _shutdown_sessions closed the agents, the SIGTERM path hard-exits after a
1s grace, and the `bash -c` + child tree survived, reparented to init.

- tools/environments/base.py: execute() records every in-flight foreground command;
  kill_live_foreground_processes() kills their trees through the backend's own
  _kill_process (the same kill an interrupt uses).
- cleanup_all_environments() (the exit funnel of the CLI, one-shot, messaging gateway
  and terminal_tool's atexit, so `hermes serve` too) kills them first.
- tui_gateway _shutdown_sessions (EOF atexit + SIGTERM handler) and the serve
  SIGTERM/SIGINT exit-flush handler interrupt running turns, wait up to 0.5s for them
  to settle so the tool call ends with a result the final persist records (no dangling
  tool_call in state.db), then kill any foreground command still alive.
- ComputeHost.close() kills them too: every caller os._exit()s right after.

Covers the case of b9dac83d366c (Desktop quit: serve SIGTERM handler) on every host.
2026-09-23 17:09:54 -07:00
teknium1
ffc44bbaf3 fix(model): prefer the current provider's alias on implicit switches too
Only the --provider path handed user_providers/custom_providers to
resolve_alias, so the provider-identity comparison behind the
current-provider preference could not see legacy custom_providers
entries on the implicit path (/model <id>) or the authenticated-provider
fallback. On custom:corp-llm, /model shared-model picked another
provider's alias for the same model id and switched to its base_url.

Pass both provider maps on every resolve_alias call site. The two
existing resolve_alias stubs in test_ollama_cloud_auth.py accept the
extra positional args; their assertions are unchanged.
2026-09-23 17:08:21 -07:00
teknium1
79d012bd25 fix(model): match alias ownership on the resolved provider id
The --provider ownership check compared normalize_provider() spellings, but a
legacy custom_providers entry resolves to id "custom:<name>". An alias with
"provider: corp-llm" under "--provider corp-llm" was therefore dropped and the
switch fell back to the provider's own base_url/key instead of the alias's.

Compare resolve_provider_full(...).id on both sides (normalize_provider when
unresolvable), in the ownership check and in resolve_alias's reverse-lookup
preference loop, which now receives the user/custom provider config.
2026-09-23 17:08:21 -07:00
teknium1
ed3e7203ee chore: map xiaoyaner0201 contributor email 2026-09-23 17:08:21 -07:00
teknium1
af80401f4e test(model-switch): pin --provider ownership of direct aliases
Two invariants against a real config.yaml: an alias bound to another
provider's endpoint never outranks an explicit --provider (host and key stay
on the named provider), and when several aliases share a model id the one
owned by the named provider wins regardless of mapping order.

Found by the C11 routing truth table E2E (lands separately with the E2E suites).
2026-09-23 17:08:21 -07:00
千乘妍 (Xiaoyaner)
1ee38c8090 fix(model): normalize direct-alias provider ownership
Normalize both provider labels at direct-alias ownership comparisons so alias
spellings and case variants match canonical target IDs without weakening
cross-provider rejection.

Maintainer finding: https://github.com/NousResearch/hermes-agent/pull/75262#discussion_r3688616856
2026-09-23 17:08:21 -07:00
千乘妍 (Xiaoyaner)
60f2bcf4ab fix(model-switch): scope direct-alias base_url to the target provider
An explicit `/model <id> --provider X` switch could report target_provider=X
while base_url still pointed at the previous provider's endpoint: resolve_alias
finds a direct alias by reverse model-ID lookup, and the base_url override
applied that alias unconditionally even when the alias belonged to a different
provider. Requests would then be sent to the old provider's URL under the new
provider's identity.

Two narrow changes:

- resolve_alias's reverse lookup now prefers an alias whose provider matches
  current_provider, falling back to first-match only when none does. Insertion
  order is not a routing decision.
- The direct-alias base_url override now ignores (and stops reporting) an alias
  whose provider differs from target_provider.

Exact alias-name lookup stays provider-agnostic, and the first-match fallback is
preserved, both pinned by tests.

[salvaged onto the refactored switch pipeline: the ownership guard now sits in
_route_explicit_provider, where resolved_alias is settled, so both the runtime
resolve (explicit_base_url) and the direct-alias override see the filtered alias]
2026-09-23 17:08:21 -07:00
teknium1
027809f88e fix(agent): don't restore a foreign Session ID or pin an empty workspace snapshot
- _stored_prompt_matches_runtime: with the Session ID trailer on
  (--pass-session-id / HERMES_TUI_PASS_SESSION_ID), a stored prompt whose
  Session ID line is not this session's is a mismatch. A /branch child now
  copies its parent's prompt bytes, and without this it told the model the
  parent's id. Gated on the flag so a "Session ID:" line in project text
  cannot force a rebuild every turn when the trailer is off.
- _seed_workspace_pin: adopt only a real workspace block. A stored prompt
  that names this cwd but carries no block (built on a messaging surface,
  resumed on the CLI in the same repo) pinned "no workspace" and dropped the
  git snapshot for the rest of the session; now the next build captures one.
2026-09-23 17:06:48 -07:00
teknium1
2ec129528d fix(sessions): a /branch child row carries the parent's system prompt (CLI + TUI/Desktop)
The branch copies the parent transcript byte-for-byte so its first turn can hit the warm
prefix cache, but the child row was created without a system prompt. The child's first
build then found nothing to restore and re-probed the workspace, rewriting the prompt at
byte 0 for any repo that moved since the parent's session start — the same rewrite this PR
removes for /compress and resume, on the one session transition it did not reach. It also
logged the "stored system prompt is null; investigate update_system_prompt" WARNING for
every branch.

Both branch writers now pass the parent's prompt into create_session: the CLI prefers the
running agent's cached bytes (what this process sends) and falls back to the parent row;
_persist_branch (TUI/Desktop, seeded and lazy) reads the parent row. With the row in place
the seeding in agent/system_prompt.py replays the snapshot on the branch's later rebuilds.

Tests: one per surface, red on the PR head (child["system_prompt"] is None), green here.
2026-09-23 17:06:48 -07:00
teknium1
921ab7a163 fix(agent): keep the session-start workspace snapshot across prompt rebuilds
The workspace git snapshot is pinned per session so a rebuild (compaction,
/compress) replays it instead of re-probing a repo that moved. Two gaps let
the rebuild re-probe anyway and rewrite the system prompt mid-session:

- The pin was keyed by resolve_context_cwd(), which is None when no cwd is
  bound (CLI launch dir) and the path once the TUI /compress binds the
  session cwd: the same directory under two keys. Key by the directory the
  probe actually inspects (resolve_context_cwd() or resolve_agent_cwd()).
- An agent that did not build the session's prompt (resumed, a fresh
  gateway/TUI agent whose first act is /compress, a kill -9 restart) had
  no pin, so its first rebuild probed git now instead of replaying the
  session-start bytes. Seed the pin from the prompt the session already
  sends (cached copy, else its persisted row), only when that prompt names
  this cwd and its snapshot's Root covers it.

reset_session_state still drops the pin, so /new, /resume and /branch
re-snapshot at their own session start.
2026-09-23 17:06:48 -07:00
teknium1
82c5afb19d fix(state): stop a waiting FTS detach once the file is quarantined
The FTS fail-open detach now waits up to the caller's write budget (20 s /
60 s) for the write lock, so the one-time quarantine check before the loop
left a long window: a sibling that quarantined the file meanwhile still got
its triggers dropped and the stale breadcrumb committed on the quarantined
handle. Re-check the handle flag and the process-wide storage latch at the
top of every attempt, via the same _raise_if_db_corrupt(storage=True) that
_execute_write runs per attempt.

Classify the retryable lock error with is_sqlite_lock_error (result code
first) instead of a locked/busy substring match, matching #120488.
2026-09-23 16:59:09 -07:00
teknium1
9ae29873f4 fix(state): a busy write lock no longer loses the write that hit a corrupt FTS index
When a canonical write trips a corrupt FTS index, SessionDB detaches the derived
indexes (breadcrumb + trigger drop) and retries the write. The detach ran one
BEGIN IMMEDIATE on the writer connection, whose busy timeout is only 1 s, and gave
up on "database is locked" — so the canonical write escaped as "database disk
image is malformed". The usual lock holder is a sibling writer (gateway + TUI)
detaching the same index, so under load the second writer's turn was lost.

The detach now waits out lock contention on the caller's write budget with the
same jittered retry as _execute_write (default _WRITE_PATIENCE_S for the search
fail-open callers).

Repro: a second process takes BEGIN IMMEDIATE the instant the corruption error
surfaces and holds it 2.5 s. Base: append raises after 1.02 s (3/3). Fixed: the
row lands after the holder releases, FTS detached (3/3). Found by the E2E sqlite
torture chamber (fts_corruption_fail_open) at load ~200.
2026-09-23 16:59:09 -07:00
teknium1
9c31215ff5 fix(desktop): a later failed turn no longer swallows a stranded member's late reply
Re-review follow-up (N1, N2) for the group room failed-turn fix.

- Stranded harvest: the retained error only counts as the stranded turn's
  own when no user row follows the turn's prompt. When a later turn in the
  same session (a gateway/CLI message, a later room turn) failed, its error
  is not the stranded turn's, and the real late reply now posts instead of
  being dropped and reported as a failure.
- Live poll: the pre-submit comparison keys the retained error on
  `turn_started_at` plus the `inflight` snapshot. The snapshot alone carries
  no start time, so a failure identical to the previous turn's read as the
  leftover and the pre-tool text was posted as the reply.
2026-09-23 16:58:39 -07:00
teknium1
9d6e4e72a4 fix(desktop): a member that fails after pre-tool text is reported, not posted as its reply
Independent-review follow-up for the group room failed-turn fix.

- Group room poll: a retained error newer than the pre-submit snapshot
  (turn start replaces it with a fresh started_at) is this turn's failure
  and wins over transcript text. The core closer writes no failed-turn row
  behind a tool row, so "said X, called a tool, provider 401" left X as the
  newest assistant row and the room posted it, dropped the 401 and re-drove
  the member to the round cap (live repro on the PR head).
- Both pickers end the turn at a failed_turn row instead of scanning past
  it; with the retained error gone (backend restart) the row's notice is
  reported through the failed path instead of a silent pass / dropped
  stranded marker. The stranded harvest treats a retained error as the
  stranded turn's own the same way.
- REST cold-load/paging (/api/sessions/{id}/messages, /messages/around)
  type legacy untyped notice rows like session.resume does, via one
  read-side helper in agent/turn_failure_copy.py.
- Gateway closer: the fresh-session closure test now asserts the row's
  display_kind through a real SessionDB round-trip (red without the
  run_turn.py stamp).
- E2E: mock trigger that says text + calls a tool, then 401s; spec asserts
  the room reports the error and never posts that text.
2026-09-23 16:58:39 -07:00
teknium1
224e7b673e fix(desktop): a failed turn's "request not processed" line is a Hermes notice, not the model speaking
The main chat painted the failed-turn boundary row as an assistant message:
"Your request was not processed. Send it again..." sat above the provider
error card in the model's own voice, live and after a reload. Rows typed
display_kind="failed_turn" now hydrate as a system row, like the other
Hermes timeline markers.

Live: a 1:1 chat with the mock provider answering 401 rendered the notice in
aui_assistant-message-root before, aui_system-message-root after (both live
and after a cold reload).
2026-09-23 16:58:39 -07:00
teknium1
502031a4a3 test(desktop-e2e): restore the group approval-click and member-failure specs
Both were deleted by #120071 as failing on main. Triage: group-member-backend-failure
was a real regression (fixed in the previous commit); group-approval-click-submits
only failed under a leaked HERMES_YOLO_MODE (harness fixed two commits up) and passes
on a clean env. Restored verbatim from d9ca819cc4.
2026-09-23 16:58:39 -07:00
teknium1
c3d4007c00 test(desktop-e2e): the sandboxed app never inherits the caller's Hermes runtime env
A spec run from inside an agent terminal passed HERMES_YOLO_MODE,
HERMES_INTERACTIVE, HERMES_SESSION_ID... straight into the sandboxed backend.
HERMES_YOLO_MODE alone made group-approval-click-submits fail (the gated
rm -rf ran without a prompt), which read as an approval regression; CI never
has these vars. buildAppEnv now drops inherited HERMES_* except the
harness's own HERMES_DESKTOP_* / HERMES_E2E_* knobs.
2026-09-23 16:58:39 -07:00
teknium1
a29cf60711 fix(desktop): a group member's failed turn is reported, not posted as its reply
Since 8f0322da5b the core loop closes a failed turn with a Hermes-authored
assistant row ("Your request was not processed..."). The group room's turn
poll took the newest assistant row as the member's answer, so a provider 401
posted that line as the bot speaking, recorded the member as replied (no
failure row, no roster badge, no error reason from #117366), re-drove it, and
the drive ended at the round cap.

pickGroupTurnReply / pickStrandedGroupTurnReply, and the external-write
mirror, now skip rows typed display_kind="failed_turn" (stamped by the
previous commit), so the retained `inflight` error surfaces through the
existing failed-turn path again. No copy of the English text on this side.

Live: the deleted group-member-backend-failure.spec.ts is red on origin/main
(3/3, "turn stopped at the round/message cap"), green with this change
("Programmer hit an error — HTTP 401: ..."). Backend-only A/B with the main
renderer: green at 8f0322da5b8^, red at 8f0322da5b.
2026-09-23 16:58:39 -07:00
teknium1
3dbb417809 fix(agent): type the failed-turn boundary row so clients never read it as the model's reply
8f0322da5b closes a failed turn with a Hermes-authored assistant row
(FAILED_TURN_NOTICE / PARTIAL_FAILED_TURN_NOTICE). It carried no marker, so
the only way a client could tell it from a real answer was matching the
English copy.

Both writers (agent/conversation_loop.py::_close_durable_failed_turn and
gateway/run_turn.py::_hmwa_close_failed_turn) now stamp
display_kind="failed_turn" (agent/turn_failure_copy.py::FAILED_TURN_DISPLAY_KIND).
display_kind is a DB/display column already stripped from every provider
request (agent/turn_context.py), so the wire bytes and prompt cache are
unchanged; the ACP loopback test asserts the replayed row is exactly
{"role": "assistant", "content": FAILED_TURN_NOTICE}.

session.resume (tui_gateway/session_history.py::_legacy_display_kind) also
types untyped rows already on disk from the last five days, matching the
Python constants in-process, so clients key on the type alone.
2026-09-23 16:58:39 -07:00
teknium1
bd14eb1c93 fix(terminal): persistent Docker keys a routed profile's cron work to its profile container
Under persistent Docker, profile B's session-bound work keys profile:B but its
session-less (cron) work keyed home:<path>, so the same profile ran two
long-lived containers. The routed no-session branch now returns the same
profile key as branch 3 when Docker is persistent (profile-scoped); other
backends keep the per-home key.
2026-09-23 16:58:12 -07:00
teknium1
02fa29943a test(gateway): a first gateway.run import under a routed override bridges the launch home
Invariant for the import-time config bridge: a multiplexed backend's first import of
gateway.run can happen inside a routed profile's session (hermes serve imports it lazily
from an agent build), and the bridge must still write the launch home's agent.max_turns and
terminal.cwd into the process env. Red on the previous gateway.run (HERMES_MAX_ITERATIONS
came from the routed profile), green with get_process_hermes_home().
2026-09-23 16:58:12 -07:00
teknium1
12ca697015 fix(gateway): anchor gateway.run's import-time home and config bridge on the process home
gateway.run bridges config.yaml into os.environ at import, keyed on
get_hermes_home(). The Desktop backend (hermes serve) first imports it lazily
from a session's agent build (tui_gateway.agent_callbacks._wire_callbacks),
under that session's routed profile override, so whichever secondary profile
built first latched its terminal.* (TERMINAL_CWD, backend...) and bridged
settings into the launch process env for every later launch-profile turn and
cron job.

Found by tests/e2e/core/tenancy/test_two_tenant_desktop_backend.py (C7 canary):
the default profile's cron env snapshot carried alpha's/beta's TERMINAL_CWD.
Use get_process_hermes_home(); identical for a standalone gateway.
2026-09-23 16:58:12 -07:00
teknium1
551497dcbf fix(terminal): session-less work for a routed profile keys its own terminal environment
A multiplexed host runs every profile's cron jobs without a session key, and
_resolve_container_task_id collapsed all of them onto the shared "default"
environment. Profile B's cron tool calls therefore reused the LocalEnvironment
the launch profile's job created: B's terminal subprocess saw the launch
profile's .env residue and bridged TERMINAL_* (TERMINAL_CWD, backend policy).

Found by tests/e2e/core/tenancy/test_two_tenant_gateway.py (C7 canary): alpha's
cron env snapshot carried default's TENANT_MARKER and TERMINAL_CWD. Session-less
work under a routed home override now keys home:<realpath>; the launch profile
and single-profile processes keep "default". The unit test that pinned the
shared key for a routed profile is updated.
2026-09-23 16:58:12 -07:00
brooklyn!
c3c03970df chore: map orangercat contributor email 2026-09-23 18:52:49 -05:00
brooklyn!
0a0e6334c6 fix(desktop): decode UTF-8 child output explicitly on the serve/agent path (#83851)
The Desktop serve backend polls /api/fs/default-cwd, whose git-branch probe
captured git output with text=True and no encoding. subprocess then decodes
with locale.getencoding() - cp936 on zh-CN Windows - so a UTF-8 branch name or
git's localized stderr raised UnicodeDecodeError inside communicate()'s
_readerthread on every poll ([gateway-crash] ... 'gbk' codec can't decode).

Pin encoding='utf-8', errors='replace' on the remaining in-process captures
whose children emit UTF-8: the git branch probe, the self-repo guard's git
alias read, the Bot Mode DM transport (a Hermes CLI child, stdio forced to
UTF-8; builds on the salvaged errors='replace'), the agent-browser npx probe,
the lightpanda help probe, and the cua-driver stderr drain thread.

Fixes #83851
2026-09-23 18:52:49 -05:00
chaochen
526a0015fb fix(decode): tolerate non-UTF-8 child output in two remaining capture paths
Same class as the cron script runner: text=True decodes with the locale codec under strict error handling, so one undecodable byte in a child's output raises UnicodeDecodeError inside subprocess.run. That is a ValueError - neither except (subprocess.SubprocessError, OSError) nor except TimeoutExpired catches it - so it escaped into user-visible paths instead of the intended error handling.

tools/bot_mode_dm._run_local_turn: a transport that exits 0 but prints a byte the locale cannot decode crashed the delivery instead of re-emitting the transport's streams (stdout is the reply text the completion notification carries back).

agent.command_token_source._mint: a key_cmd printing a non-UTF-8 byte killed token minting with a traceback instead of the documented CommandTokenError path, so provider auth failed in a way the caller could not report.

Both decode with errors=replace now: the exit code and the surrounding checks still decide what happens, and damaged bytes appear as U+FFFD in the text we carry.
2026-09-23 18:52:49 -05:00
Austin Pickett
d94769da64 fix(update): a Desktop-only host settles an inventory-less restart obligation (#120740)
* refactor(update): one predicate for serve rows outside the gateway matrix

The inventory branch of `_marker_only_restart_obsolete` inlined the rule for which
serve/dashboard rows the gateway matrix neither covers nor needs to (supervisor-owned,
or a manual serve handed to its own reminder). The inventory-less branch needs the same
rule for #118742, so it moves to `update_cmd_fleet_gatewayless.runtime_outside_gateway_evidence`
and both branches will read one definition. No behaviour change.

* fix(update): a Desktop-only host settles an inventory-less restart obligation

A host that runs no gateway (the Desktop app alone) can be left with an inventory-less
fleet-restart obligation: an updater that died before recording its inventory, or the
pre-inventory writer. With no owed set, the live gateway matrix is its only evidence, and
on that host the matrix is empty forever, so `_marker_only_restart_obsolete` never settled
and every later `hermes update` exited 1 with "gateways are still off the checkout code"
(#118742).

An empty fleet alone cannot tell that host from one whose gateway the dying update stopped,
so the inventory-less branch now asks the live host, never a historical receipt:
`host_owes_no_gateway_restart` settles only when no profile's `gateway_state.json` claims a
state other than stopped/startup_failed (a gateway that went away without a clean stop keeps
the obligation) and every live runtime sits outside the gateway matrix
(`runtime_outside_gateway_evidence`, shared with the inventory branch). HEAD must still
contain the pulled SHA (`checkout_contains`, same rule as #119367). Probe failures keep it.

Tests: the scoped-reconciliation matrix now holds its host at "the update stopped a gateway"
so it keeps pinning receipt independence; a new host-evidence matrix covers Desktop-only,
clean stop, carried commit, stopped gateway in a named profile, unclassified and
unidentified serves, a gateway row without fleet identity, and a diverged checkout. Two
manual-serve tests that assumed an empty fleet always stays pending now stub the live host
and assert the manual reminder survives the gateway obligation settling.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>

---------

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-23 18:48:33 -05:00
brooklyn!
5f8067a92b test(e2e): enforce the ACP toolset_restriction parity cell now that #74582 is fixed 2026-09-23 18:45:45 -05:00
hermes-seaeye[bot]
f799fd8578 fmt(js): npm run fix on merge (#120759)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 23:39:41 +00:00
lgy1027
093db2b190 fix(desktop): restore packaged macOS locale markers
Adapt #93931 to the current packaging lifecycle with real packager paths.
Retain the non-blocking restore from 0fbb1bc337539902d1132af4a02412b52e73d6ed,
exclude Chromium gender packs, and leave Windows stamping in afterExtract.

(cherry picked from commit c37a15c2fe5665081e7c9e6bfbe93a4c2f3aadfe)

Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-09-23 18:32:26 -05:00
brooklyn!
cb7b1b32dd fix(desktop): keep English stop defaults, Unicode-safe matching, configurable barge-in (#117801)
Builds on the salvaged voice.stop_phrases matcher:

- /api/config merges DEFAULT_CONFIG, so an untouched install reports
  stop_phrases: ["stop"] instead of omitting the key. Compare against
  /api/config/defaults so that case keeps the built-in English list
  ("goodbye", "never mind", ...) and only a list the user changed replaces
  it. Parse malformed values (mapping, null) as the backend does: default.
- Normalise transcripts and phrases with NFKC and Unicode punctuation
  (\p{P}), so NFD Cyrillic from STT, «guillemets» and CJK full stops match.
- Wire the existing voice.barge_in_threshold_multiplier into the desktop
  barge-in monitor: it scales the quiet-floor multiplier and the playback
  clamp (PLAYBACK_MIN_TRIGGER_LEVEL) by its ratio to the backend default
  (3.0), so stock behaviour is unchanged and a quiet BT HFP headset can
  interrupt with a lower value. No new setting or env var.
2026-09-23 18:30:27 -05:00
aydnOktay
802a0cf6ff fix(desktop): honour voice.stop_phrases in the spoken-stop matcher (#117801)
The desktop voice loop advertised configured stop phrases in the notice but matched only a hardcoded English list, so non-English STT phrases (e.g. отбой) were submitted as prompts. Wire the matcher to the same voice.stop_phrases config the backend uses.
2026-09-23 18:30:27 -05:00
brooklyn!
529d27e5be fix(desktop): read table headers aloud instead of dropping tables silently
Read Aloud removed Markdown tables with no spoken trace (#86602). Speak the
header row in their place ("Model, Price.") and keep skipping the body rows,
so the listener hears that a table is on screen and what it compares
without cell-by-cell data. The header is the reply's own text, so it is
always in the voice's language; a localized "table omitted" notice would
follow the UI locale instead and reintroduce the mismatch when the two
differ. A table whose header cells are all empty stays silent.

Also keep the sentence punctuation after an inline MEDIA: token on both
the desktop and backend paths ("see MEDIA:/x.py. Then" no longer loses
its full stop), per review on #89367.
2026-09-23 18:30:12 -05:00
Ricardo Mendes
25f7bac318 fix(tts): keep file links and code-block placeholders out of speech
- MEDIA:/path tokens (which render as 'Open <hyphen-slug>.xlsx') are stripped
  in both the backend normalizer and desktop sanitizer; a voice cannot say
  them and ElevenLabs-style models loop on the hyphenated slug ('eeeeee').
- Desktop no longer speaks the 'code block omitted' / 'link' placeholders;
  unspeakable tokens are silence, not English words (#86602).
- Line-final colons close to periods before newlines flatten, on both sides,
  so 'here is the list:' followed by a code block is spoken as two sane
  sentences instead of hanging the voice on an open colon-pause.

Regression tests on both sides pin each behavior, including the exact
colon-plus-code-fence reproduction.

Salvaged from #89367 (79db15c1a8). The symbol/dash/euro expansion and the
bare-path "the path" placeholder were not carried: on the desktop they put
English words into non-English voices, which is the #86602 bug class.
2026-09-23 18:30:12 -05:00
brooklyn!
58694737e2 fix(time): repair surrogate-bearing locale zone names at every strftime site
On Windows (fr-FR, es-AR, de-DE reports) the zone name arrives in the ANSI
code page but is decoded under a UTF-8 LC_CTYPE (UTF-8 mode, or Piper/espeak
flipping the process locale mid-run) with surrogateescape. datetime.strftime
splices tzname() in as UTF-8, so "%Z" raised UnicodeEncodeError while the
system prompt was being built and every new/compressed conversation died.

Rework safe_strftime into a small repair: output is untouched for valid text
(the system prompt stays byte-identical), surrogateescape'd bytes decode back
through the ANSI code page ("heure d'été"), anything else degrades to U+FFFD,
and a raising "%Z" is rendered from the repaired tzname(). hermes_time no
longer imports agent.* at module load.

Route the remaining locale-name sites through it: cron quota-hold notice,
auxiliary cooldown notice, cron session titles, session_search dates,
insights, learning graph and billing renew dates. Tests use real datetimes
with a surrogate zone name instead of stubbed strftime.

Fixes #102910

Co-authored-by: Aniruddha Adak <127435065+aniruddhaadak80@users.noreply.github.com>
2026-09-23 18:29:35 -05:00
fangliquan
85cb540e6d fix(time): tolerate Windows locale strftime errors 2026-09-23 18:29:35 -05:00
brooklyn!
b3e2e9fa25 test(desktop): unmounting I18nProvider cancels a pending locale retry
The bounded startup locale retry (#96177's second half) already landed on
main; this pins its cleanup path, which had no coverage: once the provider
unmounts, the scheduled retry is cleared instead of polling /api/config for
a tree nobody renders. Adapted from the unmount case in #96213.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-23 18:28:20 -05:00
brooklyn!
b4bad6fca1 fix(desktop): gate the boot WS probe extension on child liveness, not output
Follow-up to the progress-aware probe (#96177):

- Wait while the spawned child is alive instead of "alive and wrote output
  in the last 30s". The stall this fixes is silent: the backend answers
  HTTP, then holds the GIL importing gateway platform modules and prints
  nothing until the loop recovers (the web_server heartbeat only logs
  "event loop stalled" afterwards). Output recency counted down during
  exactly that window, so a long stall still failed. A live loopback child
  whose listener accepted the TCP connect is busy, not refusing; refusals
  and auth rejections are immediate error/close events, not timeouts.
  Drops the now-unused BackendOutputTail.lastActivityAt/backendMakingProgress.
- One policy home: spawnedBackendProbeOptions() in gateway-ws-probe.ts
  (base 10s, cap 90s = the port-announcement cold-start budget); both the
  primary and pool spawn paths use it. Remote gateways, "Test remote" and
  host-backend attach keep the fixed budget.
- Without keepWaitingWhile the probe is again a single one-shot timer (no
  1s polling during the base budget). Timeout reasons name what ended the
  wait: base budget, backend stopped after the budget, or the cap.
- Tests drive every deadline with fake timers (no wall-clock races) and
  cover the production policy: a 30s stall succeeds while alive, a child
  that exits mid-wait fails at the next check, a live backend that never
  answers fails at the cap.
2026-09-23 18:28:20 -05:00
Finn763
4fc04e5a4b fix(desktop): make boot WS probe wait on backend progress instead of a fixed 10s cap
Windows cold start can hold the backend's GIL for 12-28s while gateway
platform modules import (feishu/lark, weixin, telegram C extensions),
leaving the /api/ws upgrade unanswered well past the probe's fixed 10s
connect budget. The probe then fails, the desktop declares the backend
unhealthy, spawns a second backend, and boot lands in the
'Ignoring stale Hermes backend exit (1)' cascade (#96177).

The probe now supports backend-progress-aware waiting: once the quiet
connect budget passes, a keepWaitingWhile callback is polled and the
deadline only fires when the backend stops reporting progress (hard
capped by maxConnectWaitMs so a hung backend still fails the boot).

- gateway-ws-probe.ts: keepWaitingWhile / progressCheckIntervalMs /
  maxConnectWaitMs options; behavior is unchanged when the callback is
  omitted (remote-gateway probes keep the fixed timeout).
- main.ts: both local-backend boot paths (primary window + pool) pass
  keepWaitingWhile = backendMakingProgress(child, outputTail) with a
  90s cap aligned with the existing port-announcement cold-start budget.
- backend-claim.ts: BackendOutputTail tracks lastActivityAt; new
  backendMakingProgress() signal: alive child + (no output yet, i.e.
  inside the block-buffered import window, or recent output within 30s).
- Regression tests: probe waits past the base budget while progress is
  reported, fails once progress stops, respects the hard cap, fails
  closed on a throwing callback; tail activity + progress-signal units.

Benchmark (simulated 13s import stall, production budgets):
  before: FAIL at 10.0s -> boot cascade     after: OK at 13.8s
  warm start: OK 0.96s                      after: OK 0.96s (no delta)
  dead backend: FAIL 10.0s                  after: FAIL 10.1s (same)
  hung backend (alive, never ready):        after: FAIL 90.8s (hard cap)

Closes #96177

(cherry picked from commit e8dbe9e0a2bfb6853401d55a2bbda02393560de0)
2026-09-23 18:28:20 -05:00
teknium1
8aa1e450b8 fix(anthropic): keep OAuth rotation on accepted native proxies
The anthropic.com-only refresh guard stopped token rotation on hosts the
resolver itself accepts as native Anthropic (*.claude.com, /anthropic
proxies), which already receive the Anthropic token at startup. Long
sessions there hit 401 and the recovery refresh was refused too.

Official hosts (anthropic.com, claude.com) rotate as before; any other
host rotates only when the current key is already an Anthropic
credential (sk-ant- or OAuth), so the #17829 case (custom key swapped
for ANTHROPIC_API_KEY / the OAuth token) stays blocked. Azure keeps its
static-key exclusion. Also corrects the _resolve_openrouter_runtime
docstring, which still claimed OPENAI_BASE_URL is never consulted.
2026-09-23 16:27:37 -07:00
teknium1
fd1f47d8f4 chore: map limu.mzm@bytedance.com to leemove for attribution 2026-09-23 16:27:37 -07:00
teknium1
da031c31c1 fix(runtime_provider): don't send an OPENAI_BASE_URL-bound OPENAI_API_KEY to OpenRouter
The OpenRouter key ladder falls back to OPENAI_API_KEY (legacy home of an
OpenRouter key). When OPENAI_BASE_URL binds that key to another host, a
bare `provider: custom` with no endpoint fell through to openrouter.ai and
shipped the OpenAI-proxy key there instead of failing fast. Use it for
OpenRouter only when OPENAI_BASE_URL is unset or names the same host.

Hosts compare through utils.base_url_hostname, so a scheme-less
OPENAI_BASE_URL (proxy.corp:8080/v1) still counts as binding the key.

Found by the C11 routing truth table (tests/e2e/core/tenancy).
2026-09-23 16:27:37 -07:00
teknium1
a4d8f173a9 fix(anthropic): never refresh Anthropic credentials onto a foreign host
_try_refresh_anthropic_client_credentials runs before every request and on
401 for provider 'anthropic'. A URL-bearing model alias under the built-in
'anthropic' label (#28660) points that provider at a foreign host with no
key; the pre-request refresh then swapped in ANTHROPIC_API_KEY / the OAuth
token and sent it to that host.

The third-party guard above matched "anthropic.com" as a substring of the
URL, so a host like 127.0.0.1:8080/anthropic.com still passed as native and
got the refresh. Refresh only when the endpoint's hostname is Anthropic's own
(or no endpoint is set).

Found by the C11 routing truth table (tests/e2e/core/tenancy): a /model
switch to such an alias carried the vendor key to the alias host.
2026-09-23 16:27:37 -07:00
leemove
96ff42ff66 fix(anthropic): skip credential refresh for third-party endpoints
With provider 'anthropic' pointed at a third-party Anthropic-compatible
endpoint, _try_refresh_anthropic_client_credentials only skipped Azure and
otherwise re-resolved native Anthropic credentials (ANTHROPIC_API_KEY, the
stored OAuth token) and rebuilt the client with them, so the endpoint got
a credential it was never given. Skip the refresh for any endpoint
build_anthropic_client already classifies as third-party.

Ported from run_agent.py onto agent/client_lifecycle.py, where the method
lives now; the separate Azure test is dropped (Azure is one such endpoint).

Salvaged from #17829.
2026-09-23 16:27:37 -07:00
teknium1
94c0e49307 docs: scope the ACP toolset parity wording
ACP resolves toolsets like the gateway for the same platform config,
which includes plugin toolsets. An empty platform_toolsets.acp list still
adds enabled plugin toolsets, so the docs point at agent.disabled_toolsets
for removing them.
2026-09-23 16:27:05 -07:00
teknium1
316d6fb8f7 fix(acp): admit config MCP servers by platform_toolsets.acp like the gateway
A fresh ACP agent appended mcp-<server> for every enabled config MCP server
unconditionally, so a platform_toolsets.acp allowlist of server names and the
no_mcp sentinel were ignored on ACP while the gateway honoured both.

The MCP half now comes from the same _get_platform_tools(config, "acp") call
as the base toolsets: its server names (default every enabled server, a listed
allowlist, or none for no_mcp) are keyed as mcp-<server>. Editor-provided
session/new servers are unchanged.

Docs: `hermes tools` has no ACP platform entry, so drop the claim that it
configures platform_toolsets.acp; document the MCP rules with a config example.
2026-09-23 16:27:05 -07:00
teknium1
cccea211a9 chore: map tymrabchuk@gmail.com to 5uck1ess 2026-09-23 16:27:05 -07:00
teknium1
5fb5e743b3 fix(acp): resolve ACP toolsets through the shared platform resolver
A fresh ACP agent hardcoded enabled_toolsets=["hermes-acp"], so
platform_toolsets.acp never narrowed the editor tool surface, unlike the
gateway, cron and api_server which all resolve via
hermes_cli.tools_config._get_platform_tools. Resolve the ACP base the same
way (ACP keeps appending its own mcp-<server> entries), and treat only None,
not an explicit empty list, as "use the hermes-acp default" in the /tools
and MCP-refresh rebuilds so a deny-all list cannot re-widen mid-session.

With the unconfigured default the resolved tool definitions are
byte-identical to the hermes-acp composite, so existing sessions keep the
same tool list and prompt cache.

The fresh-session assertion in test_make_agent_prefers_passed_toolsets_over_config_servers
now checks membership of the config MCP entry: the resolver returns the
expanded toolset keys rather than the bare composite name.

Refs #74582, #79516. Credit: #64045 (@israellot), #80309 (@thatssoheil),
#106834 (@nicolasramos) proposed the resolver routing.
2026-09-23 16:27:05 -07:00
teknium1
9a07debfbf test(acp): stub load_config on the real hermes_cli.config module
test_acp_real_agent_gets_session_db_for_recall replaced the whole
hermes_cli.config module with a one-attribute stub, so any real import
from it inside _make_agent (now hermes_cli.tools_config, for the shared
toolset resolver) died with ImportError. Patch the one function instead;
the assertions are unchanged.
2026-09-23 16:27:05 -07:00