Commit Graph

1317 Commits

Author SHA1 Message Date
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
teknium1
5be70b5e25 docs(mcp): note Google-hosted OAuth servers get access_type=offline for a refresh token 2026-09-20 18:22:18 -07:00
teknium1
f568b860d7 fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
  (production entry) on a real AIAgent with a one-entry chain, so dropping the
  reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
  rate-limited primary's declared reset is sooner than N seconds,
  try_activate_fallback returns False and no cooldown is armed; docs row added.
2026-09-20 17:00:43 -07:00
teknium1
97beaeeeb3 feat(delegate): warn the child at 80% of its inactivity window before abandoning it
A child that stalls under a configured delegation.child_timeout_seconds used to
learn about the budget only by dying, losing its whole context. The liveness
wait now queues a one-line "[delegation budget warning]" through the child's
steer channel once the idle window is 80% spent (delivered at the child's next
iteration boundary), so a slow-but-recoverable child can wrap up and return its
summary. The warning fires once per idle window and re-arms when progress
resets the window; a progressing child never sees it.

Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged.
2026-09-20 15:50:20 -07:00
finn763
03973bd02b fix(delegate): child_timeout_seconds bounds inactivity, not total runtime
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a
plain `settled.wait(timeout=child_timeout)`, so any child that outlived the
budget was abandoned even while the provider was actively serving it.

The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could
not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint
— but the mechanism it names is exactly this gate: a child waiting on an
in-flight LLM completion, killed with a nearly-finished context. A slow child is
already bounded elsewhere (the per-call stale watchdog, the heartbeat's
staleness verdict), so this cap could only ever kill children the runtime had
judged healthy.

`child_timeout_seconds` now measures time with NO progress: the wait runs in
slices and restarts the window on the same signals the heartbeat's stale verdict
reads (completed call, tool change, activity-clock tick). A frozen child is
still abandoned when the window elapses; a progressing one is never killed for
taking long.

Timeout entries also carry `last_event_age` (how long the child had been silent),
so operators can tell a slow provider from a runaway without transcript
forensics.

Fixes the mechanism reported in #116001. The budget warning and
continuation-respawn items in that issue are separate features and are not part
of this change.
2026-09-20 15:50:20 -07:00
John Paul Soliva
786c0e3f9d fix(cron): a bot-chat delivery builds its child env for the DESTINATION profile, not the sender
`deliver: bot-chat:<other profile>` spawned the destination's agent turn with the sending
gateway's whole environment: its `.env` settings, bridged `TERMINAL_*` policy, platform
authorization gates and provider credentials. The lane called
`strip_launch_profile_env(env)` with no target, so the strip resolved against the ambient home
override — which is the SENDER's home, never the destination's. On an ordinary root-profile
gateway (`hermes gateway run`, no `-p`) `_is_routed_home` is then false and the strip is a
complete no-op, including the #113270 gate strip that lives after its early return.

This is the only cron child built for a profile other than the one whose tick spawned it; the
worker lane (`scheduler.py`) targets its own home, so its no-target call is correct. Build this
one through `served_profile_child_env(target_home=home, inherit_credentials=True)` — the helper
`kanban_db_dispatch` and `web_server_gateway` already use for cross-profile spawns: it strips the
launch residue against the real target, scrubs credentials the launch process was given by
systemd/Compose/the shell (which no name-based strip can see), points TMPDIR at the destination's
scratch, and overlays the destination's own secrets, as a standalone `hermes -p <profile>` has.

A failure to build that environment (an unreadable target home under per-user 0700, a broken
secret source) is reported as a refusal string like every other failure in this lane rather than
raised: `_deliver_result`'s fan-out does not catch, unlike the deferred drain.

Regressions drive the real `_deliver_to_bot_chat`; removing the fix fails the two leak witnesses
(`HERMES_MODEL leaked from the launch profile`, and the firing profile's gate reaching another
profile's turn under an active override) and leaves the four guard tests green.

Fixes #117220

(cherry picked from commit 6cc81d7ddad8c9793f21fdda7c2e260c3ee1ca44)
2026-09-20 15:33:28 -07:00
teknium1
f9d6d04332 fix(computer-use): doctor and status report a configured cua-driver daemon whose serve is not listening
A hand-written Linux daemon unit (systemd user unit / XDG autostart entry
running `cua-driver serve`) can be dead for days — crash loop, stopped, never
started — while `hermes computer-use doctor` and `status` report a healthy
binary: the runtime contract only checks the binary (`manifest`), and nothing
ever connected to the daemon socket (#114748).

- tools/computer_use/cua_backend.py::cua_daemon_listening — socket-level
  liveness via `cua-driver status [--socket PATH]` (rc 0 = a daemon answered,
  "not running" = dead, anything else = unknown). Never raises.
- tools/computer_use/doctor.py::cua_daemon_units — one scan of the units that
  run cua-driver (kind, unit, exec target, runs `serve`, `--socket` path with
  `%h` expanded); the pruned-Exec guard now filters that list instead of
  re-scanning.
- doctor.py::_apply_daemon_liveness_guard — per `serve` unit: `pass` when its
  socket answers, `fail` (degrading `ok`) when not, with the hint that a driver
  reinstall does not start the daemon. Unconfigured daemons are never probed:
  on Linux the MCP runtime needs none, so a silent default socket is normal.
- `hermes computer-use status` prints the dead-daemon line and exits 1.
- docs: the Linux daemon-unit paragraph under the doctor section.

Not changed: _maybe_repair_runtime_contract. The repair is gated on the
binary-level contract only; a dead daemon never enters it, so a reinstall was
never triggered by the daemon (the `.release_installed/<version>` marker the
reporter saw is written by cua-driver itself on any first run of the binary —
observed via strace of `cua-driver status` under a fresh HOME).

Supersedes #114928 (@Finn763): same finding, ~600-line implementation with a
new module and a repair-gate rewrite; this is the ~90-line version on the
existing doctor seams.
2026-09-20 15:19:55 -07:00
teknium1
bdd7192bcf fix(kanban): a profile_routes-pinned profile without this platform's adapter delivers via the primary bot
`_adapter_for_subscription` fail-closed on ANY connected secondary adapter of
the pinned profile: a profile that ran Signal/Home Assistant bots but held no
Telegram token (removed on purpose to avoid a duplicate-credential collision
with the shared bot) could never receive kanban notifications in a Telegram
group that `gateway.profile_routes` pins to it — the claim rewound every tick
and the docs' route-only promise could not be met because the profile is never
route-only.

Only an adapter for the subscription's OWN platform is a credential boundary
(`_authorization_adapter` already answered for it). Adapters on other
platforms no longer gate delivery; the exact-route match still authorizes the
primary solely for a chat `profile_routes` pins to that served profile — the
same authority the primary already exercises for that chat's inbound turns.
The "other-platform adapters but none for X" warning added for this branch is
superseded by delivery; the stamped-with-the-wrong-profile warning stays.

Fixes #115460
2026-09-20 14:22:58 -07:00
teknium1
d9fdaeb434 feat(website): catalog cards open the Desktop Install Plugin dialog by catalog name
Every catalog card gets an 'Open in Hermes Desktop' link:
hermes://plugin/install?catalog=<name>. The app resolves the reviewed pin
itself, so the page hands it a catalog name and never a repo URL; the CLI
install command stays in the expanded card for people without the app.
In the in-app picker embed the existing '+ Add to this Agent' button is
unchanged.
2026-09-20 13:56:30 -07:00
teknium1
aae3c737ad feat(desktop): hermes://plugin/install?catalog=<name> opens the reviewed catalog install dialog
A `catalog=<name>` deep link now resolves the name against the live plugin
catalog feed (the same `/docs/api/plugins.json` the Capabilities → Plugins
picker renders) and opens the Install Plugin dialog in its reviewed/pinned
catalog mode — identical to an in-app catalog pick, via one shared
`openCatalogPluginInstall` helper the Plugins tab now uses too.

The `catalog` param claims the link outright: an unknown, invalid, or
unresolvable name is a clear error toast and nothing else. It is never
reinterpreted as a git identifier, so a link cannot smuggle an unreviewed
repo behind a familiar-looking name (a `repo=` riding along is ignored).
2026-09-20 13:56:30 -07:00
teknium1
066c9dca9f docs: describe the Processes block in the CLI/TUI live-work dock 2026-09-20 13:55:03 -07:00
teknium1
2c0b2a980c docs(memory): add a troubleshooting section for memory that does not survive a new session
A 42-comment r/hermesagent thread ("Frustrated.") shows the recurring
shape: the model answers "Done, I will remember that", never calls the
memory tool, and the next session knows nothing. The page had no place
that tells a user to open MEMORY.md and check, or lists the other
reasons a write can be invisible (staged approval, another profile,
memory disabled, frozen snapshot). Add an ordered checklist and say
plainly that .env variables are not memory.
2026-09-20 13:13:19 -07:00
teknium1
41b6ba09d9 docs: name the registered tool in prose that still says cronjob/todo/process
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
2026-09-20 12:56:25 -07:00
teknium1
a4f9857ff5 fix(mcp): OAuth login waits for oauth.timeout and a probe timeout names itself
`hermes mcp login`, the dashboard re-auth (web_server_mcp.py) and the Desktop
re-auth (tui_gateway/mcp_oauth_sessions.py) each bounded the login probe at
`max(connect_timeout, 315)`. The 315 was the default 300 s callback window
plus headroom, frozen: a user who set `oauth: {timeout: 3600}` still had the
probe cancelled at 315 s. That expiry was a bare `asyncio.TimeoutError`, whose
`str()` is '', so the CLI printed a blank `✗ Authentication failed:` line
(#116278, reporter's steps 5).

`tools/mcp_oauth.py::login_connect_timeout(config)` computes the bound once —
`max(connect_timeout, oauth.timeout + 15)` — and all three call sites use it.
`_probe_single_server` re-raises its wait_for expiry as a TimeoutError naming
the server, the elapsed bound and both governing knobs, so every caller's
`humanized or exc` renders a reason.

Slimmer redo of the blank-line half of #114527 (@liuhao1024): that PR races
the connect against `asyncio.wait` so an inner exception can win; with the
window now sized from oauth.timeout the callback waiter's own
`OAuth callback timed out` message wins on its own, and the described
TimeoutError covers the remaining case in one place.

Completes the timeout atom of #116278 (closed by #116658); refs #103633

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:40:36 -07:00
teknium1
e6bb65aa2f fix(mcp): startup summary names every failed server with its connect error
`MCP: registered N tool(s) from M server(s) (2 failed)` left the failing
identity diagnosable only by elimination from the per-server `registered`
lines. The per-server WARNING fires only on the immediate-failure path; a
candidate skipped for its retry cooldown (a failure from an earlier pass) is
counted as failed with no line of its own at all.

`_connected_summary` now returns `(name, reason)` pairs and `_log_summary`
prints them inline: `(2 failed: github (Connection closed); notion (HTTP 401
...))`. The reason is the recorded `_server_connect_errors` entry (already
credential-scrubbed by `_format_connect_error`); a candidate never attempted
this pass reads `not attempted (in retry cooldown)`. One line, no second
WARNING per server on the path that already warns.

Slimmer redo of #114794 (@liuhao1024, earliest) and #114872 (@Finn763): both
name the failures via an extra WARNING per server; inline on the summary keeps
the immediate-failure path at its current two WARNINGs and still covers the
cooldown case. #114872's extra `_sanitize_error` pass is redundant with the
recorder.

Fixes #114746

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
2026-09-20 12:19:28 -07:00
teknium1
17958bfef3 fix(doctor): validate GITHUB_TOKEN/GH_TOKEN against api.github.com and name the .env file
`hermes doctor` reported "GitHub token configured (authenticated API access)" for any
value in .env, including an expired classic PAT — the very token that shadowed a working
`gh` login for every git-auth clone (#115257). The resolver half (fall through to the gh
CLI when GitHub refuses the .env token) landed separately; this adds the diagnostic half:
a "GitHub token" row under API Connectivity that sends the configured token to
api.github.com/user and, on 401, names the variable and the .env file so the user can
remove or replace it. No token configured: the row is skipped (the Skills Hub section
already reports gh-CLI / no-token state).
2026-09-20 12:15:07 -07:00
teknium1
6d335529ea fix(api_server): stamp the drain boundary on restart too; one HTTP-level invariant; docs
- request_restart() opens the same drain window as stop() (new turns refused,
  in-flight work awaited), so pollers of GET /v1/runs/{id} need the
  shutdown_requested_at marker from that moment as well; the marker is idempotent
  so stop() re-marking after the restart wait is a no-op.
- Collapse the three adapter-level tests into one invariant that drives the real
  GET /v1/runs/{id} route: live run keeps status=running but carries the marker
  (durably), a status set after the boundary inherits it, a terminal run never does.
- Shutdown-path fakes gain _mark_api_runs_shutdown_requested (the real mixin method
  where the fake already borrows _api_server_hook, a 0-stub where it has no adapter).
- Document the field on the runs API page.

Refs #115133.
2026-09-20 12:13:44 -07:00
Aaron08140
efa31fff8c docs(hooks): state the Windows interpreter routing the command contract now has
Line 1687 documents 'runs via shlex.split, shell=False', which after the
preceding commit is no longer the whole story on Windows: the bare script paths
every example on this page use are routed through their interpreter there. Left
unstated, the page keeps describing the behaviour that caused the bug.
2026-09-20 11:27:39 -07:00
teknium1
21e96c9593 docs(kanban): document done-column completion order and the completed-desc sort
Follow-up to the salvaged #116041: the board's `done` column is now newest-
completed-first and `hermes kanban list --sort completed-desc` exists, so the
feature doc and the CLI usage block say so; the stale "per-column ordering
comes from list_tasks" comment in `get_board` now describes the queue
columns only (wording from #116051).

Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-20 10:18:14 -07:00
teknium1
0469740ab3 feat(cron): jobs follow the main agent model at fire time; pinned locks it on request
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.

`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.

Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).

Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
2026-09-20 09:20:51 -07:00
joaomarcos
f7ff7d3f4f fix(config): register tool search deferral settings 2026-09-19 23:51:25 -07:00
teknium1
729b1c3d56 fix(cron): bound the standalone send inside the coroutine so the thread fallback keeps it
Wrapping the coroutine in `asyncio.wait_for` outside `asyncio.run` left the
running-loop fallback (which closes the unstarted coroutine and retries in a
thread) with a never-awaited wait_for wrapper and an unbounded inner send.
Awaiting `wait_for` inside `_send` bounds every runner of the coroutine.

The invariant tests now drive the production entry point `_deliver_result`
with no live adapters (the standalone lane) instead of `_standalone_send`
directly, and the user-visible knob `cron.standalone_send_timeout_seconds`
is documented (#115469).
2026-09-19 23:50:50 -07:00
teknium1
0818892db3 fix(mcp): accept an origin-issued metadata document for a path-scoped OAuth authorization server
A protected resource may advertise a path-scoped authorization server
(`https://www.strava.com/mcp-issuer`) whose RFC 8414 document, served from
`/.well-known/oauth-authorization-server/mcp-issuer`, declares the origin
(`https://www.strava.com`) as its issuer. The SDK's exact-string check
(`validate_metadata_issuer`, RFC 8414 §3.3) rejected that document with
"Authorization server metadata issuer mismatch" and the connection parked
before registration or login (#116233).

`metadata_issued_by_origin` accepts exactly that shape and nothing else: the
document must have been read from the well-known URL derived from the
advertised identifier (so a redirect target, the root document or an OIDC
fallback never qualify) and its issuer must be the advertised identifier's
origin. Only the origin's operator controls that location, so a party
controlling a path or a sibling host cannot use it to make the client accept
another server's endpoints; whoever could would already control the
exact-match document too. The browser flow applies it in the mixin's request pump
(installs the document, hands the SDK a 204 so its loop stops) without touching
`auth_server_url`, so SEP-2352 credential binding keeps the advertised
identifier while the RFC 9207 `iss` check and refresh-token issuer binding use
the document's issuer. The device flow applies the same rule in its own
discovery and now binds registered credentials the way the SDK's Step 4 does,
so the runtime flow reuses them instead of discarding them on the next 401.

Supersedes #116359: its rule accepted the origin issuer from any discovery URL
(root and OIDC fallbacks, redirect targets) and fabricated a 500 response.

Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
2026-09-19 23:40:35 -07:00
teknium1
e84f0a1c5b fix(codex): a codex app-server thread started from scratch is seeded with the session's prior turns
A codex thread that codex hands back via thread/resume already holds the conversation, but a
thread started fresh did not: a session that ran on another provider before /model switched to
openai-codex, a stored thread codex could not resume, or a thread retired mid-session (prompt
composition change, wedged client) answered the first turn blind. The prior user/assistant text,
tool names and tool-result previews (most recent 32K chars) now ride once on
thread/start.developerInstructions after the prompt composition; thread/resume never carries them,
and the recorded composition stays the bare prompt so the seed cannot make the next turn retire the
thread.

Direction from #26081 (first-turn seeding of the Hermes transcript); redone on the extracted
agent/codex_runtime.py path with the system prompt sent once (#115759) instead of duplicated.

Completes #26035 / #74712 (closed by #115759 for the prompt half; this is the history half).
Co-authored-by: LeonSGP43 <154585401+LeonSGP43@users.noreply.github.com>
2026-09-19 20:44:24 -07:00
Teknium
8a92051f20 Merge pull request #116328 from NousResearch/fix/boa-w3-new-reports-cron-codex
fix(cron): missing-credential preflight verdict names the profile and HERMES_HOME it read (#116213)
2026-09-19 14:33:58 -07:00
Teknium
b5fcf635dc Merge pull request #116343 from NousResearch/boa-w3-codex-thread
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
2026-09-19 14:31:55 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
teknium1
e77e24a6a6 fix: persist the codex thread id per session and thread/resume it across an API-server restart
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.

Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.

Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
2026-09-19 12:27:01 -07:00
teknium1
fc49f7619d fix(cron): a missing-credential preflight verdict names the profile and HERMES_HOME it read
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.

Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.

Part of #116213
2026-09-19 12:14:32 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
d19963782b fix(api-server): anchor the Responses current turn on this turn's user row, not a history prefix match
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.

Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.

The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).

Fixes #89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
2026-09-19 12:11:23 -07:00
Teknium
f0d8efe4e7 Merge pull request #115905 from NousResearch/fix/boa-res-R2-app-server-customprov
feat(codex): named custom providers work with the codex_app_server runtime (#75186, salvage #75191)
2026-09-19 11:35:36 -07:00
teknium1
049a62ab3d Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	website/docs/user-guide/features/codex-app-server-runtime.md
2026-09-19 11:23:47 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
Teknium
271cf9e2d5 Merge pull request #115938 from NousResearch/fix/boa-res-R8-openai-native-search
feat(web): openai-native backend lets Codex Responses turns use the server-side web_search built-in (#19320, salvage #107377)
2026-09-19 11:15:31 -07:00
Teknium
3df8ad0d7f Merge pull request #115903 from NousResearch/fix/boa-res-R2-app-server-interim
fix(api-server): codex commentary reaches streaming clients on session SSE, /v1/runs and /v1/responses (#67580, salvage #67593 #67613)
2026-09-19 11:14:10 -07:00
Teknium
c99d016388 Merge pull request #115857 from NousResearch/fix/boa-response-store-api-server
fix(api-server): chained /v1/responses turns store each message once instead of doubling history every turn (#95137, #101644, salvage #85678)
2026-09-19 11:11:48 -07:00
Teknium
d6ff3bcfee Merge pull request #115840 from NousResearch/fix/boa-tts-stt-voice-sample-rate
fix(tts): OpenAI-compatible streaming TTS plays at the endpoint's reported sample rate (#76466, salvage #76501)
2026-09-19 11:10:21 -07:00
Teknium
61db846f19 Merge pull request #115831 from NousResearch/fix/boa-desktop-openai-codex-fallback
fix(desktop): open chats switch to a fallback provider added after they were opened (#95066, salvage #95139)
2026-09-19 11:09:59 -07:00
Teknium
438d2f061a Merge pull request #115763 from NousResearch/fix/boa-codex-oauth-refresh-401-soft-failure
fix(codex): usage-limit soft failures rotate the credential pool before provider fallback (#24159, salvage #24173)
2026-09-19 11:07:49 -07:00
teknium1
ec2c4eeb80 chore: merge origin/main (resolve apps/desktop/src/lib/voice-client-direct.ts, hermes_cli/config_defaults.py, tools/voice_client_config.py, website/docs/user-guide/features/tts.md) 2026-09-19 10:54:38 -07:00
teknium1
12cc17fd76 chore: merge origin/main (resolve gateway/platforms/api_server.py, gateway/platforms/api_server_openai_routes.py) 2026-09-19 10:54:05 -07:00
teknium1
349f67778f chore: merge origin/main (resolve agent/transports/codex.py) 2026-09-19 10:51:50 -07:00
teknium1
c5490df54b chore: merge origin/main (resolve hermes_cli/config_defaults.py) 2026-09-19 10:51:43 -07:00
teknium1
6b415f8754 chore: merge origin/main (resolve gateway/platforms/api_server_openai_routes.py) 2026-09-19 10:51:28 -07:00
teknium1
54c01bc19a chore: merge origin/main (resolve hermes_cli/runtime_provider.py, website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:51:04 -07:00
teknium1
4171ad5967 chore: merge origin/main (resolve website/docs/user-guide/features/fallback-providers.md) 2026-09-19 10:47:48 -07:00
teknium1
68f3ff788a chore: merge origin/main (resolve website/docs/user-guide/features/credential-pools.md) 2026-09-19 10:47:48 -07:00
teknium1
bf6977fe16 chore: merge origin/main (resolve website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:47:48 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00