Commit Graph

2750 Commits

Author SHA1 Message Date
Aniruddha Adak
6eef5ea0bb fix: fallback activation picks the OpenCode per-model wire (muse-spark → Responses)
A `fallback_providers` entry on opencode-go / opencode-zen / opencode-free (or a
custom provider hosted on opencode.ai) landed on `/chat/completions` unless the
user pinned `api_mode:` by hand: `_fallback_api_mode_resolved` re-detected the
wire for openai-codex, Nous, Anthropic URLs, Azure, direct OpenAI, gpt-5 and
Bedrock but never consulted `opencode_model_api_mode`, which the primary
`/model` path already uses. Responses-only models (muse-spark, gpt-*, grok-*)
500'd three times per activation and the chain advanced past a healthy entry;
Anthropic-wire models (minimax, qwen) on Go were misrouted the same way.

Resolve the OpenCode family through the same predicate the custom-provider
runtime uses (`_opencode_family_for_custom`: provider name or opencode.ai host)
and return the per-model table's mode. An explicit `api_mode:` still wins (the
resolver is only called for the chat_completions default), and chat-completions
OpenCode models keep their wire.

Fixes #102148
Salvages #102229 (@aniruddhaadak80)
2026-09-16 14:24:30 -07:00
teknium1
4a1daecc18 fix: seed Responses budget escalation from the observed ceiling when no cap is configured
With max_tokens unset the escalation armed nothing, so a default-config user hitting the
provider ceiling resent the same absent budget. The usage.output_tokens of the exhausted
response is that ceiling; seed from it (else 4096). Test control swapped to the reachable
invariant (non-empty fragment keeps today's replay) per the independent review.
2026-09-16 14:23:55 -07:00
teknium1
92c24d410d fix: Responses-wire length continuation raises the output cap and drops reasoning after an empty fragment
When a Responses-API call (muse-spark, gpt-*, grok-* routes) ends with
status=incomplete / incomplete_details.reason=max_output_tokens and no visible
text, reasoning consumed the whole output budget. continue_codex_incomplete
re-sent the identical request three times: same max_output_tokens, same
reasoning effort, plus a nudge. The model re-burned the same budget each time
and the turn ended as "Codex response remained incomplete after 3 continuation
attempts" with nothing for the user (#90393; measured in #103483 at xhigh:
out_tokens == max_output_tokens, reasoning_tokens == out - 3).

The chat-completions length path already handles this with two one-shot
overrides (_ephemeral_reasoning_off, _ephemeral_max_output_tokens); reuse them:
on a budget-exhausted empty fragment set reasoning off and double the cap
(2x, 4x, capped at 32768 / the configured cap) for the next attempt, and make
_build_codex_kwargs consume the ephemeral cap it never read before. A
reasoning-only status=completed response (Codex "still thinking") is untouched.

Live repro (fake Responses provider, real loop): before 2000/high, 2000/high,
2000/high -> partial; after 2000/high, 4000/no reasoning, 8000/no reasoning.
Chat-completions control unchanged (2000->4000->8000->16000, effort none).
2026-09-16 14:23:55 -07:00
teknium1
2b4ca80f5f test: trim Responses blank-carrier coverage to two invariants
Keep the fixture-driven test on the real failing turn shape from #103483
(no invented carrier, every reasoning item followed by its function_call,
call/output pairs intact) and the control that a lone reasoning item still
gets a non-empty follower. Drop the synthetic parametrized round: the fixture
already covers that shape once per tool round, and the repo caps a fix at two
invariant tests.

Part of #103483
2026-09-16 14:23:20 -07:00
actualrat1984
f33624c58a test(responses): cover real failing-turn shape from #103483
Fixture by @cristianbdev (content-redacted replay of the failing turn):
no invented blank carrier, every reasoning followed by its
function_call, all five call/output pairings preserved.
2026-09-16 14:23:20 -07:00
actualrat1984
1c7bc44eda fix(responses): avoid invented empty assistant turns before tool replay 2026-09-16 14:23:20 -07:00
NanPan
dceae1a766 test(agent): pin periodic scheduler context propagation
(cherry picked from commit 972a0444032ec7671ec30007ed883e0ef165e0f6)
2026-09-16 11:06:06 -07:00
NanPan
f1b7ae7206 fix(usage): preserve routed profile context in Nous account fetches
(cherry picked from commit 1675c92c8b0101315a9e18823a7c2e444be29259)
2026-09-16 11:06:06 -07:00
teknium
9977df914d fix(kanban): delegate_task children never act on the parent's HERMES_KANBAN_TASK
In-process delegate_task children (and cron runs fired from a worker) inherit
the dispatcher worker's HERMES_KANBAN_TASK via os.environ. Three readers still
gated on the bare env var instead of is_dispatcher_owned_worker_context():

- agent/turn_finalizer.py: a child exhausting ITS iteration budget recorded
  `timed_out` against the parent's task and released the parent's claim.
- tools/kanban_tools.py::inject_new_comments_from_env: a child polled operator
  notes addressed to the worker, steered itself with them, and advanced the
  shared per-task watermark so the worker never saw them.
- agent/title_generator.py::_kanban_task_title: a child's session was titled
  after the parent's card.

Each now uses the single identity predicate. Dispatcher-owned workers are
unchanged (existing #87096 tests still pass).

Fixes #112817
2026-09-16 10:43:48 -07:00
kshitijk4poor
b6bba97f7b fix(state): pin the shared storage-failure action copy to the failing profile
hermes_state_user_copy.py is the one table feeding the CLI banner, gateway
warning and TUI/Desktop RPC errors, and its action strings were still bare
(`hermes doctor --fix`, `hermes gateway stop`, `hermes sessions recover`).
Substitute {profile_arg} once in describe_storage_failure so those surfaces
get the same pin as the turn explainer.

Also keep final_response unbound in the turn_finalizer error fallback: it
feeds the external-memory sync and the background-review gate, which must
still see an empty response on a persistence-failed turn. One test binds the
fallback (explainer stubbed empty) and the untouched final_response.
2026-09-16 21:15:30 +05:30
kshitijk4poor
60945c70bd fix(agent): pin the turn-explainer hermes commands to the failing profile
The replaced / deleted_wal / default persistence explanations printed bare
`hermes gateway stop` and `hermes doctor`, while `{home}` in the same sentence
was already profile-aware. On a multi-profile backend (Desktop serve) the
session whose state.db failed is not the process default, and a bare `hermes`
follows the sticky active_profile — so the copy-pasteable command stops or
inspects the wrong profile's database. corrupt / fts_index were pinned by #105887;
this applies the same `{profile_arg}` substitution to every cause
instead of the two-cause tuple.

Spotted via #110073 (@JoaoMarcos44), whose explainer hunk added the selector
to the since-rewritten deleted_wal runbook.
2026-09-16 21:15:30 +05:30
brooklyn!
b4e4b9a650 test(approval): align CI with transcript hosting and terminal scope 2026-09-16 06:27:23 -05:00
brooklyn!
42e933b411 fix(approval): prepare batches for desktop session sources 2026-09-16 06:27:23 -05:00
brooklyn!
dd68d17567 feat(approval): prepare GUI terminal asks before ordered execution 2026-09-16 06:27:23 -05:00
kshitijk4poor
5d07f7fe66 test(agent): fold the aux create_client() seam tests to two contracts
Two parametrized invariant tests (native client sync/async; None or raising
profile falls back to the standard client) replace four near-duplicates, and
the hook's kwargs are pinned exactly to the mapping openai.OpenAI would have
received. The fixture isolates both provider registries through monkeypatch
instead of a hand-rolled snapshot/restore, drops the HERMES_HOME override that
tests/conftest.py already provides, and replaces the tuple-truthiness lambda
with a plain function. Call-site comment trimmed to the ordering WHY.
2026-09-16 11:57:56 +05:30
liuhao1024
4cd2eb013c fix(agent): honor ProviderProfile.create_client() for api_key aux routes
The auxiliary api_key branch built openai.OpenAI directly, so an
out-of-tree provider registered with auth_type="api_key" lost its
native transport for auxiliary tasks even though the main-agent path
(_provider_supplied_client) and the external_process branch honor the
same hook. Consult the profile's create_client() before the built-in
gemini/OpenAI ladder; None (the default) falls through untouched, and
a raising profile is logged and skipped (#112384).
2026-09-16 11:57:56 +05:30
teknium1
c2e5c94cd7 fix(tui): tolerate agents without session_cwd in _register_session_cwd; adapt stubs to the cwd kwarg
Workspace moves stamp agent.session_cwd so a lazily started Codex thread
starts in the moved-to directory. Agents that never had the attribute
(test doubles, slotted objects) must keep working, so stamp only when the
attribute exists. Test stubs of _set_session_context mirror the new cwd
kwarg, and the Codex gateway test asserts the contract (no pinned session
cwd) instead of the attribute's absence.
2026-09-15 22:30:11 -07:00
teknium1
3317b8c1e1 test(honcho): trim salvage coverage to the invariants; drop duplicate ACP session_cwd stamp
Salvage of #93452 (@outpoints). Keep one invariant per fix:
resolver-level "automatic title never remaps a strategy session",
integration "provider routes by logical workspace, not process cwd",
agent-level "title provenance + cwd reach the provider", deferred
Desktop/TUI build threads the session cwd, seeded branch titles are
derived, and the workspace-move E2E. Drop the plumbing/legacy-shape
tests that re-assert the same contract.

acp_adapter: AIAgent(cwd=...) now stamps session_cwd itself, so the
direct assignment after construction was a duplicate.
2026-09-15 22:30:11 -07:00
outpoints
44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints
5237cab756 fix(honcho): thread logical cwd through agent construction
(cherry picked from commit b1d7207c45311be658592c6ad34ee84634fed0ee)
2026-09-15 22:30:11 -07:00
outpoints
3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
teknium1
51a2f4878f test: pin the non-SDK facade gate alongside the escape hatch
The MoA aggregator and test stand-ins never merge extra_body; the bypass
must hand them the kwargs untouched or the conversation would be sent
empty. Fold that control into the existing rail test (still two tests).
2026-09-15 19:23:53 -07:00
kshitijk4poor
a12b3c7aa3 refactor(agent): import the transform bypass from its defining module, no re-export shim
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
2026-09-15 19:23:53 -07:00
kshitijk4poor
af7b60e8ec fix(agent): keep moved chat fields as slot placeholders so the wire body is byte-identical; one shared escape hatch
The cherry-picked helper deleted 'tools' from the typed kwargs, so the SDK's
post-transform extra_body merge appended it after the caller's extra_body keys —
equal dict, different bytes (byte-keyed prompt caches would miss). keep_slots=True
leaves [] placeholders that the merge overwrites in place. Drop the invented
HERMES_CHAT_SDK_TRANSFORM env var; the pre-existing HERMES_CODEX_SDK_TRANSFORM
hatch from #93650 now disables both API families. Tests trimmed to the two
invariants (byte-identity incl. caller extra_body precedence; escape hatch).
2026-09-15 19:23:53 -07:00
John Paul Soliva
1e39c93710 perf(agent): keep bulk chat-completions payloads out of the SDK request transform
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.

Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:

    101 msgs /  76 KB   13.6 ms -> 1.2 ms
    401 msgs / 190 KB   48.6 ms -> 2.0 ms
   1601 msgs / 650 KB  188.7 ms -> 5.8 ms

and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.

The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.

Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.

Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.

The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
2026-09-15 19:23:53 -07:00
KoNit-K
a457e91a50 fix(secrets): preserve OP_CONFIG_DIR for 1Password 2026-09-15 19:10:00 -07:00
teknium1
204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
Zheqing Zeng
9b003f201f fix(curator): seed shared read-marks store in the LLM consolidation fork
The read-before-write guard requires a skill_view mark from the SAME review
run before any skill_manage write. mark_background_review_skill_read
auto-creates a store when the ContextVar is unset, but tool workers run on
copied contexts, so marks recorded in one worker stayed invisible to the
others: every patch was refused with "current SKILL.md content has not been
loaded in this review turn" even after fresh full reads, and the
consolidation pass burned its iterations retrying a dead-end write.

The background-review fork already seeds a shared store before
run_conversation (agent/background_review.py); do the same in the curator
fork so every copied worker context shares one store.
2026-09-15 19:01:11 -07:00
teknium1
60f436b5f6 fix: redact '/'- and '~'-led secrets whose segments cannot be a path
Review finding (agent/redact.py::_should_redact_assignment): a '/'-prefixed
secret with a second '/' (`AWS_SECRET_ACCESS_KEY=/wJalrXUtnFEMIK7MDENG/bPx…`)
still parsed as a multi-segment path and leaked under a strong key; the same
held for a '~'-led value. Apply the opaque bar per segment (16+ chars, no
'.', mixed case and digits) to every '/' or '~' value instead of only to
single-segment ones, so `/home/u/.docker`, `~/.ssh/id_rsa` and
`S.gpg-agent.ssh`-style paths stay readable while base64 secrets mask.
2026-09-15 18:59:55 -07:00
teknium1
34067a7b6d fix: keep $(cmd) substitutions readable under strong-key assignments
Review finding (agent/redact.py::_PATH_OR_VAR_VALUE_RE): anchoring the
path/var exemption regressed `export SSH_AUTH_SOCK=$(gpgconf --list-dirs
agent-ssh-socket)` vs main — the `$(gpgconf` token no longer parsed as a
reference and was masked. Accept a leading `$(` as a reference atom in the
grammar and pin the gpg-agent line in
test_real_path_and_var_references_stay_readable.
2026-09-15 18:59:55 -07:00
teknium1
93a269c026 fix(redact): keep $VAR interpolations inside rc path values readable
The anchored path/variable grammar from the salvaged fix only allowed one
leading $VAR; a strong-key rc line such as SSH_AUTH_SOCK=/run/user/$UID/ssh or
SSH_AUTH_SOCK=$XDG_RUNTIME_DIR/agent.$USER.sock no longer parsed as a reference
and was masked, undoing the readability contract from 979576d938 for exactly
the lines it was written for.

Allow $VAR / ${VAR...} anywhere in the value (and ':' list separators). Crypt
digests still fall through to the credential checks: their '$' fields start
with a digit or carry '=' / ',', which the grammar rejects.
2026-09-15 18:59:55 -07:00
beardthelion
93a0ec705a fix(redact): anchor the path/var exemption so leading-/ and $ secrets still mask
_PATH_OR_VAR_VALUE_RE was an unanchored character class, so re.match made it a
first-character test: any assignment value beginning with '$', '/', or '~'
returned early from _should_redact_assignment, ahead of the strong-key and
opaque-credential checks. AWS secret access keys (~1 in 64 begin with '/') and
argon2/bcrypt digests (always '$'-prefixed) leaked verbatim through
redact_sensitive_text.

Anchor the pattern on both ends so the exemption only fires on a complete
$VAR/${VAR}/~/path//abs/path reference, and require a single-segment absolute
path — indistinguishable by shape from a high-entropy secret — to clear the
opaque-credential bar first. $VAR and ~/ references stay exempt
unconditionally, preserving the rc-readability contract that motivated the
exemption (SSH_AUTH_SOCK=$HOME/.ssh/agent.sock,
DOCKER_AUTH_CONFIG=/home/u/.docker).
2026-09-15 18:59:55 -07:00
KoNit-K
6c05cbf4bd test(agent): make aux timeout FD test deterministic 2026-09-15 18:48:35 -07:00
teknium1
1d14418ea2 fix(agent): concurrent worker survives a dict error result
_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
2026-09-15 18:42:10 -07:00
teknium1
339fa6d918 fix(gateway): bounded redacted result preview on tool.completed run events
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
2026-09-15 18:42:10 -07:00
KoNit-K
c89209d564 fix(gateway): report structured tool failures in run events 2026-09-15 18:42:10 -07:00
jerryhjones
2bd0f1c59e fix(kanban): scope the stop nudge to the dispatcher-owned worker
agent/kanban_stop.py::kanban_stop_nudge_enabled tested only HERMES_KANBAN_TASK,
which in-process delegate_task children (and cron runs fired inside a worker)
inherit from the worker's process environment. Those executions own no board
task and have the kanban toolset withheld, so the turn-end nudge ordered them to
call a tool they cannot reach — burning attempts, and in production driving
children to complete the parent's card through the CLI.

Gate on agent/delegation_context.py::is_dispatcher_owned_worker_context, the
predicate every other HERMES_KANBAN_* identity gate already uses. The real
worker and the HERMES_KANBAN_STOP_NUDGE opt-out are unchanged.

Salvaged from PR #84656 by @jerryhjones (re-applied onto the current facade
shape); the same gate was first proposed in PR #80023 by @webdevfrancisco
using the narrower delegated-child predicate.

Co-authored-by: webdevfrancisco <franciscombautista2015@gmail.com>
2026-09-15 18:41:19 -07:00
teknium1
efdf766cad test: trim concurrent multimodal log test to one invariant, reuse the stub
Keep only the invariant the fix owns (an envelope dict logs its serialized
size on the concurrent path); the plain-string control was already covered
by the pre-existing behaviour and doubled the file. Import the AIAgent stub
and fake tool-call shapes from test_start_order_gate.py instead of copying
them, so the concurrent-executor stub has one home.

Part of #112095. Salvage of #112104 (@kokhlo).
2026-09-15 18:40:03 -07:00
Konstantin Khlopkov
699037176c fix(agent): log the serialized size of multimodal results in the concurrent executor
The concurrent completion line logged len(result) directly, so a native-path
vision_analyze envelope dict reported "4 chars" — its key count — while the
sequential path already logs the serialized length. Mirror the sequential
measurement so parallel multimodal calls stop looking truncated in logs.
2026-09-15 18:40:03 -07:00
teknium1
996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
KoNit-K
c93f2e1d59 fix(agent): update zai vision fallback 2026-09-15 18:23:55 -07:00
teknium1
cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1
ee49b7d25d fix(agent): file-mutation footer states failed edits, not "files were NOT modified"
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.

- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
  and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
  (mtime_ns, size) snapshot; a later success clears every entry with the same
  identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
  entries whose file changed since the failed call, so a receipt-less
  mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
  way the file tools resolved them.

Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.

Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:22:12 -07:00
teknium1
0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1
66c9440826 fix: stream retry no longer replays the old model on the switched-to provider
After a mid-turn /model switch while a stream was stalled, the streaming
retry loop re-sent the request it had captured at construction time. That
payload still named the OLD model, but every stream (re)open builds its
request client from the LIVE agent, so the new provider's base_url received
a foreign model slug: 404 "Not found the model ...", then the turn sat in
the provider's rate-limit hold (#112121).

_StreamingCall now records the route (model, provider, base_url, api_mode)
its api_kwargs were built for. When a retry is about to be issued and the
live route differs, the streamer stops and hands the transient error back
to the turn loop instead. The turn loop already rebuilds the request per
attempt for the CURRENT route (turn_api_request.build_api_request: model,
wire shape, prompt-cache decoration, provider request overrides), so a
re-keyed model alone would still have shipped a payload shaped for the old
provider. Non-streaming requests have no in-process retry, and the
fallback / restore-primary paths go through the same turn-loop rebuild, so
this is the only site that replayed a captured route.

Fixes #112121

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:21:17 -07:00
teknium1
d54ae02606 fix(pricing): custom-provider /models prices already per-million are no longer inflated 1e6x
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.

Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.

The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:20:51 -07:00
teknium1
5435ac8cc4 fix(aux): clamp extra_body.reasoning effort too so auxiliary.<task>.reasoning_effort ultra never reaches the wire 2026-09-15 18:20:13 -07:00
teknium1
9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K
c02db64077 fix(agent): clamp auxiliary ultra reasoning effort 2026-09-15 18:20:13 -07:00
teknium1
f55d4f6747 fix(aux): bare-custom AuthError yields no custom endpoint instead of a stale env OPENAI_BASE_URL 2026-09-15 18:19:50 -07:00