Commit Graph

4568 Commits

Author SHA1 Message Date
Teknium
520e63661c fix: keep command-auth model discovery lazy across config and setup 2026-09-07 21:22:49 -07:00
Hayden Moulds
c111ede3e5 fix(picker): resolve key_cmd credentials for model discovery
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted
by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The
request path has honoured it since it landed, but the picker resolved probe
credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed
`/v1/models` with an EMPTY key.

Against an authenticated endpoint the probe 401s, discovery returns nothing,
and the provider falls back to its single configured default model. The
picker shows ONE model, indistinguishable from an endpoint that genuinely
serves one — while inference keeps working, because that path mints
correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models
discovered with an empty key, 26 with the minted token.

Both picker probe sites already funnel through `_entry_credentials()`, so
the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and
each site falls back to `resolve_probe_token()` after api_key/key_env. An
explicit static key still wins, so existing configs are unaffected.

The identity is keyed on the COMMAND, never the minted token: the token
rotates on every refresh, so keying on its value would change the group
fingerprint constantly and force a re-probe on every open. Two entries on
one URL with different helpers still get distinct rows.

`resolve_probe_token()` lives in agent.command_token_source, which already
owns key_cmd minting, and shares the CommandTokenSource cache with the
request path — a cache read, not a fresh sign-in. Fail-closed: a helper
needing an interactive sign-in degrades to today's empty-key behaviour
rather than taking down every other provider's row.

`_model_flow_named_custom` (the `hermes model` setup flow) is the sibling
path — it builds its own `Authorization: Bearer` from the same incomplete
resolution — and is fixed the same way, with one ordering constraint: the
value persisted to config.yaml is computed BEFORE the mint, so a short-lived
bearer can never be written back to shadow the key_cmd meant to re-mint it.

Tests drive the real code paths and assert on the credential each probe
receives rather than on function source, so a semantics-preserving refactor
does not fail them. Verified they fail with the fix reverted.
2026-09-07 21:22:49 -07:00
Teknium
bbcf1ee180 fix: preserve native Gemini union constraints
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.

Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
2026-09-07 21:10:50 -07:00
Max Freedom Pollard
6a04ea67c0 fix(gemini): collapse array-typed tool schemas instead of crashing translation
The enum-compatibility check evaluated `[...] in {...}` on an array `type`,
raising TypeError: unhashable type: 'list' and aborting translation of the whole
tool catalog rather than the one offending tool.

Addresses both review points:

- Reuses tools.schema_sanitizer._normalize_type_array instead of picking the
  first non-null member, so a real union becomes an anyOf of single-type
  branches and no branch is dropped.
- Derives the type outside the key loop and sets nullable after it, so the
  flag implied by "null" in the array beats an input nullable: false whichever
  key the producer emitted first. Both orders are pinned by a parametrized test.
2026-09-07 21:10:50 -07:00
Teknium
5280fe9987 fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.

Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
2026-09-07 16:48:29 -07:00
Teknium
e9313f6458 fix: let provider evidence adjudicate past-window preflight estimates 2026-09-07 14:11:41 -07:00
Teknium
3114916ee4 fix(gateway): carry accepted-input ownership through persistence
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.

Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
2026-09-07 14:11:18 -07:00
kshitijk4poor
ed063034ad test(compressor): pin custom_providers threading at the get_model_context_length seam
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
2026-09-08 02:27:01 +05:30
dhruv kejriwal
71516214c3 fix(compressor): thread custom_providers into context-length resolution
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
2026-09-08 02:27:01 +05:30
kshitijk4poor
b6852995ed refactor(openai): one is_astra_model predicate; gate Astra at the untrusted Codex inputs
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.

``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.

``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.
2026-09-07 21:43:54 +05:30
kshitijk4poor
170a73637d fix(openai): drop prompt_cache_options from Astra requests — not an SDK kwarg, 30m is the server default
Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.

OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.

Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.

Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
2026-09-07 21:43:54 +05:30
Eva
299d86851c fix(openai): keep Astra 900K alias gated and wire-compatible
(cherry picked from commit c7cd27d7f050598b9dc052c73fa1dc78c045d3c6)
2026-09-07 21:43:54 +05:30
Michael Steuer
17c7be1485 feat(models): preserve live-verified Astra 900K opt-in from #103132
Retain the two context-variant metadata additions by Michael Steuer. Keep the dedicated Astra reasoning contract already present in #103057 rather than replacing it with the GPT-5.6 vocabulary.

(cherry picked from commit add3a4fa6f31ea3f5fdad701a840d4f8530eb30e)
(cherry picked from commit 1672f12c260260ea38ed636528c02aee734d62b1)
2026-09-07 21:43:54 +05:30
Eva
3825d25191 fix(openai): cap Astra Codex OAuth fallback
(cherry picked from commit bb4156c3881cbaf20736f0fcfd6c3cc5c861a6fb)
2026-09-07 21:43:54 +05:30
Eva
850680fcdf fix(openai): require canonical host for Astra cache
(cherry picked from commit 6e73cc5cc66e4003276c4472e6ce2d77e602f130)
2026-09-07 21:43:54 +05:30
Eva
5a82258626 fix(openai): keep Astra rules on eligible routes
(cherry picked from commit f92fb9d9964e6e8ea118035c548aae812054f504)
2026-09-07 21:43:54 +05:30
Eva
c990017482 test(openai): close Astra baseline review gaps
(cherry picked from commit 8c27b7c9316ba675e4d9ebcc7a659754f37f18ef)
2026-09-07 21:43:54 +05:30
Eva
2c315ff59b feat(openai): add GPT-6 Astra baseline support
(cherry picked from commit a8c53d20c6b16cc35745e364e16bb7259166a1d3)
2026-09-07 21:43:54 +05:30
Teknium
a7198a8855 fix: keep budget checkpoints out of cancelled tool results
Skip checkpoint evaluation when a turn is interrupted so cancellation rows
remain durable without urging continued execution. The existing minimal-agent
interrupt regression also avoids dereferencing an absent iteration budget.

Consolidate the warning coverage into two invariants, including real SQLite
readback and dispatcher/child scope controls. Cold-start tool availability
between construction cases to model independent worker processes. Place ratio
normalization beside the existing iteration budget instead of growing init.

Real cancelled-tool A/B against current main, draft, and fix: three cancelled
rows and zero writes on all arms; persisted checkpoint notices 0 / 1 / 0.
Repeated scripted HTTP/SQLite loop A/B preserves completion opportunity,
ordinary default-off behavior, and blocked/two-failure exhaustion behavior.

Local targeted run initially passed 15 cases with one fixture cache-isolation
failure; corrected target and inherited affected suites remain queued behind
the campaign lock. This commit is not a CI-green or merge-ready claim.
2026-09-07 08:28:43 -07:00
Teknium
93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium
37fb7adfd6 fix: structured reasoning no longer breaks chat consumers
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.

Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.

Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.

Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:24:54 -07:00
Teknium
bbbcde8173 fix(process): stop late forks escaping deadline tree cleanup 2026-09-07 08:19:32 -07:00
Teknium
ab98a92a45 fix(notifications): report applied skill batch operations
Use successful applied result records rather than requested operations, and keep staged writes silent. Include legacy delete/write messages.

Fixes #104506
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 08:14:14 -07:00
Teknium
62f64ae5b0 fix(display): retain provider provenance for persisted fallback readings 2026-09-07 08:13:01 -07:00
Teknium
b4e2ae9e76 test(display): exercise real usage and usage-less provider controls 2026-09-07 08:13:01 -07:00
Teknium
04767e7aaa fix(display): distinguish estimated context from provider usage 2026-09-07 08:13:01 -07:00
Teknium
1eb1b795b5 test(gemini): verify alias routing and compatible endpoint controls 2026-09-07 08:10:36 -07:00
Charles Ji
a745101e5f fix(gemini): route google-alias fallback providers through GeminiNativeClient
fallback_providers entries using the "google" alias for the gemini
profile were falling through to the generic OpenAI-SDK client because
the native-client gate only matched the literal string "gemini". That
client posts straight to the raw REST endpoint, sending thinking_config
as an unnested top-level field, which Gemini rejects with:

  Invalid JSON payload received. Unknown name "thinking_config": Cannot find field.

Broaden the check to the same alias set the gemini profile itself
registers (mirrors _GEMINI_NATIVE_PROVIDER_NAMES already used for this
in auxiliary_client.py).

Fixes #104583
2026-09-07 08:10:36 -07:00
Teknium
a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium
7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium
bbbccd3935 fix: failed probe credentials cannot fall through to configured auth 2026-09-07 08:08:04 -07:00
Teknium
7d44fe9c74 fix: capability probes send minted credentials instead of callable representations
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:08:04 -07:00
Teknium
bdf45abd7c fix: prefer owned Anthropic grants and bind auxiliary refresh to request credentials
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>

Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
2026-09-07 08:07:26 -07:00
Brian Le
32a59f3bf7 feat: refresh one pooled OAuth grant from the CLI 2026-09-07 08:06:48 -07:00
Brian Le
1a4bb74a40 feat: choose pooled credential priority from the CLI 2026-09-07 08:06:48 -07:00
Brian Le
b86fb277b8 fix: count credential selections across every pool strategy 2026-09-07 08:06:48 -07:00
Brian Le
53221df05f feat: reset one pooled credential without clearing sibling cooldowns 2026-09-07 08:06:48 -07:00
Teknium
12ad29a48a refactor: isolate credential pool administration methods 2026-09-07 08:06:48 -07:00
Teknium
b578261584 fix: keep Kanban worker scope out of descendant processes
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.

Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.

Refs #103974, #104058, #104904
2026-09-07 07:10:28 -07:00
Teknium
fc70d05bb1 fix: arm primary cooldown once while walking fallback candidates 2026-09-07 07:08:25 -07:00
Teknium
68b16aa7e3 fix: retain exhausted-chain cooldown classification after extraction 2026-09-07 07:08:25 -07:00
Teknium
fb2f66f586 fix: describe remaining retry eligibility without promising recovery 2026-09-07 07:08:25 -07:00
Halldrix
ae4f777c69 fix(agent): surface armed rate-limit cooldown in fallback notice (#104120)
_arm_rate_limit_cooldown now returns the armed backoff seconds so
try_activate_fallback can append them to the user-facing notice
(Primary retried in ~N min/h) instead of discarding the one number
that decides whether the user waits or re-plans. Duration, never
wall-clock: the stored deadline is monotonic-based. Non-rate-limit
reasons and chain-switches from an active fallback arm nothing, so
no suffix is printed there.
2026-09-07 07:08:25 -07:00
Teknium
cb1a42d33b refactor: isolate custom health identity and trim redundant alias tests 2026-09-07 07:06:58 -07:00
fangliquanflq
84bbf176a3 fix(agent): preserve route URL path identity 2026-09-07 07:06:58 -07:00
fangliquanflq
e3ff3e78d1 fix(agent): quarantine failed fallback destination 2026-09-07 07:06:58 -07:00
fangliquanflq
e033bbaa3c fix(agent): quarantine bare custom aliases by endpoint 2026-09-07 07:06:58 -07:00
fangliquanflq
3fcbf13ddb fix(agent): scope named custom health by endpoint 2026-09-07 07:06:58 -07:00
fangliquanflq
b40998bc3c fix(agent): isolate custom endpoint billing health 2026-09-07 07:06:58 -07:00
Teknium
27f32bd50b test: exercise output-cap removal across native and child surfaces 2026-09-07 06:15:43 -07:00