Commit Graph

1929 Commits

Author SHA1 Message Date
kshitijk4poor
8a949659c3 refactor(credits): fold review findings for stealth free-tier fix
- credits_tracker: trim inline comment block (duplicated docstring) and
  correct its safety claim - a paid model under stealth/ would fail
  closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
  recognize stealth/ prefix (same bug class as #91843: free_only=true
  wrongly skipped the OpenRouter fallback and the paid-lane warning
  fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
  my-stealth/model not).
2026-08-22 04:20:48 +05:30
kshitijk4poor
ad96d2e2d9 fix(credits): suppress depleted banner on stealth-preview models
Stealth-preview SKUs (e.g. stealth/ox-alpha) are free-tier but carry no
:free suffix, so is_free_tier_model() returned False for them.  On gateway
sessions (which never run the model picker's pricing fetch), the free-model
suppression of the credits.depleted banner never engaged, and any response
carrying paid_access:false triggered a false "Credit access paused" notice.

Add stealth/ prefix detection to is_free_tier_model() as a zero-network
signal, same design as the existing :free suffix check.  Fail-open to
False (banner still shows) if the prefix changes — recoverable noise,
never a masked depletion on a paid model.

Closes #91843
2026-08-22 04:20:48 +05:30
Teknium
334bcbac93 fix(desktop): error card honors the classifier's retry verdict + failing-session identity (review feedback)
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
  verdict) next to failure_reason; error_surface prefers it and only falls
  back to the reason set for older results. Fallback set corrected to match
  classify_api_error (auth, format_error, billing_unverified now
  non-retryable).
- The descriptor carries the failing session's provider/model captured at
  classification time; Copy error details prefers them over the foreground
  composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
  the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
  requests/aiohttp so other adapter SDKs don't misclassify as gateway.
2026-08-21 15:24:03 -07:00
Teknium
98f6fc549a feat(desktop): failed turns name the failing layer with recovery actions
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.

Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
2026-08-21 15:24:03 -07:00
Teknium
42dd219f46 test: trim salvage of #65076 to a lean regression set
Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
2026-08-21 15:02:29 -07:00
Vinay Shah
41ca67c5b1 fix(bedrock): align auxiliary region resolution with runtime + document Mantle route
Address review feedback on #65076:

- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
  config-first region resolution (bedrock.region in config.yaml, then
  AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
  runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
  branch) to the new helper. Previously it derived its region with bare
  resolve_bedrock_region() (env-first), so when config.yaml pinned
  bedrock.region to a different region than the ambient AWS env, auxiliary
  calls (compression, memory, vision) left the primary runtime's region.
  Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
  Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
  for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
  uses the OpenAI-compatible endpoint, which the Mantle route made stale.
  Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
  Converse), the Mantle auth model (bearer token or SigV4), and add the
  GPT-5.5/5.6 model IDs to the models table.
2026-08-21 15:02:29 -07:00
Vinay Shah
16476fad10 feat(bedrock): add OpenAI GPT-5.6 family (Sol/Terra/Luna) to Mantle Responses routing
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:

- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
  so runtime resolution, auxiliary calls, and MoA slots all take the
  SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
  Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
  BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
  additions do not require test surgery; add routing, picker, and
  context-length coverage for the 5.6 family.

Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
2026-08-21 15:02:29 -07:00
Nathaniel Branscum
e5b96fcb10 feat(bedrock): support OpenAI Responses models
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
2026-08-21 15:02:29 -07:00
unsupportedpastels
ac8dff4fbc fix(compression): auto-raise Daybreak Codex threshold 2026-08-21 13:19:07 -07:00
unsupportedpastels
f8e5949f61 fix(model_metadata): add Daybreak Codex 900K context 2026-08-21 13:19:07 -07:00
abundantbeing
6b7aee2f80 fix(agent): preserve live merged tool-call carriers
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
2026-08-21 21:52:55 +05:30
abundantbeing
fdf0111410 fix(agent): guard merged assistant compaction handoffs
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.

Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.

Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
2026-08-21 21:52:55 +05:30
Troy Rowe
6e53628338 fix(classifier): retry provider-injected parameter 400s instead of aborting
The Codex OAuth backend (chatgpt.com/backend-api/codex) intermittently
injects prompt_cache_retention into its own upstream call and then rejects
it, returning HTTP 400 invalid_parameter. Hermes never sends that field on
this route (see agent/transports/codex.py::_default_prompt_cache_retention_
for_request, which only sets it for api.meta.ai and bedrock-mantle hosts).

Reproduced live: a minimal 1-message request carrying no cache parameters
at all failed 4/20 (20%) with this error, so the rejection is not
deterministic and retrying the identical request is the correct recovery.

Previously the catch-all in _classify_400 returned format_error/
retryable=False, which tripped the is_client_error abort gate in
conversation_loop and killed the turn on the first attempt - burning an
entire large-context request (~550k tokens) per failure.

Classify these as retryable server_error (should_compress=False - the
request shape was never the problem). The same guard is applied to the
sibling 5xx request-validation branch, where a fronting proxy can surface
the identical rejection.

Deliberately narrow: keyed on parameters we only send on specific routes,
and skipped when the current provider is one that legitimately sends them,
so a genuine client-side bad parameter (max_tokens on GPT-5) still fails
fast as a format_error.
2026-08-21 21:43:39 +05:30
Adolanium
fc9cbc872d fix(cron): do not load MEMORY.md into scheduled jobs
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.

Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
2026-08-21 13:24:43 +05:30
Teknium
ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal
28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
kshitijk4poor
8e77d03188 test(prompt_caching): home empty-tools fallback test in TestPromptCachePlan
It exercises build_prompt_cache_plan's direct_native_tool_cache fallback,
not repeated apply on pre-decorated input, so it belongs with the other
plan-layout tests rather than in TestApplyIdempotency.
2026-08-21 12:23:47 +05:30
kshitijk4poor
c26357ad6a refactor(prompt_caching): shallow strip copy, exact-count guards, dedupe idempotency tests
Follow-up to the #90972 salvage:

- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
  is copy-on-write on content parts by contract (pops the top-level key,
  rebuilds content lists/part dicts fresh), so a shallow top-level copy
  preserves the caller-non-mutation guarantee — verified for all four
  marker shapes — and removes the redundant second deepcopy the re-mark
  path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
  tests/agent/test_prompt_caching.py (where this module's tests live) as
  TestApplyIdempotency; dropped the three tests that duplicated existing
  coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
  test_copies_sections_and_keeps_canonical_tools_plain which already
  asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
  never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
  no-tools fallback pins == 3 (marker loss can no longer masquerade as
  safety); added the one new _can_carry_marker assertion (native=True
  empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
  where part-dict aliasing is the risk); fails on pre-fix base with
  marker accumulation (9 > 4), passes with the fix.
2026-08-21 12:23:47 +05:30
Ben Barclay
efb6b40f94 Merge pull request #91237 from NousResearch/fix/relay-env-exclusive-messaging
fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
2026-08-21 14:33:29 +10:00
Ben Barclay
f8c3635d40 fix(gateway): classify relay routing stamps as process-global deployment config
sol-reviewer round-2 IMPORTANT: relay env vars had no scope
classification, so the two readers disagreed under a multiplexed
profile scope — gateway/config.py (scope-aware getenv) dropped a
process-env GATEWAY_RELAY_URL during the scoped runner reload while
gateway/relay's relay_url()/register_relay_adapter()/self-provision
(direct os.environ reads) still saw it. Result: adapter registered but
Platform.RELAY absent from config, so the connect loop never dialed
and direct adapters stayed up. The inverse split (profile-only stamp:
config enables RELAY, registration finds no URL) was equally dead.

GATEWAY_RELAY_* ROUTING stamps (URL, ENDPOINT, ALLOW_DIRECT_PLATFORMS,
PLATFORMS, BOT_IDS, ROUTE_KEYS, INSTANCE_ID, WAKE_URL, DISPLAY_NAME)
are now in _GLOBAL_ENV_EXACT: deployment config read from os.environ
under any scope, exactly like the API_SERVER listener settings
(#69379), so every reader resolves the same value. Relay AUTH material
(SECRET, ID, DELIVERY_KEY, IDP_*) is deliberately NOT global — it
stays profile-scoped with the fail-closed multiplex guard, mirroring
the non-secret/secret line the terminal env blocklist already draws
(tools/environments/local.py).

The round-1 multiplex regression test asserted the now-rejected
semantic (profile-scoped stamps win); it is inverted to pin the
global-stamp contract: a process-env stamp survives the profile scope
(sweep runs, matching registration), and a profile-only .env stamp
does NOT activate relay.
2026-08-21 13:21:46 +10:00
Teknium
a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Teknium
7c9285aa14 fix(memory): bound invalid-target error and restore recovery hint
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
2026-08-20 20:20:23 -07:00
kshitijk4poor
2cf7b36e11 fix(memory): enforce independent built-in store permissions
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
2026-08-20 20:20:23 -07:00
kshitijk4poor
c809d964d4 fix(memory): parse boolean config values consistently
Use Hermes's shared truthy-value parser so quoted false memory flags disable both built-in stores as expected.
2026-08-20 20:20:23 -07:00
kshitijk4poor
5a5d6b966d fix(memory): unify store flags and bypass stale availability cache
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
2026-08-20 20:20:23 -07:00
Brooklyn Nicholson
c670464ca7 fix(anthropic): send an explicit thinking disable on the native Messages wire
Adaptive Claude models think by default, so omitting the `thinking`
parameter left thinking ON for users who had turned it off. Send
`thinking: {"type": "disabled"}` instead, and keep the omission for
reasoning-mandatory families that answer a disable with HTTP 400.
2026-08-20 12:29:11 -05:00
Teknium
761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
kshitijk4poor
cef999c5cf refactor(agent): consolidate uncompressed-overflow guard to one warn site + re-arm
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
  unconditionally computed request_pressure_tokens (zero marginal cost,
  covers turn-start AND mid-turn growth) — drops the duplicate every-turn
  estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
  under the window, so warn -> /compress -> regrow warns again (the dedup
  was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
  len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
  ~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
  sites (context_compressor always exists; its context_length property
  hard-floors positive; the fallback would have been a synchronous
  network probe mid-turn that also bypassed config overrides) and the
  undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
  fake reimplementation left them uncovered) and add dedup, re-arm,
  no-rearm-while-over, and multimodal-gate coverage.
2026-08-20 12:35:20 +05:30
joaomarcos
db5d5dffea fix(agent): guard against uncompressed session overflow when compression is disabled (#89297)
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.

Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
2026-08-20 12:35:20 +05:30
kshitijk4poor
0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs
fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30
liuhao1024
62016a1b0a fix(compression): count a would-grow refusal as an ineffective strike
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).

Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.

Fixes #88568
2026-08-20 11:36:54 +05:30
HexLab98
a969c5a93d test(memory): cover the disabled built-in memory surface
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).

Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.

The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
2026-08-19 22:59:17 -07:00
Teknium
4999af5cc1 fix: K3 plan-variant slugs (k3-256k) now get K3's effort vocabulary
kimi_supported_efforts() used exact/prefix matching and missed Kimi
Coding plan variants like k3-256k, which fell back to the K2-era
low/medium/high set and mistranslated efforts on a K3 wire. Replaced
with the boundary-token regex from #76427 (credit @ruizanthony), which
matches k3/k3-256k/kimi-k3* without matching kimi-k2.6 or mk3000.
2026-08-19 22:56:04 -07:00
Teknium
7bf66ec39b fix(compression): /compress refusal no longer reports a successful rewrite 2026-08-19 22:41:05 -07:00
Teknium
d0573880a8 fix: Codex Responses effort vocabulary is now per-model — gpt-5.5 no longer 400s on 'max' (#68365 confirmed live)
Live probes against api.openai.com/v1/responses (Aug 2026):
- gpt-5.6: accepts none/low/medium/high/xhigh/max; rejects minimal, ultra
- gpt-5.5: accepts none/low/medium/high/xhigh; rejects max ('Unsupported
  value'), minimal, ultra

So #68365's premise was half right: 'max' does 400 — but only on pre-5.6
models; blanket-clamping max->xhigh on gpt-5.6 (its fix) would have capped
the one model that supports max. The declared-vocabulary design absorbs
this as data: codex_supported_efforts(model) picks CODEX_GPT56_EFFORTS or
CODEX_LEGACY_EFFORTS, and the shared clamp does the rest. Both the main
Codex transport and the auxiliary client's Responses path use it.

Wire outcomes: ultra -> max on gpt-5.6, ultra/max -> xhigh on gpt-5.5/o5,
minimal -> low everywhere.
2026-08-19 22:37:46 -07:00
Brooklyn Nicholson
d39a031329 fix(nous): stop dropping "thinking off" on Portal models that can honor it
reasoning: {enabled: false} is the only shape the Portal honors, and the
profile refused to send it for every model. Sending nothing means the
upstream default instead, which on a thinking-first route like
deepseek/deepseek-v4-pro (catalog: default_effort high) is thinking ON — so
turning thinking off kept billing reasoning tokens on every turn.

The blanket omission was over-broad. The Portal only rejects a disable on
reasoning-mandatory routes ("Reasoning is mandatory for this model"), which
its catalog flags per model, so that flag now gates the omission. Models the
catalog can't speak to keep the old behavior rather than risk the 400.

extra_body.thinking, DeepSeek's own disable shape, is not forwarded upstream
by the Portal and is not an option here.
2026-08-19 22:14:56 -05:00
Jeff Mettel
22f66de638 fix(skills): publish and read the (map, platform) pair under one lock
Review feedback: publishing the map and its platform tag as two separate
global assignments is not atomic. A reader landing between them sees the
NEW map still carrying the OLD tag, and if that stale tag matches its own
platform it accepts the map without rescanning — serving another
platform's disabled-skill view, the leak #14536 closed.

Guard the pair with a module lock. scan_skill_commands publishes both
under it; get_skill_commands resolves its platform first, then reads the
map and tag together under the same lock to make the freshness decision.
Scanning stays outside the lock — it does file I/O and deferred imports,
and concurrent scans are already independent after the local-map change.

get_skill_commands now returns the scan's own completed map rather than
re-reading the global, so a concurrent publish cannot swap the result
between the decision and the return.

Adds a regression test that holds the publish lock and asserts a reader
cannot complete its lookup until it is released.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:41:42 -07:00
Teknium
f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Jeffrey Quesnelle
612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
Teknium
449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium
803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
Teknium
09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00
Teknium
d762ed9b3c feat: execution-discipline guidance now reaches all tool-capable models (config model.execution_guidance)
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.

Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.

The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
  file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
  format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
  broader retry before concluding)
- completion gated on verification (done = every named acceptance
  criterion verified, never a plausible subset)

The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.

Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.

Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).

Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
2026-08-19 16:16:37 -07:00
Teknium
9ec5750aca fix: widen reasoning-effort wire translation to sibling sites (#89503 class)
The chat_completions chokepoint fix (ultra->max for every model,
cherry-picked from #89509) has siblings with the same bug shape:

- codex.py: ultra->max was gated on gpt-5.6 only; now baseline for all
  Responses-API models (backend-specific branches still override).
- Kimi top-level reasoning_effort: K3 accepts low/high/max only —
  'medium' and upper-ladder levels were dropped to the medium default
  (400s on K3, ladder inversion on K2). Full ladder mapped per family,
  mirroring the kimi-coding plugin's K3 map.
- TokenHub: 'minimal' fell through to the 'high' default (asked least,
  got most); full ladder now mapped onto low/medium/high.
- auxiliary_client Responses path: ultra->max alongside the existing
  minimal->low clamp.
- custom provider plugin: ultra capped at max instead of forwarded
  verbatim to GLM/vLLM/SGLang backends that reject it.
- copilot plugin: ad-hoc downgrade rules replaced with the shared
  clamp_reasoning_effort_to_supported ladder walk so ultra/max resolve
  to the strongest supported level instead of medium (#74295).

Sabotage-verified: new sibling-site tests fail 6/10 without the fixes.
2026-08-19 16:04:22 -07:00
liuhao1024
a664429192 fix(agent): cap the ultra reasoning level at the wire vocabulary
Hermes' internal effort vocabulary extends the wire set with ultra
(documented by /reasoning as none..xhigh|max|ultra). OpenAI-compatible
wires — OpenRouter chief among them — accept exactly
max|xhigh|high|medium|low|minimal|none and reject the extension with
HTTP 400, so an ultra configured while the default model was Anthropic
worked (the Anthropic adapter maps its own levels) but leaked
untranslated the moment a per-job override pinned a non-Anthropic
model, failing every call for that job.

The wire-compat chokepoint for this transport previously mapped
ultra to max only for gpt-5.6; generalize the cap to every model.
2026-08-19 16:04:22 -07:00
Falko
8430c1b4da test(codex): cover nested retention entry paths 2026-08-20 03:15:14 +05:30
Falko
8e2949495a fix(codex): strip unsupported cache retention at wire 2026-08-20 03:15:14 +05:30
Brooklyn Nicholson
df9e2251f8 test(agent): give the MoA switch fake a reasoning_echo reader
663fa68cd4 added an unguarded `agent._read_reasoning_echo_from_config()` call
to switch_model's core field swap. The fake agent here carries only the
attributes switch_model touches, so the new call raises AttributeError inside
the rollback-protected block — every field the test asserts on gets restored
to its pre-swap value and all four cases fail on main with
`assert 'opencode-go' == 'moa'`.

Teach the fake the reader, matching the production AIAgent shape.
2026-08-19 13:56:11 -05:00
Yingliang Zhang
73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30