MiniMax M2.x reasoning models emit reasoning_content blocks before
their first content token (#17924). During extended thinking phases,
they routinely exceed the default 180s chat-model stale-stream
timeout, causing the stale-stream detector to kill the connection
mid-think.
This adds minimax-m2 to _REASONING_STALE_TIMEOUT_FLOORS with a
300s floor — generous enough to cover the documented 240s stall
(test_streaming.py:1270-1278) with margin.
Fixes#62353
(cherry picked from commit d66b2a75dee78f583edb75b699ba055ffcb93013)
Follow-up to the salvaged #103818 (@fangliquanflq): gpt-6-astra (and its -900k alias) is the
same named-reasoning line as gpt-5.6-sol/-terra/-luna, so it joins the 600s floor; parametrized
cases pin -sol-900k / -astra-900k through the separator anchor and a control asserts the
non-reasoning gpt-4.x / gpt-5.1-chat line still gets the plain defaults.
WHY: sub-10k-token requests sit below the Codex context-size floor, so without a family entry
these models fall to the 90s non-stream / 180s stream stale defaults (#103802). gpt-5.5 is left
on the effort-tier mechanism from #112909 on purpose (its control tests pin medium effort at 90s).
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:
* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
omitted `extra_body.thinking`. The server then defaults to thinking-on, so
the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
V-series regex), so the id a user picked never reached the wire and the
config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
instead of the real 1M window, capping the model at an eighth of its
context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
detector at its 180s default instead of the 600s reasoning-model floor.
Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".
Adds the id to all four gates plus regression coverage for each site.
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:
- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
OpenRouter live /api/v1/models; without this the slug fell through
to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
OpenCode Zen twin slug x-preview-f-free (reasoning model,
long-horizon agentic work per its model card)
- model-catalog.json regenerated
Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
Review-pass findings on the lazy-init deferral:
- No-op guard: the codex app-server usage callback assigns
compressor.context_length on EVERY response (same window each time).
The setter unconditionally invalidated the derived budgets, wiping
runtime corrections applied directly to threshold_tokens /
tail_token_budget (aux-context threshold sync) — those persisted on
main's eager init. Same-value assignment is now a no-op.
- Re-floor on genuinely new window: the setter invalidates budgets but
previously kept the stale threshold_percent, so a codex window switch
recomputed threshold_tokens from the new window with the old model's
floored percent. Re-apply the raise-only small-context floor from
_base_threshold_percent so percent and tokens derive from the same
window (guarded with getattr for object.__new__ test instances).
- Init log extracted to _emit_init_summary_once() and also fired from
the setter path, so a consumer assigning context_length before any
read no longer strands the startup line forever.
- threshold_tokens getter resolves the window into a local before
reading threshold_percent — correctness no longer depends on
left-to-right argument evaluation order.
- reasoning_timeouts: fix inaccurate 'tuples are immutable' comment
(the container is a list; safety comes from build-once-at-import),
document why the slug stays in the tuple.
Adds TestContextLengthSetterCoherence (3 tests): same-value assignment
preserves overrides; new-window assignment re-floors both directions.
_match_any() was re-sorting _REASONING_STALE_TIMEOUT_FLOORS (21 elements)
on every call. This function runs per API turn via error_classifier,
chat_completion_helpers, and thinking_timeout_guidance.
Also fixes thread-safety: the old _PATTERN_CACHE was a mutable dict
accessed from multiple threads without locking. Pre-compiling all
patterns at module load time eliminates both the per-call sort and
the TOCTOU race condition. The resulting list is effectively
immutable after import, safe for free-threaded Python 3.13+.
(cherry picked from commit e8b006b853dea8b28725d755f08750022f258c9d)
claude-fable-5 is a Mythos-class reasoning model (1M context, 128K output,
adaptive thinking per anthropic_adapter.py) but was missing from the
_REASONING_STALE_TIMEOUT_FLOORS table. Without a floor entry it got the
default 180s stale timeout (300s with context scaling), which is too short
for fable-5's thinking phase on large contexts.
Each stale kill bumped the cross-turn circuit breaker streak; after 5
consecutive kills _check_stale_giveup() fired immediately (elapsed: 0.00s),
aborting all calls with "Provider has been unresponsive for 5 consecutive
stale attempts." Users with 191K-token contexts hit this reliably.
Add ("claude-fable", 600) — deep-reasoning tier alongside o1/deepseek-r1/
nemotron-3-ultra. The claude-fable slug matches claude-fable-5 and future
variants via the existing right-anchor regex.
Anthropic released Claude Opus 5 (+ -fast variant) — both are live on
OpenRouter and the Nous Portal /models endpoint (verified against both
live APIs). Opus 4.8 entries are kept.
- hermes_cli/models.py: opus-5 + opus-5-fast in OPENROUTER_MODELS;
opus-5 in _PROVIDER_MODELS[nous] (Portal serves both, curated list
carries the base model like the rest of the Nous Anthropic block).
Ordering: below fable-5 flagship, above opus-4.8.
- agent/model_metadata.py: claude-opus-5 -> 1M context (matches live
OpenRouter metadata).
- agent/reasoning_timeouts.py: claude-opus-5 -> 240s stale-timeout
floor (same as the opus-4.x thinking family).
- website/static/api/model-catalog.json: regenerated via
scripts/build_model_catalog.py.
Both providers bill via official_models_api (live pricing), so no
_OFFICIAL_DOCS_PRICING snapshot entry is needed for these routes.
grok-4.5 is xAI's newest release (their versioning is non-monotonic:
4.5 > 4.20) and is the model xAI's own docs use for the server-side
x_search tool. Users who explicitly pinned x_search.model keep their
choice; everyone else picks up the new default via the config
deep-merge — no _config_version bump needed.
- tools/x_search_tool.py: DEFAULT_X_SEARCH_MODEL
- hermes_cli/config.py: DEFAULT_CONFIG x_search.model + comment
- agent/reasoning_timeouts.py: 300s stale-timeout floor entry for
grok-4.5 (grok-4.20-reasoning entry kept for pinned users)
- docs: x-search.md en + zh-Hans (config sample + troubleshooting)
- tests: default-model assertion + timeout-floor positive case
DeepSeek V4 models (deepseek-v4-flash, deepseek-v4-pro) emit
reasoning_content in a separate delta field before final content,
requiring the same 600s stale timeout floor as R1. Without this,
streams hang for 30–50s with APITimeoutError on providers like
opencode-go while direct calls succeed in ~3s.
Fixes#60338.
Two-part fix:
Part 1 (classifier override at agent/error_classifier.py:720-738):
A transport disconnect on a reasoning model — even on a large session —
now routes to FailoverReason.timeout instead of context_overflow. Without
this, large-session reasoning-model disconnects route to the compression
branch and silently delete conversation history on a phantom
context-length error. The override is strictly targeted: non-reasoning
models (gpt-4o, claude-3-5-sonnet, llama-3.3-70b, etc.) still route to
context_overflow on large sessions — the existing intentional behavior
for chat models whose proxy doesn't idle-kill during prefill/generation.
Part 2 (new agent/thinking_timeout_guidance.py + integration at
agent/conversation_loop.py:3488-3567):
New is_thinking_timeout() and build_thinking_timeout_guidance() helpers.
When a known reasoning model (NVIDIA Nemotron 3 Ultra, OpenAI o1/o3,
Anthropic Opus 4.x thinking, DeepSeek R1, Qwen QwQ, xAI Grok reasoning)
hits a transport-kill on a small session (classifier says timeout
directly) or after Part 1 routes correctly (large session), the user
now sees reasoning-specific guidance with three actionable workarounds
in priority order:
1. Set providers.<provider>.models.<model>.stale_timeout_seconds: 900
in ~/.hermes/config.yaml (Hermes's built-in floor is already 600s
for known reasoning models; raise further if upstream is even
tighter).
2. Lower reasoning_budget or set reasoning_effort: medium on this
model if the provider supports it.
3. Use a smaller / faster reasoning model if the task doesn't
require deep thinking.
The new guidance takes precedence via if/elif over the existing
_is_stream_drop block, so a reasoning-model user with a transport-kill
message sees actionable advice instead of the misleading "try
execute_code with Python's open() for large files" advice (which is
correct for the unrelated large-file-write stream-drop case but
actively wrong for the thinking-timeout case).
Verified:
- 478 tests passing across 9 directly-relevant files (49 new + 429
existing, zero regressions).
- Ruff lint clean on all 4 modified/new files.
- Negative test: 6 parametrized regression guards confirm non-reasoning
models still route to context_overflow on large sessions; 4
parametrized gates confirm non-timeout classifier reasons never
trigger the guidance; 5 parametrized cases confirm non-transport
messages never trigger it.
- Regression guard: new guidance message does NOT contain
"execute_code" or "open()" — the misleading advice is fully
replaced, not appended alongside.
- Cross-vendor dual review via agy -p:
- Gemini 3.5 Flash (Medium) — passed: true, zero blockers, one
SHOULD-FIX (vprint block duplication — fixed by extracting
detection into a helper module).
- GPT-OSS 120B (Medium) — passed: true, zero blockers, two nits
(test placement — adopted at tests/agent/test_thinking_timeout_guidance.py;
primary-model capture — accepted as non-issue per Flash's nit).
Dependency note for maintainers:
This PR includes agent/reasoning_timeouts.py (the reasoning-model
allowlist module from PR #52238) because the Layer 1 override is
load-bearing on get_reasoning_stale_timeout_floor(). After PR #52238
lands on main, this PR's duplicate agent/reasoning_timeouts.py should
be rebased away. Either PR can land first; the other rebase is
mechanical.
Fixes#52271.