Review findings on the first cut:
* A hosted provider with a stale model.ollama_num_ctx passed the floor for
a 40K model although nothing ever raises a hosted window. The served
window now counts only when the endpoint is local (is_local_endpoint),
the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
compressor clamp ahead of it, which flipped two observables: a
model.context_length above a sub-64K num_ctx was rejected instead of
constructed-and-clamped, and the floor's message reported the clamped
value with advice (set model.context_length) that could not help. The
phase is split: resolution runs before the floor, the clamp
(_clamp_compressor_to_ollama_num_ctx) stays at its original position,
so every case main constructed still constructs with identical
compressor numbers and the floor's message is unchanged.
Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".
num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.
Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review findings on the mirror, all verified live against throwaway items:
* Identity gate. The service name is fixed but the file path honours
CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
item under the service can hold a different login's pair or a newer
rotation. Overwriting it would be the bug in the other direction. The
refresh token that was just POSTed is threaded through
`_write_claude_code_credentials(spent_refresh_token=...)` from both
callers (the singleton refresher and the pool commit) and the mirror
only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
as `0x<HEX> "<echo>"` otherwise. The old regex modelled `\"` escaping
that never happens: a `"` in the account truncated it and the mirror
created a second item; any non-ASCII byte made it return "" and the
mirror silently no-oped. One `find-generic-password -g` call now yields
account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
succeeded when the mirror runs, and a raise here made the refresher
mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
`-w` returns non-ASCII payloads as hex.
Two pre-existing test doubles for the writer accept the new keyword.
Two defects in the cherry-picked mechanism, both found live on macOS:
* `add-generic-password -w` with no value prompts on /dev/tty when a
terminal exists, so the CLI refresh path hung for the 10 s timeout and
wrote nothing; with no terminal it read one line from stdin, hit EOF on
the confirmation read and stored an EMPTY password — bricking the very
item this exists to keep fresh. The command line now goes to
`security -i` on stdin with the payload hex-encoded (`-X`): no argv, no
tty prompt, no quoting of the JSON.
* `-a getpass.getuser()` assumed the item's account is the login user;
`-U` matches on account + service, so a mismatch would have created a
second item Claude Code never reads. The account is read from the
existing item's `acct` attribute and the mirror is skipped without one.
The command builder is a pure function so the no-argv / hex / account
contract is tested on every lane; the Darwin-gated no-entry no-op test is
`macos_only` instead of faking `platform.system`. Tests trimmed to the
invariant bar; the merge test pins that `mcpOAuth` siblings survive.
Live E2E on a throwaway Keychain item seeded with 24 mcpOAuth entries:
triple rotated, scopes/subscriptionType and all 24 siblings preserved,
one item, updated under the seeded account.
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").
_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).
Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).
Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).
With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.
Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).
Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
The independent-account test drives the real load_pool() -> _refresh_entry() path
and asserts the second account POSTs its own refresh token and keeps its principal
(red on base: the row silently became account A). The alias test pins both halves
of the same-account rule: never fall back onto an older singleton, always follow a
newer re-auth.
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).
Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.
Reactive, destination-scoped recovery instead:
- error_classifier: new `FailoverReason.role_alternation` for the vendor
alternation wordings (checked before the request-validation table since
the body also carries `invalid_request_error`); same abort+fallback hints
as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
loop's `_merge_user_content`), `destination_key` (base_url|provider,
model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
adjacent user turns merged, remember the destination on the facade for
the session so later iterations pre-merge, never touch destinations that
accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.
Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.
Fixes#112358
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.
- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
current users). When false:
- `read_claude_code_credentials()` — the only reader of the borrowed Claude
Code login — returns None, so the resolver fallback, the expired-token
refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
file; the pool prunes a `claude_code` row an earlier adopting process
persisted.
- `_recover_codex_tokens_from_cli` returns None for both automatic recovery
paths (rejected refresh, half-empty singleton); the real AuthError is
surfaced instead. The interactive import offer in `hermes auth add
openai-codex` still asks first and is unaffected.
- One INFO line per process the first time adoption would have happened;
`hermes auth list` / `hermes auth status anthropic|openai-codex` print the
same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.
Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.
Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).
Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
test_summarizer_output_think_block_stripped_before_store fed the summarizer output a "## Active Task"
heading and asserted it survived. That only held because grounding could not match the alias and
prepended a second task section next to it (#114479). With the alias now normalised the invariant is:
canonical heading present, alias absent.
Follow-up to the salvaged #114480 (@jonameijers) and #114553 (@JoaoMarcos44).
A small summarizer can emit both the canonical "## Historical Task Snapshot"
and the legacy "## Active Task" heading in one summary (the 4B model in the
issue "duplicates the section set"). With count=1 the grounding pass replaced
only the first match and left the second as a live-looking, undisclaimed
task section. Replace the first task section with the deterministic snapshot
and drop every later one.
Tests: fold the two contributor test files into one file with two invariants
(prompt names the emitted heading; grounding collapses alias + duplicate
sections). The original #114480 test asserted the pre-fix prepend behaviour,
which #114553 makes false.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Align the iterative update instruction with HISTORICAL_TASK_HEADING and
treat leftover ## Active Task sections as the same snapshot so grounding
replaces them instead of prepending a second live-looking task heading.
Fixes#114479
The iterative-update instruction told the summarizer to update
"## Active Task", but the template it is given emits HISTORICAL_TASK_HEADING
("## Historical Task Snapshot"). Leftover from #44454, which renamed the
heading in the template and SUMMARY_PREFIX but not in this literal.
A summarizer that follows the literal emits a section
_HISTORICAL_TASK_SECTION_RE cannot match, so
_ground_historical_task_snapshot prepends instead of replacing and the
summary carries two task sections. Only the grounded one is disclaimed by
SUMMARY_PREFIX, leaving the stale one readable as live work - the hijack
class #44454 closed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
build_converse_kwargs gated cachePoint markers on the raw model id, so a profile ARN wrapping
Claude matched nothing in _CACHE_POINT_PATTERNS and silently lost Bedrock prompt caching.
_model_supports_prompt_cache now resolves the profile through the per-process-cached
_resolve_inference_profile_model_id first; the request still targets the profile ARN.
Part of #114476
The cherry-picked #114482 never resolved a profile in production: the only
caller, agent/model_metadata.py::_resolve_bedrock_context_length, invokes
get_bedrock_context_length(model, probe=False) with no region, and the
resolver was gated on `region`; and it called
get_inference_profile(inferenceProfileId=...) where botocore requires
`inferenceProfileIdentifier`, so even with a region the call raised
ParamValidationError, was swallowed, and the 128k default applied with only
the new warning. Live against a botocore Stubber: 128000 before, 1000000 after.
- resolve in the ARN's own region (field 4), then the passed region, then
the standard AWS chain; the runtime region / base_url may differ
- inferenceProfileIdentifier is the only request parameter GetInferenceProfile has
- match only `application-inference-profile/`: system-defined
`inference-profile/us.anthropic...` ARNs embed the model id and need no call
- no nested-profile recursion: GetInferenceProfile lists foundation-model ARNs
and the wrapped ARN itself satisfies the static-table substring match
- cache both outcomes per process (runs on every context-length resolution),
cleared by reset_client_cache()
- tests trimmed to two invariants on the production call shape; restore the
TestBedrockContextProbe class header the cherry-pick clobbered
- docs: bedrock:GetInferenceProfile IAM permission and the fallback WARNING
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
An application-inference-profile ARN names no model, so the context probe
error text and the static substring table both miss and the 128k default
silently applies: a profile wrapping a 1M-context Sonnet compacts at ~96k.
Resolve the wrapped foundation-model id via GetInferenceProfile (reusing
the control-plane client; an application profile's id IS its ARN, so the
call works on both API shapes), then let the existing probe/table paths
run on that id. Resolution failures (missing bedrock:GetInferenceProfile,
no creds) keep today's behaviour but now log a WARNING naming the
explicit model.context_length escape hatch instead of staying silent.
Fixes#114476
The remediation ("set compression.checkpoint_required: false, or switch
provider") is right only when the gate can never pass with the active provider
set. Attached to every _checkpoint_blocked() it also decorated the transient
"provider checkpoint API v2 failed" path, where the provider IS capable and the
correct move is a retry once the store recovers, and the codex_app_server
refusal, where switching memory provider changes nothing. A small
_checkpoint_incapable() wrapper carries it on the two capability branches only.
Tests: six near-duplicates collapse into three invariants (warn on mismatch,
parametrized over v1 provider / no manager; silent when off or capable; the
compress-time refusal names the flag). Proven red with the fix neutralized.
Consolidates #106879 (KoNit-K) and #106882 (kokhlo); fixes#106870.
When compression.checkpoint_required is set but the active memory provider
lacks checkpoint API v2, keep fail-closed compress and surface a startup
warning plus actionable remediation on the block error.
Co-authored-by: Cursor <cursoragent@cursor.com>
Refusing every apostrophe on the right also protected the possessive, leaking
the raw slug in ordinary prose. Only an apostrophe that does not start 's'
joins (quoted identifiers such as name='hermes-agent' stay). The test counts
rewrites in the caller's block, not the Claude Code prefix.
Two boundary cases from review: excluding every `.` from the right boundary let
"…run by hermes-agent." leak the raw slug, so only a dot FOLLOWED by a word char
(host, extension) now counts as a joiner. And the shipped prompt says
`skill_view(name='hermes-agent')` / "the `hermes-agent` skill" — a quoted slug is
a name the model dereferences, so quote characters join both boundaries. The
regression covers a mailbox, the quoted identifier and the sentence-final dot.
The docs host was one instance of a class: every place the ``hermes-agent`` slug
is joined to an address the model dereferences. Subagent briefs naming
``~/.hermes/hermes-agent/venv/bin/python`` came back as ``~/.hermes/claude-code/...``
(a file that does not exist) and ``NousResearch/hermes-agent`` became a repo
nobody owns. Rewrite the slug only when it stands alone in prose: a lookbehind
rejects ``/ . : @ -`` and word characters on the left, a lookahead rejects
``/ . @ -`` and word characters on the right. Product-name prose ("the
hermes-agent skill") is still rewritten as before.
Path class reported by @chazmaniandinkle and @christophcemper on #48860.
The helper wrote current_block_index for every event and the stop branch
immediately undid it; make it a pure resolver and assign in the start/delta
branches only. The test's replay tail re-covered ordering the sidecar
assertion already pins — dropped.
In the arrival-order fallback (events without contentBlockIndex) a text delta
following contentBlockStop still continued the closed tool's slot and bolted a
stray `text` key onto the toolUse dict — the same mixed block the indexed path
just stopped producing, and one `_replay_ordered_blocks` drops the toolUse from.
Clear the current index at every stop so the next index-less delta starts fresh;
the no-index regression now covers text on both sides of the tool.
Also collapses the dead bool guard (boto3 indices are ints) and trims the
docstring to the one-line WHY.
Claude 5 on Bedrock (global.anthropic.claude-opus-5 observed) killed every
tool-continuation turn that followed an assistant reply containing BOTH text
and a tool call:
ValidationException: This model does not support assistant message prefill.
The conversation must end with a user message.
The payload did end with a user turn. The rejected shape was the assistant
turn before it, replayed from the bedrock_content_blocks sidecar as
assistant[text, text, toolUse, text x16, cachePoint]
i.e. the model's text shredded into one block per streamed delta, with the
toolUse spliced into the middle and one fragment missing.
Cause: stream_converse_with_callbacks() only read contentBlockIndex on
contentBlockStart. Bedrock emits NO contentBlockStart for text blocks (only
for toolUse), so current_block_index stayed None and every text delta was
keyed by len(stream_blocks) — a fresh slot per delta. When the toolUse block
then started at its real index it overwrote one text fragment and sorted into
the middle of the text on replay. Claude 5 rejects text trailing a tool_use in
the last assistant turn as prefill; text-only or toolUse-only replies were
unaffected, which is why the failure looked intermittent. The joined
`content` string was correct, so the transcript looked fine while the wire
payload was not.
Fix: read contentBlockIndex off every contentBlockStart/Delta/Stop and key
stream_blocks by it, so all deltas of a block merge into one entry and blocks
keep Bedrock's order. Events without an index (test doubles, proxies) fall
back to arrival order: a start opens a new slot, a delta/stop continues the
current one.
Verified against Bedrock: replaying the dumped failing request reproduces the
400; the same request with the assistant text ahead of the toolUse (merged or
still fragmented) is accepted; dropping only the cachePoint still fails, so
caching is not involved. The added tests use the captured live event sequence
and fail on main.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The #113719 report asked that, when the request does fail, the error
names the endpoint that was actually contacted. classify_api_error now
appends "(endpoint: <host>)" to auth-class messages when base_url is set
and is not the provider's own route (provider_owns_route is not True),
so a credential posted to a stale model.base_url reads as a wrong
endpoint rather than a bad key. Stock routes and empty base_url keep
the plain message.
The #64507 test already assumed 'no override' without clearing the shell; fold
the delenv into _make_codex_agent (as the reasoning-effort sibling does) and
drop the per-test helper. The explicit-cap test pins reasoning off too, so the
effort floor can never mask a cap regression.
Docstring tunables line, the cap comment and its log hint still described a
120s default ceiling. The large-request test pins reasoning off so the effort
floor cannot mask a cap regression.
Regression pair for the scale-then-recap contradiction: with no TTFB overrides a
>100K-token openai-codex request keeps the 180s no-event cutoff, and an explicit
HERMES_CODEX_TTFB_MAX_SECONDS still bounds it. Taken from PR #104339 (the same
fix proposed with a 180s default cap); the fix itself lands as the earlier #91635.
set/get used only the lineage filter while take_unseen also required
(active = 1 OR compacted = 1), so a rewound row could be reacted to but never
announced; _DISPLAY_META_ROW_SQL now carries the shared _DISPLAY_ACTIVE_CLAUSE.
take_unseen_reactions scans the whole lineage each turn, so it now filters on
json_extract(display_metadata, '$.reactions') in SQL instead of decoding every
metadata-bearing row in Python. Tests share the compacted-lineage fixture.
A display resume materializes the whole compression lineage with row ids
(`get_resume_conversations(include_ancestors=True)`), so the desktop shows —
and lets the user react to — rows that live in an ended parent segment. The
gateway's `session_key` is re-anchored to the continuation after every
compaction, and `set_message_reaction` scoped the row by that exact key, so
every reaction on a pre-compaction message returned None and the desktop
surfaced RPC 4040 "message not found in this session" (#80670: the 802
compacted-row repro; the agent-side `react_to_message` tool hit the same
wall in #108633).
`set_message_reaction` / `get_message_reactions` / `take_unseen_reactions`
now scope by `_resume_lineage_ids(session_id)` — the same set the resume
loads, so an explicit /branch copy still owns only its own rows and an
unrelated session's row stays foreign. The RPC handler and the tool are
unchanged: ownership is decided once, at the row.
Lineage ownership was first identified in #108635 by @KoNit-K (tool path);
the compacted-row half of the unseen-reaction scan is @Liuzikaii's #108542,
cherry-picked ahead of this commit.
Two invariants, both red on the hardcoded "primary" base: the scheduler's and
delegate_task's platforms map to their own context while interactive surfaces stay
primary, and the real bundled supermemory provider switches writes off for a cron
session's kwargs. Distilled from the regression file in #107045.
An OAuth auxiliary provider with no login now raises 'no credentials were found'
(#114405 / #78996); the compressor must classify it as permanent instead of
retrying. Extend the existing missing-credential classifier test with that
wording so removing the marker from _SUMMARY_MISSING_CREDENTIAL_MARKERS goes red.
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.
One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.
Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.
Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
The titler receives the opening message AFTER @-reference expansion, so
the Desktop's generated pasted_content @file: ref arrives with a
'--- Context Warnings ---' (or '--- Attached Context ---') footer. That
footer made the ref-only check in build_title_input fail, so the live
wire capture showed the path + warning leading the title-model input
with 'Pasted content: ...' appended after it. Strip the expander footer
when a preview is present, so a paste-only opener lets the preview lead.
Invariant: test_expanded_paste_ref_footer_does_not_demote_the_preview
(red on a5dac801, green here). Re-checked on the wire against the stub
model: title input now starts with the pasted topic, no @file:/warning.
The PR widened tui_gateway/methods_prompt.py::_run_after_agent_ready with a
display_metadata positional; tests/tui_gateway/test_submit_time_user_row.py
(already on main) still called the six-arg form and failed CI with a
TypeError. Update the call.
Add one turn_context-level invariant: a user message carrying
display_metadata.title_preview reaches maybe_auto_title(title_preview=...).
Red with agent/turn_context.py swapped to origin/main, green on this head;
without it the whole prompt.submit -> display_metadata -> titler seam had
no test.
Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:
- `apply_instant_title` never received the preview, so the instant title of a
paste-only opener was the generated `@file:` path — and stayed that way,
because the upgrade thread's `derive_title` fallback writes `derived`
provenance, which never replaces the `derived` title already stored.
Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
but the generated attachment ref; the preview now leads for a ref-only
opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
manual attachments stay unread).
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.
Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.
Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).
Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
A session already on the fallback (preferred entry benched by another session) that
rotates UP to the preferred entry once its window reopened must not be pulled back
DOWN when the fallback's own cooldown lifts. Compare priorities in _rotate_and_swap
before arming _credential_pool_revert_id; cover the two-session interleaving as the
control case inside the existing positive test.
Reshape of the salvaged fix from #114513 (@whyyagswhy):
- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
or round-robin order. The contributor's version called the private
``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.