Commit Graph

5161 Commits

Author SHA1 Message Date
kshitijk4poor
cf13448580 refactor(agent): name the tool-call XML namespace prefix once
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
2026-09-19 10:41:08 +05:30
DoGMaTiiC
60735cbf81 fix(agent): strip namespace-prefixed text-channel tool-call XML from visible text
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).

Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).

With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.

Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).

Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
2026-09-19 10:41:08 +05:30
Tranquil-Flow
20512a47a6 fix(agent): refuse stale singleton adoption over rotated manual:device_code entries (#106705)
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).

Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
2026-09-18 20:57:07 -07:00
Marc Caelier
46503f1672 fix(auth): isolate Codex singleton sync by principal 2026-09-18 20:57:07 -07:00
teknium1
76724f3dc2 fix(agent): terminal copy for the new role_alternation failure reason
Reclassifying alternation 400s away from format_error left non-MoA callers
on the generic default text.
2026-09-18 20:56:35 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
5e95050608 fix(agent): a server context rejection the transcript cannot explain is no longer "conversation too long"
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).

_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.

Fixes #114644
2026-09-18 19:18:57 -07:00
teknium1
a6ad3cc3ff fix(review): derive the default input budget from the fork's resolved context window
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.

Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).

Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
2026-09-18 19:18:57 -07:00
KoNit-K
f797a23c09 fix(review): cap default background input budget 2026-09-18 19:18:57 -07:00
teknium1
dbf19a6cfa fix(compression): collapse duplicate task sections during snapshot grounding
Follow-up to the salvaged #114480 (@jonameijers) and #114553 (@JoaoMarcos44).

A small summarizer can emit both the canonical "## Historical Task Snapshot"
and the legacy "## Active Task" heading in one summary (the 4B model in the
issue "duplicates the section set"). With count=1 the grounding pass replaced
only the first match and left the second as a live-looking, undisclaimed
task section. Replace the first task section with the deterministic snapshot
and drop every later one.

Tests: fold the two contributor test files into one file with two invariants
(prompt names the emitted heading; grounding collapses alias + duplicate
sections). The original #114480 test asserted the pre-fix prepend behaviour,
which #114553 makes false.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 19:18:23 -07:00
joaomarcos
d295ef02ee fix(compression): replace alias task headings during snapshot grounding
Align the iterative update instruction with HISTORICAL_TASK_HEADING and
treat leftover ## Active Task sections as the same snapshot so grounding
replaces them instead of prepending a second live-looking task heading.

Fixes #114479
2026-09-18 19:18:23 -07:00
openclaw
4ab7dec756 fix(compaction): name the emitted task heading in the update instruction
The iterative-update instruction told the summarizer to update
"## Active Task", but the template it is given emits HISTORICAL_TASK_HEADING
("## Historical Task Snapshot"). Leftover from #44454, which renamed the
heading in the template and SUMMARY_PREFIX but not in this literal.

A summarizer that follows the literal emits a section
_HISTORICAL_TASK_SECTION_RE cannot match, so
_ground_historical_task_snapshot prepends instead of replacing and the
summary carries two task sections. Only the grounded one is disclaimed by
SUMMARY_PREFIX, leaving the stale one readable as live work - the hijack
class #44454 closed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 19:18:23 -07:00
teknium1
a51143fbbe fix(bedrock): resolve application-inference-profile ARNs before the prompt-cache allowlist match
build_converse_kwargs gated cachePoint markers on the raw model id, so a profile ARN wrapping
Claude matched nothing in _CACHE_POINT_PATTERNS and silently lost Bedrock prompt caching.
_model_supports_prompt_cache now resolves the profile through the per-process-cached
_resolve_inference_profile_model_id first; the request still targets the profile ARN.

Part of #114476
2026-09-18 15:15:02 -07:00
teknium1
11445cc54e fix(bedrock): application inference profile ARNs size from the wrapped model on the real path
The cherry-picked #114482 never resolved a profile in production: the only
caller, agent/model_metadata.py::_resolve_bedrock_context_length, invokes
get_bedrock_context_length(model, probe=False) with no region, and the
resolver was gated on `region`; and it called
get_inference_profile(inferenceProfileId=...) where botocore requires
`inferenceProfileIdentifier`, so even with a region the call raised
ParamValidationError, was swallowed, and the 128k default applied with only
the new warning. Live against a botocore Stubber: 128000 before, 1000000 after.

- resolve in the ARN's own region (field 4), then the passed region, then
  the standard AWS chain; the runtime region / base_url may differ
- inferenceProfileIdentifier is the only request parameter GetInferenceProfile has
- match only `application-inference-profile/`: system-defined
  `inference-profile/us.anthropic...` ARNs embed the model id and need no call
- no nested-profile recursion: GetInferenceProfile lists foundation-model ARNs
  and the wrapped ARN itself satisfies the static-table substring match
- cache both outcomes per process (runs on every context-length resolution),
  cleared by reset_client_cache()
- tests trimmed to two invariants on the production call shape; restore the
  TestBedrockContextProbe class header the cherry-pick clobbered
- docs: bedrock:GetInferenceProfile IAM permission and the fallback WARNING

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 15:15:02 -07:00
liuhao1024
4c28be68c2 fix(bedrock): size inference-profile ARNs from the wrapped foundation model
An application-inference-profile ARN names no model, so the context probe
error text and the static substring table both miss and the 128k default
silently applies: a profile wrapping a 1M-context Sonnet compacts at ~96k.
Resolve the wrapped foundation-model id via GetInferenceProfile (reusing
the control-plane client; an application profile's id IS its ARN, so the
call works on both API shapes), then let the existing probe/table paths
run on that id. Resolution failures (missing bedrock:GetInferenceProfile,
no creds) keep today's behaviour but now log a WARNING naming the
explicit model.context_length escape hatch instead of staying silent.

Fixes #114476
2026-09-18 15:15:02 -07:00
teknium1
96030a3536 fix(compression): scope the checkpoint remediation to capability refusals; trim tests
The remediation ("set compression.checkpoint_required: false, or switch
provider") is right only when the gate can never pass with the active provider
set. Attached to every _checkpoint_blocked() it also decorated the transient
"provider checkpoint API v2 failed" path, where the provider IS capable and the
correct move is a retry once the store recovers, and the codex_app_server
refusal, where switching memory provider changes nothing. A small
_checkpoint_incapable() wrapper carries it on the two capability branches only.

Tests: six near-duplicates collapse into three invariants (warn on mismatch,
parametrized over v1 provider / no manager; silent when off or capable; the
compress-time refusal names the flag). Proven red with the fix neutralized.

Consolidates #106879 (KoNit-K) and #106882 (kokhlo); fixes #106870.
2026-09-18 13:58:52 -07:00
KoNit-K
8e7120e2fd fix(compression): debug-log checkpoint capability probe failures
Keep init fail-open on a broken probe, but emit DEBUG so flaky probes are diagnosable.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
KoNit-K
ef7c7f9e08 fix(compression): warn and remediate opaque checkpoint_required blocks
When compression.checkpoint_required is set but the active memory provider
lacks checkpoint API v2, keep fail-closed compress and surface a startup
warning plus actionable remediation on the block error.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
kshitijk4poor
9f5098c36c refactor(anthropic): quote characters only mark an identifier from the left
A quoted slug never matches because the lookbehind rejects the opening quote,
so the quote classes in the lookahead were dead; drop them.
2026-09-19 01:44:10 +05:30
kshitijk4poor
ae62e832cc fix(anthropic): a possessive 'hermes-agent's' is prose and is still rewritten
Refusing every apostrophe on the right also protected the possessive, leaking
the raw slug in ordinary prose. Only an apostrophe that does not start 's'
joins (quoted identifiers such as name='hermes-agent' stay). The test counts
rewrites in the caller's block, not the Claude Code prefix.
2026-09-19 01:44:10 +05:30
kshitijk4poor
f7676e7ff1 fix(anthropic): quoted slugs stay identifiers; a sentence-final dot is prose
Two boundary cases from review: excluding every `.` from the right boundary let
"…run by hermes-agent." leak the raw slug, so only a dot FOLLOWED by a word char
(host, extension) now counts as a joiner. And the shipped prompt says
`skill_view(name='hermes-agent')` / "the `hermes-agent` skill" — a quoted slug is
a name the model dereferences, so quote characters join both boundaries. The
regression covers a mailbox, the quoted identifier and the sentence-final dot.
2026-09-19 01:44:10 +05:30
kshitijk4poor
75b6281394 fix(anthropic): OAuth slug rewrite skips paths, repo slugs and mailboxes too
The docs host was one instance of a class: every place the ``hermes-agent`` slug
is joined to an address the model dereferences. Subagent briefs naming
``~/.hermes/hermes-agent/venv/bin/python`` came back as ``~/.hermes/claude-code/...``
(a file that does not exist) and ``NousResearch/hermes-agent`` became a repo
nobody owns. Rewrite the slug only when it stands alone in prose: a lookbehind
rejects ``/ . : @ -`` and word characters on the left, a lookahead rejects
``/ . @ -`` and word characters on the right. Product-name prose ("the
hermes-agent skill") is still rewritten as before.

Path class reported by @chazmaniandinkle and @christophcemper on #48860.
2026-09-19 01:44:10 +05:30
Gaurav Saxena
e218f3641c fix(anthropic): preserve docs URL in OAuth sanitizer 2026-09-19 01:44:10 +05:30
kshitijk4poor
e5af917aaa refactor(bedrock): block_index resolves, callers own the current-index state
The helper wrote current_block_index for every event and the stop branch
immediately undid it; make it a pure resolver and assign in the start/delta
branches only. The test's replay tail re-covered ordering the sidecar
assertion already pins — dropped.
2026-09-19 01:35:23 +05:30
kshitijk4poor
5d4a206c92 fix(bedrock): index-less text after a tool stop opens its own block
In the arrival-order fallback (events without contentBlockIndex) a text delta
following contentBlockStop still continued the closed tool's slot and bolted a
stray `text` key onto the toolUse dict — the same mixed block the indexed path
just stopped producing, and one `_replay_ordered_blocks` drops the toolUse from.
Clear the current index at every stop so the next index-less delta starts fresh;
the no-index regression now covers text on both sides of the tool.

Also collapses the dead bool guard (boto3 indices are ints) and trims the
docstring to the one-line WHY.
2026-09-19 01:35:23 +05:30
EricN
8b28df7b1f fix(bedrock): key streamed content blocks by contentBlockIndex, not arrival count
Claude 5 on Bedrock (global.anthropic.claude-opus-5 observed) killed every
tool-continuation turn that followed an assistant reply containing BOTH text
and a tool call:

  ValidationException: This model does not support assistant message prefill.
  The conversation must end with a user message.

The payload did end with a user turn. The rejected shape was the assistant
turn before it, replayed from the bedrock_content_blocks sidecar as

  assistant[text, text, toolUse, text x16, cachePoint]

i.e. the model's text shredded into one block per streamed delta, with the
toolUse spliced into the middle and one fragment missing.

Cause: stream_converse_with_callbacks() only read contentBlockIndex on
contentBlockStart. Bedrock emits NO contentBlockStart for text blocks (only
for toolUse), so current_block_index stayed None and every text delta was
keyed by len(stream_blocks) — a fresh slot per delta. When the toolUse block
then started at its real index it overwrote one text fragment and sorted into
the middle of the text on replay. Claude 5 rejects text trailing a tool_use in
the last assistant turn as prefill; text-only or toolUse-only replies were
unaffected, which is why the failure looked intermittent. The joined
`content` string was correct, so the transcript looked fine while the wire
payload was not.

Fix: read contentBlockIndex off every contentBlockStart/Delta/Stop and key
stream_blocks by it, so all deltas of a block merge into one entry and blocks
keep Bedrock's order. Events without an index (test doubles, proxies) fall
back to arrival order: a start opens a new slot, a delta/stop continues the
current one.

Verified against Bedrock: replaying the dumped failing request reproduces the
400; the same request with the assistant text ahead of the toolUse (merged or
still fragmented) is accepted; dropping only the cachePoint still fails, so
caching is not involved. The added tests use the captured live event sequence
and fail on main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 01:35:23 +05:30
teknium1
b7b203cda0 fix(skills): built-in name collisions show a note in /skills, /help skills and the palette
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).

Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
  "slash command /<name> unavailable — name taken by built-in; use /skill <name>"

- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
  `hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
  skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
  carries the notice in its existing `warning` field (rendered by the Desktop
  `/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
  listing (Telegram/Discord/…) appends one ⚠ line per colliding skill

Fixes #113560
2026-09-18 12:51:22 -07:00
teknium1
7530f40383 fix(errors): auth refusals from a non-stock route name the contacted host
The #113719 report asked that, when the request does fail, the error
names the endpoint that was actually contacted. classify_api_error now
appends "(endpoint: <host>)" to auth-class messages when base_url is set
and is not the provider's own route (provider_owns_route is not True),
so a credential posted to a stale model.base_url reads as a wrong
endpoint rather than a bad key. Stock routes and empty base_url keep
the plain message.
2026-09-18 12:49:40 -07:00
kshitijk4poor
e3bde9d9c5 docs(agent): TTFB cap is opt-in — say so where operators read it
Docstring tunables line, the cap comment and its log hint still described a
120s default ceiling. The large-request test pins reasoning off so the effort
floor cannot mask a cap regression.
2026-09-19 01:05:09 +05:30
fangliquanflq
aa7af52a85 fix(agent): preserve scaled codex ttfb timeout 2026-09-19 01:05:09 +05:30
RelaxJonh
6f305f3dc1 fix: derive agent_context from platform for memory provider context-skip
agent_context was hardcoded to "primary" in init_agent(), making the
provider context-skip logic (cron/flush/subagent) dead code. Memory
providers like supermemory and honcho check agent_context to decide
whether to skip writing to memory for cron/flush/subagent sessions.

Derive agent_context from the platform parameter: "cron" for cron jobs,
"subagent" for subagent sessions, "primary" otherwise.

Fixes #80646
2026-09-19 00:55:05 +05:30
liuhao1024
a0562d17d8 fix(auth): missing-credential hints name the real env var or the OAuth login (#114405, #78996)
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.

One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.

Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.

Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
2026-09-18 11:11:44 -07:00
teknium1
85f5f560c3 fix(agent): a paste-only opener's expansion footer must not demote the title preview
The titler receives the opening message AFTER @-reference expansion, so
the Desktop's generated pasted_content @file: ref arrives with a
'--- Context Warnings ---' (or '--- Attached Context ---') footer. That
footer made the ref-only check in build_title_input fail, so the live
wire capture showed the path + warning leading the title-model input
with 'Pasted content: ...' appended after it. Strip the expander footer
when a preview is present, so a paste-only opener lets the preview lead.

Invariant: test_expanded_paste_ref_footer_does_not_demote_the_preview
(red on a5dac801, green here). Re-checked on the wire against the stub
model: title input now starts with the pasted topic, no @file:/warning.
2026-09-18 10:56:00 -07:00
teknium1
34ba61bf67 fix(agent): paste title hint reaches the instant title and the prompt.submit contract
Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:

- `apply_instant_title` never received the preview, so the instant title of a
  paste-only opener was the generated `@file:` path — and stayed that way,
  because the upgrade thread's `derive_title` fallback writes `derived`
  provenance, which never replaces the `derived` title already stored.
  Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
  but the generated attachment ref; the preview now leads for a ref-only
  opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
  TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
  manual attachments stay unread).
2026-09-18 10:56:00 -07:00
KoNit-K
8d25e69b6e fix(desktop): use generated paste previews for titles 2026-09-18 10:56:00 -07:00
teknium1
1b083b85a3 fix(checkpoints): drop the defence layers around the never-raising backend predicate
`container_backend_for_task` already catches everything inside `_terminal_env_type_for_task`
and `_uses_container_paths` and returns defaults, and every CheckpointManager has
`unsupported_backend_reason`: call the method directly (no getattr fallback), drop the
try/except in `unsupported_backend_reason` and the suppress() around the import+call in
turn_explainers (the one around record_agent_write stays). The rollback.restore fakes in the
server tests gain the method they now must provide.
2026-09-18 10:38:08 -07:00
teknium1
dac4c60fbb fix(checkpoints): refuse host rollback and session diff from container sessions on the gateway too
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.

Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-18 10:38:08 -07:00
Sora-bluesky
3220b9ed2f fix(checkpoints): classify the backend on every /rollback instead of remembering it
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
2026-09-18 10:38:08 -07:00
Sora-bluesky
10ee9819f9 fix(checkpoints): do not feed container paths into host checkpoint storage
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.

Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).

Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
2026-09-18 10:38:08 -07:00
teknium1
2d333e68b2 fix(compression): read the aux compression budget unguarded in resolve_context_compression_timeouts
_effective_aux_timeout is a pure config read used unguarded elsewhere; a
swallowed failure would silently fall back to the 120s idle window this
change calls structurally wrong.
2026-09-18 10:36:11 -07:00
teknium1
303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00
joaomarcos
f2bd6b4431 fix(compression): host idle watchdog never undercuts the aux summary request budget
The owned compress_context wait judged silence at compression.context_timeout_seconds (120s)
while the auxiliary compression request itself is allowed _effective_aux_timeout("compression")
(floor 300s). A summariser that legitimately produces no streamed token for >120s (non-streaming
route, reasoning model thinking before its first delta, long prompt build over thousands of rows)
was always cancelled as "no progress" before the provider layer would have given up.

Raise-only: the idle window is floored at the aux compression budget (never above the ceiling,
which is itself raised when the aux budget exceeds it); an explicit larger value is kept.

Salvaged from PR #114635 (@JoaoMarcos44); the emergency-prune hunk of that PR is dropped because
prune_tool_results_only is gated on proactive_prune_tokens (0 = off by default).
2026-09-18 10:36:11 -07:00
teknium1
4c5a70065c fix: arm the pool revert only when the benched credential outranks the one rotated to
A session already on the fallback (preferred entry benched by another session) that
rotates UP to the preferred entry once its window reopened must not be pulled back
DOWN when the fallback's own cooldown lifts. Compare priorities in _rotate_and_swap
before arming _credential_pool_revert_id; cover the two-session interleaving as the
control case inside the existing positive test.
2026-09-18 10:35:37 -07:00
teknium1
92bb5b92b8 fix: revert a quota-benched credential through the pool, not a private probe; cover /model and chained rotations
Reshape of the salvaged fix from #114513 (@whyyagswhy):

- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
  benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
  refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
  or round-robin order. The contributor's version called the private
  ``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
  and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
  429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
  the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
  the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
  by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
  hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
  ``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
  fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
  and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
2026-09-18 10:35:37 -07:00
Yagna Vudathu
0dd153ba2e fix(agent): revert pool rotation after transient quota cooldown expires (#114501) 2026-09-18 10:35:37 -07:00
teknium1
8f0322da5b fix(agent): close a failed turn's durable user tail so the next prompt is not merged into it
Terminal-failure paths (HTTP-200 content-policy refusal, ``_Trunc.end_turn``, retry
exhaustion, interrupt before any assistant text) persist the accepted user row and return
before ``finalize_turn``, so ``user`` stays the durable conversation tail. The next prompt
appends a second user row, ``repair_message_sequence`` merges the pair, and the provider
is asked to act on the failed request again. The gateway compensates with
``_hmwa_close_failed_turn`` (#108033); standalone ACP, the CLI and the TUI/Desktop hand
``result["messages"]`` straight back as history and had no closer.

Close it once, at ``agent/conversation_loop.py::run_conversation`` — the seam every
envelope leaves through — with a Hermes-authored assistant boundary
(``agent/turn_failure_copy.py::FAILED_TURN_NOTICE`` / ``PARTIAL_FAILED_TURN_NOTICE``,
which the gateway now aliases instead of keeping its own copy). Idempotence is keyed on
``SessionDB.latest_conversation_role`` (durable state, not content), so a redelivery or a
tail another writer already closed is a no-op and the gateway's closer no-ops in turn.
The context-pressure classes (``compression_exhausted``, ``compression_deferred``,
``failure_reason == "context_overflow"``) are excluded: appending to an oversized session
is the #1630 growth loop; their repair is rotation.

Adjacent defect from the same report: ``acp_adapter/server.py::_finish_turn`` called
``final_response.startswith`` on ``None`` for an interrupted turn — the same one-line fix
PR #64471 by @israellot filed first (its wider prompt()-restructure is superseded by the
current ``_finish_turn`` shape).

Slimmer redo of #114168 by @kendrickkester (same seam and invariants; the +1023-line
PR carried a new copy module, an accepted-turn re-anchoring scan and an 859-line suite).
Two invariant tests: the real ACP path (loopback provider, refusal then a new prompt) and
the durable-tail idempotence / overflow exclusion.

Co-authored-by: Kendrick Kester <kendrick.kester@gmail.com>
Co-authored-by: Israel Lot <israel.lot@gmail.com>
2026-09-18 10:33:21 -07:00
teknium1
1764508c16 fix(tui): project-local skills register and dispatch for the session's repo
TUI/desktop RPC handlers (commands.catalog, command.dispatch, slash.exec's
skill routing guard) run on the socket thread with no session context; the
terminal scope (and gateway.run's import-time bridge) resolve a placeholder
terminal.cwd to $HOME, so project skills under <repo>/.hermes/skills never
registered and `/<name>` failed with "not a quick/plugin/bundle/skill
command" even though the session's workspace was the repo (#114359).

- tui_gateway/methods_tools.py: _session_home_scope also binds the session's
  cwd (runtime_cwd.set_session_cwd) for the block; commands.catalog binds the
  calling session's workspace (_completion_cwd: its record, else the cwd a
  new session would be seeded with) around skill discovery.
- agent/skill_commands.py: the cached registry is tagged by project root too,
  so two sessions in two repos in one process each see their own skills
  instead of the first scan's.
- agent/runtime_cwd.py: reset_session_cwd(token) twin for set_session_cwd.

The multiplex guarantee from 9c9e7ab6e5 is untouched: a served profile's
turn still resolves ITS scope cwd, and the process cwd is only ever the last
fallback; the session cwd that now wins is the session's own workspace.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: sum117 <75037449+sum117@users.noreply.github.com>
2026-09-18 10:22:47 -07:00
liuzikaii
3316a1224b fix(skills): resolve project roots from session cwd 2026-09-18 10:22:47 -07:00
teknium1
b44a481334 fix(video-gen): trust the operator-configured origin on the first download hop
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.

save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.

security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
2026-09-18 10:22:14 -07:00