Commit Graph

2966 Commits

Author SHA1 Message Date
kshitijk4poor
afaa53e5fe fix(agent_init): judge the served window only for local endpoints; clamp after the floor as before
Review findings on the first cut:

* A hosted provider with a stale model.ollama_num_ctx passed the floor for
  a 40K model although nothing ever raises a hosted window. The served
  window now counts only when the endpoint is local (is_local_endpoint),
  the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
  compressor clamp ahead of it, which flipped two observables: a
  model.context_length above a sub-64K num_ctx was rejected instead of
  constructed-and-clamped, and the floor's message reported the clamped
  value with advice (set model.context_length) that could not help. The
  phase is split: resolution runs before the floor, the clamp
  (_clamp_compressor_to_ollama_num_ctx) stays at its original position,
  so every case main constructed still constructs with identical
  compressor numbers and the floor's message is unchanged.

Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
2026-09-19 11:44:15 +05:30
kshitijk4poor
3de7140bcd fix(agent_init): the 64K context floor judges the window Ollama serves, not the GGUF metadata
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".

num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.

Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-19 11:44:15 +05:30
kshitijk4poor
14a3463454 fix(anthropic): mirror only the Keychain item that held the spent pair; parse both security attribute encodings
Review findings on the mirror, all verified live against throwaway items:

* Identity gate. The service name is fixed but the file path honours
  CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
  item under the service can hold a different login's pair or a newer
  rotation. Overwriting it would be the bug in the other direction. The
  refresh token that was just POSTed is threaded through
  `_write_claude_code_credentials(spent_refresh_token=...)` from both
  callers (the singleton refresher and the pool commit) and the mirror
  only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
  is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
  as `0x<HEX>  "<echo>"` otherwise. The old regex modelled `\"` escaping
  that never happens: a `"` in the account truncated it and the mirror
  created a second item; any non-ASCII byte made it return "" and the
  mirror silently no-oped. One `find-generic-password -g` call now yields
  account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
  succeeded when the mirror runs, and a raise here made the refresher
  mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
  `-w` returns non-ASCII payloads as hex.

Two pre-existing test doubles for the writer accept the new keyword.
2026-09-19 11:25:23 +05:30
kshitijk4poor
bf5a6f6ae6 fix(anthropic): mirror the Keychain item through security -i with a hex payload, under its own account
Two defects in the cherry-picked mechanism, both found live on macOS:

* `add-generic-password -w` with no value prompts on /dev/tty when a
  terminal exists, so the CLI refresh path hung for the 10 s timeout and
  wrote nothing; with no terminal it read one line from stdin, hit EOF on
  the confirmation read and stored an EMPTY password — bricking the very
  item this exists to keep fresh. The command line now goes to
  `security -i` on stdin with the payload hex-encoded (`-X`): no argv, no
  tty prompt, no quoting of the JSON.
* `-a getpass.getuser()` assumed the item's account is the login user;
  `-U` matches on account + service, so a mismatch would have created a
  second item Claude Code never reads. The account is read from the
  existing item's `acct` attribute and the mirror is skipped without one.

The command builder is a pure function so the no-argv / hex / account
contract is tested on every lane; the Darwin-gated no-entry no-op test is
`macos_only` instead of faking `platform.system`. Tests trimmed to the
invariant bar; the merge test pins that `mcpOAuth` siblings survive.

Live E2E on a throwaway Keychain item seeded with 24 mcpOAuth entries:
triple rotated, scopes/subscriptionType and all 24 siblings preserved,
one item, updated under the seeded account.
2026-09-19 11:25:23 +05:30
salch-cred
6d4ca15d38 fix(agent): mirror Claude Code OAuth refresh into the macOS Keychain (#98334)
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").

_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).

Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
2026-09-19 11:25:23 +05:30
kshitijk4poor
cf13448580 refactor(agent): name the tool-call XML namespace prefix once
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
2026-09-19 10:41:08 +05:30
DoGMaTiiC
60735cbf81 fix(agent): strip namespace-prefixed text-channel tool-call XML from visible text
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).

Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).

With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.

Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).

Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
2026-09-19 10:41:08 +05:30
teknium1
c1ad9daa8f test: two invariant tests for Codex singleton adoption, replacing the salvaged suites
The independent-account test drives the real load_pool() -> _refresh_entry() path
and asserts the second account POSTs its own refresh token and keeps its principal
(red on base: the row silently became account A). The alias test pins both halves
of the same-account rule: never fall back onto an older singleton, always follow a
newer re-auth.
2026-09-18 20:57:07 -07:00
Tranquil-Flow
20512a47a6 fix(agent): refuse stale singleton adoption over rotated manual:device_code entries (#106705)
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).

Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
2026-09-18 20:57:07 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
a6ad3cc3ff fix(review): derive the default input budget from the fork's resolved context window
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.

Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).

Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
2026-09-18 19:18:57 -07:00
KoNit-K
f797a23c09 fix(review): cap default background input budget 2026-09-18 19:18:57 -07:00
teknium1
ffe69fd711 test(compression): sibling asserts the grounded task heading, not the legacy alias
test_summarizer_output_think_block_stripped_before_store fed the summarizer output a "## Active Task"
heading and asserted it survived. That only held because grounding could not match the alias and
prepended a second task section next to it (#114479). With the alias now normalised the invariant is:
canonical heading present, alias absent.
2026-09-18 19:18:23 -07:00
teknium1
dbf19a6cfa fix(compression): collapse duplicate task sections during snapshot grounding
Follow-up to the salvaged #114480 (@jonameijers) and #114553 (@JoaoMarcos44).

A small summarizer can emit both the canonical "## Historical Task Snapshot"
and the legacy "## Active Task" heading in one summary (the 4B model in the
issue "duplicates the section set"). With count=1 the grounding pass replaced
only the first match and left the second as a live-looking, undisclaimed
task section. Replace the first task section with the deterministic snapshot
and drop every later one.

Tests: fold the two contributor test files into one file with two invariants
(prompt names the emitted heading; grounding collapses alias + duplicate
sections). The original #114480 test asserted the pre-fix prepend behaviour,
which #114553 makes false.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 19:18:23 -07:00
joaomarcos
d295ef02ee fix(compression): replace alias task headings during snapshot grounding
Align the iterative update instruction with HISTORICAL_TASK_HEADING and
treat leftover ## Active Task sections as the same snapshot so grounding
replaces them instead of prepending a second live-looking task heading.

Fixes #114479
2026-09-18 19:18:23 -07:00
openclaw
4ab7dec756 fix(compaction): name the emitted task heading in the update instruction
The iterative-update instruction told the summarizer to update
"## Active Task", but the template it is given emits HISTORICAL_TASK_HEADING
("## Historical Task Snapshot"). Leftover from #44454, which renamed the
heading in the template and SUMMARY_PREFIX but not in this literal.

A summarizer that follows the literal emits a section
_HISTORICAL_TASK_SECTION_RE cannot match, so
_ground_historical_task_snapshot prepends instead of replacing and the
summary carries two task sections. Only the grounded one is disclaimed by
SUMMARY_PREFIX, leaving the stale one readable as live work - the hijack
class #44454 closed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 19:18:23 -07:00
teknium1
a51143fbbe fix(bedrock): resolve application-inference-profile ARNs before the prompt-cache allowlist match
build_converse_kwargs gated cachePoint markers on the raw model id, so a profile ARN wrapping
Claude matched nothing in _CACHE_POINT_PATTERNS and silently lost Bedrock prompt caching.
_model_supports_prompt_cache now resolves the profile through the per-process-cached
_resolve_inference_profile_model_id first; the request still targets the profile ARN.

Part of #114476
2026-09-18 15:15:02 -07:00
teknium1
11445cc54e fix(bedrock): application inference profile ARNs size from the wrapped model on the real path
The cherry-picked #114482 never resolved a profile in production: the only
caller, agent/model_metadata.py::_resolve_bedrock_context_length, invokes
get_bedrock_context_length(model, probe=False) with no region, and the
resolver was gated on `region`; and it called
get_inference_profile(inferenceProfileId=...) where botocore requires
`inferenceProfileIdentifier`, so even with a region the call raised
ParamValidationError, was swallowed, and the 128k default applied with only
the new warning. Live against a botocore Stubber: 128000 before, 1000000 after.

- resolve in the ARN's own region (field 4), then the passed region, then
  the standard AWS chain; the runtime region / base_url may differ
- inferenceProfileIdentifier is the only request parameter GetInferenceProfile has
- match only `application-inference-profile/`: system-defined
  `inference-profile/us.anthropic...` ARNs embed the model id and need no call
- no nested-profile recursion: GetInferenceProfile lists foundation-model ARNs
  and the wrapped ARN itself satisfies the static-table substring match
- cache both outcomes per process (runs on every context-length resolution),
  cleared by reset_client_cache()
- tests trimmed to two invariants on the production call shape; restore the
  TestBedrockContextProbe class header the cherry-pick clobbered
- docs: bedrock:GetInferenceProfile IAM permission and the fallback WARNING

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 15:15:02 -07:00
liuhao1024
4c28be68c2 fix(bedrock): size inference-profile ARNs from the wrapped foundation model
An application-inference-profile ARN names no model, so the context probe
error text and the static substring table both miss and the 128k default
silently applies: a profile wrapping a 1M-context Sonnet compacts at ~96k.
Resolve the wrapped foundation-model id via GetInferenceProfile (reusing
the control-plane client; an application profile's id IS its ARN, so the
call works on both API shapes), then let the existing probe/table paths
run on that id. Resolution failures (missing bedrock:GetInferenceProfile,
no creds) keep today's behaviour but now log a WARNING naming the
explicit model.context_length escape hatch instead of staying silent.

Fixes #114476
2026-09-18 15:15:02 -07:00
teknium1
96030a3536 fix(compression): scope the checkpoint remediation to capability refusals; trim tests
The remediation ("set compression.checkpoint_required: false, or switch
provider") is right only when the gate can never pass with the active provider
set. Attached to every _checkpoint_blocked() it also decorated the transient
"provider checkpoint API v2 failed" path, where the provider IS capable and the
correct move is a retry once the store recovers, and the codex_app_server
refusal, where switching memory provider changes nothing. A small
_checkpoint_incapable() wrapper carries it on the two capability branches only.

Tests: six near-duplicates collapse into three invariants (warn on mismatch,
parametrized over v1 provider / no manager; silent when off or capable; the
compress-time refusal names the flag). Proven red with the fix neutralized.

Consolidates #106879 (KoNit-K) and #106882 (kokhlo); fixes #106870.
2026-09-18 13:58:52 -07:00
KoNit-K
8e7120e2fd fix(compression): debug-log checkpoint capability probe failures
Keep init fail-open on a broken probe, but emit DEBUG so flaky probes are diagnosable.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
KoNit-K
ef7c7f9e08 fix(compression): warn and remediate opaque checkpoint_required blocks
When compression.checkpoint_required is set but the active memory provider
lacks checkpoint API v2, keep fail-closed compress and surface a startup
warning plus actionable remediation on the block error.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
kshitijk4poor
ae62e832cc fix(anthropic): a possessive 'hermes-agent's' is prose and is still rewritten
Refusing every apostrophe on the right also protected the possessive, leaking
the raw slug in ordinary prose. Only an apostrophe that does not start 's'
joins (quoted identifiers such as name='hermes-agent' stay). The test counts
rewrites in the caller's block, not the Claude Code prefix.
2026-09-19 01:44:10 +05:30
kshitijk4poor
f7676e7ff1 fix(anthropic): quoted slugs stay identifiers; a sentence-final dot is prose
Two boundary cases from review: excluding every `.` from the right boundary let
"…run by hermes-agent." leak the raw slug, so only a dot FOLLOWED by a word char
(host, extension) now counts as a joiner. And the shipped prompt says
`skill_view(name='hermes-agent')` / "the `hermes-agent` skill" — a quoted slug is
a name the model dereferences, so quote characters join both boundaries. The
regression covers a mailbox, the quoted identifier and the sentence-final dot.
2026-09-19 01:44:10 +05:30
kshitijk4poor
75b6281394 fix(anthropic): OAuth slug rewrite skips paths, repo slugs and mailboxes too
The docs host was one instance of a class: every place the ``hermes-agent`` slug
is joined to an address the model dereferences. Subagent briefs naming
``~/.hermes/hermes-agent/venv/bin/python`` came back as ``~/.hermes/claude-code/...``
(a file that does not exist) and ``NousResearch/hermes-agent`` became a repo
nobody owns. Rewrite the slug only when it stands alone in prose: a lookbehind
rejects ``/ . : @ -`` and word characters on the left, a lookahead rejects
``/ . @ -`` and word characters on the right. Product-name prose ("the
hermes-agent skill") is still rewritten as before.

Path class reported by @chazmaniandinkle and @christophcemper on #48860.
2026-09-19 01:44:10 +05:30
Gaurav Saxena
e218f3641c fix(anthropic): preserve docs URL in OAuth sanitizer 2026-09-19 01:44:10 +05:30
kshitijk4poor
e5af917aaa refactor(bedrock): block_index resolves, callers own the current-index state
The helper wrote current_block_index for every event and the stop branch
immediately undid it; make it a pure resolver and assign in the start/delta
branches only. The test's replay tail re-covered ordering the sidecar
assertion already pins — dropped.
2026-09-19 01:35:23 +05:30
kshitijk4poor
5d4a206c92 fix(bedrock): index-less text after a tool stop opens its own block
In the arrival-order fallback (events without contentBlockIndex) a text delta
following contentBlockStop still continued the closed tool's slot and bolted a
stray `text` key onto the toolUse dict — the same mixed block the indexed path
just stopped producing, and one `_replay_ordered_blocks` drops the toolUse from.
Clear the current index at every stop so the next index-less delta starts fresh;
the no-index regression now covers text on both sides of the tool.

Also collapses the dead bool guard (boto3 indices are ints) and trims the
docstring to the one-line WHY.
2026-09-19 01:35:23 +05:30
EricN
8b28df7b1f fix(bedrock): key streamed content blocks by contentBlockIndex, not arrival count
Claude 5 on Bedrock (global.anthropic.claude-opus-5 observed) killed every
tool-continuation turn that followed an assistant reply containing BOTH text
and a tool call:

  ValidationException: This model does not support assistant message prefill.
  The conversation must end with a user message.

The payload did end with a user turn. The rejected shape was the assistant
turn before it, replayed from the bedrock_content_blocks sidecar as

  assistant[text, text, toolUse, text x16, cachePoint]

i.e. the model's text shredded into one block per streamed delta, with the
toolUse spliced into the middle and one fragment missing.

Cause: stream_converse_with_callbacks() only read contentBlockIndex on
contentBlockStart. Bedrock emits NO contentBlockStart for text blocks (only
for toolUse), so current_block_index stayed None and every text delta was
keyed by len(stream_blocks) — a fresh slot per delta. When the toolUse block
then started at its real index it overwrote one text fragment and sorted into
the middle of the text on replay. Claude 5 rejects text trailing a tool_use in
the last assistant turn as prefill; text-only or toolUse-only replies were
unaffected, which is why the failure looked intermittent. The joined
`content` string was correct, so the transcript looked fine while the wire
payload was not.

Fix: read contentBlockIndex off every contentBlockStart/Delta/Stop and key
stream_blocks by it, so all deltas of a block merge into one entry and blocks
keep Bedrock's order. Events without an index (test doubles, proxies) fall
back to arrival order: a start opens a new slot, a delta/stop continues the
current one.

Verified against Bedrock: replaying the dumped failing request reproduces the
400; the same request with the assistant text ahead of the toolUse (merged or
still fragmented) is accepted; dropping only the cachePoint still fails, so
caching is not involved. The added tests use the captured live event sequence
and fail on main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 01:35:23 +05:30
teknium1
7530f40383 fix(errors): auth refusals from a non-stock route name the contacted host
The #113719 report asked that, when the request does fail, the error
names the endpoint that was actually contacted. classify_api_error now
appends "(endpoint: <host>)" to auth-class messages when base_url is set
and is not the provider's own route (provider_owns_route is not True),
so a credential posted to a stale model.base_url reads as a wrong
endpoint rather than a bad key. Stock routes and empty base_url keep
the plain message.
2026-09-18 12:49:40 -07:00
kshitijk4poor
b1968795b2 test(agent): codex TTFB fixture starts from a clean HERMES_CODEX_TTFB_* env
The #64507 test already assumed 'no override' without clearing the shell; fold
the delenv into _make_codex_agent (as the reasoning-effort sibling does) and
drop the per-test helper. The explicit-cap test pins reasoning off too, so the
effort floor can never mask a cap regression.
2026-09-19 01:05:09 +05:30
kshitijk4poor
e3bde9d9c5 docs(agent): TTFB cap is opt-in — say so where operators read it
Docstring tunables line, the cap comment and its log hint still described a
120s default ceiling. The large-request test pins reasoning off so the effort
floor cannot mask a cap regression.
2026-09-19 01:05:09 +05:30
Justin Wilson
7126830d7b test(agent): large codex requests keep the scaled TTFB cutoff; an explicit cap still binds
Regression pair for the scale-then-recap contradiction: with no TTFB overrides a
>100K-token openai-codex request keeps the 180s no-event cutoff, and an explicit
HERMES_CODEX_TTFB_MAX_SECONDS still bounds it. Taken from PR #104339 (the same
fix proposed with a 180s default cap); the fix itself lands as the earlier #91635.
2026-09-19 01:05:09 +05:30
kshitijk4poor
d695359d9e refactor(state): one visibility clause for every reaction path; SQL-side reaction prefilter
set/get used only the lineage filter while take_unseen also required
(active = 1 OR compacted = 1), so a rewound row could be reacted to but never
announced; _DISPLAY_META_ROW_SQL now carries the shared _DISPLAY_ACTIVE_CLAUSE.
take_unseen_reactions scans the whole lineage each turn, so it now filters on
json_extract(display_metadata, '$.reactions') in SQL instead of decoding every
metadata-bearing row in Python. Tests share the compacted-lineage fixture.
2026-09-19 00:59:32 +05:30
kshitijk4poor
fc5424d00e fix(state): reactions resolve rows across the compression lineage, not the tip alone
A display resume materializes the whole compression lineage with row ids
(`get_resume_conversations(include_ancestors=True)`), so the desktop shows —
and lets the user react to — rows that live in an ended parent segment. The
gateway's `session_key` is re-anchored to the continuation after every
compaction, and `set_message_reaction` scoped the row by that exact key, so
every reaction on a pre-compaction message returned None and the desktop
surfaced RPC 4040 "message not found in this session" (#80670: the 802
compacted-row repro; the agent-side `react_to_message` tool hit the same
wall in #108633).

`set_message_reaction` / `get_message_reactions` / `take_unseen_reactions`
now scope by `_resume_lineage_ids(session_id)` — the same set the resume
loads, so an explicit /branch copy still owns only its own rows and an
unrelated session's row stays foreign. The RPC handler and the tool are
unchanged: ownership is decided once, at the row.

Lineage ownership was first identified in #108635 by @KoNit-K (tool path);
the compacted-row half of the unseen-reaction scan is @Liuzikaii's #108542,
cherry-picked ahead of this commit.
2026-09-19 00:59:32 +05:30
Tranquil-Flow
e105d4589d test(agent): pin agent_context derivation and the supermemory cron write-off (#80646)
Two invariants, both red on the hardcoded "primary" base: the scheduler's and
delegate_task's platforms map to their own context while interactive surfaces stay
primary, and the real bundled supermemory provider switches writes off for a cron
session's kwargs. Distilled from the regression file in #107045.
2026-09-19 00:55:05 +05:30
teknium1
cab9a27954 fix(tests): cover the 'no credentials were found' permanent-failure marker
An OAuth auxiliary provider with no login now raises 'no credentials were found'
(#114405 / #78996); the compressor must classify it as permanent instead of
retrying. Extend the existing missing-credential classifier test with that
wording so removing the marker from _SUMMARY_MISSING_CREDENTIAL_MARKERS goes red.
2026-09-18 11:11:44 -07:00
liuhao1024
a0562d17d8 fix(auth): missing-credential hints name the real env var or the OAuth login (#114405, #78996)
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.

One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.

Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.

Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
2026-09-18 11:11:44 -07:00
teknium1
85f5f560c3 fix(agent): a paste-only opener's expansion footer must not demote the title preview
The titler receives the opening message AFTER @-reference expansion, so
the Desktop's generated pasted_content @file: ref arrives with a
'--- Context Warnings ---' (or '--- Attached Context ---') footer. That
footer made the ref-only check in build_title_input fail, so the live
wire capture showed the path + warning leading the title-model input
with 'Pasted content: ...' appended after it. Strip the expander footer
when a preview is present, so a paste-only opener lets the preview lead.

Invariant: test_expanded_paste_ref_footer_does_not_demote_the_preview
(red on a5dac801, green here). Re-checked on the wire against the stub
model: title input now starts with the pasted topic, no @file:/warning.
2026-09-18 10:56:00 -07:00
teknium1
490bd2093b fix: cover the title_preview wire seam and the widened _run_after_agent_ready call
The PR widened tui_gateway/methods_prompt.py::_run_after_agent_ready with a
display_metadata positional; tests/tui_gateway/test_submit_time_user_row.py
(already on main) still called the six-arg form and failed CI with a
TypeError. Update the call.

Add one turn_context-level invariant: a user message carrying
display_metadata.title_preview reaches maybe_auto_title(title_preview=...).
Red with agent/turn_context.py swapped to origin/main, green on this head;
without it the whole prompt.submit -> display_metadata -> titler seam had
no test.
2026-09-18 10:56:00 -07:00
teknium1
34ba61bf67 fix(agent): paste title hint reaches the instant title and the prompt.submit contract
Build on #114129 (@KoNit-K), which carries a Desktop-generated large-paste
preview from the composer through `prompt.submit` -> `display_metadata` ->
turn context -> the shared title input. Two gaps closed:

- `apply_instant_title` never received the preview, so the instant title of a
  paste-only opener was the generated `@file:` path — and stayed that way,
  because the upgrade thread's `derive_title` fallback writes `derived`
  provenance, which never replaces the `derived` title already stored.
  Thread the hint into the instant stage too.
- `build_title_input` let the `@file:` ref lead when the opener was nothing
  but the generated attachment ref; the preview now leads for a ref-only
  opener (an instruction still leads when the user typed one).
- `prompt.submit` gains `title_preview` in the contract (regenerated shared
  TS/OpenRPC); documented as title-only input in the configuration guide.
- Tests trimmed to two invariants (shared input reaches both stages; budget +
  manual attachments stay unread).
2026-09-18 10:56:00 -07:00
KoNit-K
8d25e69b6e fix(desktop): use generated paste previews for titles 2026-09-18 10:56:00 -07:00
teknium1
dac4c60fbb fix(checkpoints): refuse host rollback and session diff from container sessions on the gateway too
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.

Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-18 10:38:08 -07:00
Sora-bluesky
3220b9ed2f fix(checkpoints): classify the backend on every /rollback instead of remembering it
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
2026-09-18 10:38:08 -07:00
Sora-bluesky
10ee9819f9 fix(checkpoints): do not feed container paths into host checkpoint storage
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.

Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).

Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
2026-09-18 10:38:08 -07:00
teknium1
303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00
teknium1
4c5a70065c fix: arm the pool revert only when the benched credential outranks the one rotated to
A session already on the fallback (preferred entry benched by another session) that
rotates UP to the preferred entry once its window reopened must not be pulled back
DOWN when the fallback's own cooldown lifts. Compare priorities in _rotate_and_swap
before arming _credential_pool_revert_id; cover the two-session interleaving as the
control case inside the existing positive test.
2026-09-18 10:35:37 -07:00
teknium1
92bb5b92b8 fix: revert a quota-benched credential through the pool, not a private probe; cover /model and chained rotations
Reshape of the salvaged fix from #114513 (@whyyagswhy):

- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
  benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
  refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
  or round-robin order. The contributor's version called the private
  ``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
  and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
  429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
  the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
  the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
  by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
  hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
  ``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
  fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
  and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
2026-09-18 10:35:37 -07:00
Yagna Vudathu
0dd153ba2e fix(agent): revert pool rotation after transient quota cooldown expires (#114501) 2026-09-18 10:35:37 -07:00