Commit Graph

37924 Commits

Author SHA1 Message Date
teknium1
c8ecc3db64 fix: drop replayed reasoning_details on every chat-completions route that does not read it
OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning
continuity; every other OpenAI-compatible route either ignores the field or, when
its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole
request with 400/422 once an earlier reasoning turn is in history — wedging the
session after an in-session model switch (#70233). Strip the field from the wire
copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url),
mirror it in the auxiliary wire boundary and the iteration-summary path; state.db
history keeps the field so switching back to OpenRouter/Nous replays it again.
2026-09-19 09:28:45 -07:00
teknium1
0837f549be fix: drop base64 imports left unused by the shared codex_account_headers helper
ruff --select F401 flagged both after the per-module _extract_chatgpt_account_id
copies were removed.
2026-09-19 09:28:13 -07:00
teknium1
59e858c20c test: pin the residency header on both models-catalog probes' wire requests
test_residency_header_from_jwt_claims asserted the shared helper directly, so
reverting the hermes_cli/codex_models.py or agent/model_metadata.py call sites
to base stayed green. Drive _fetch_models_from_api (httpx.get) and
_fetch_codex_oauth_context_lengths_with_source (requests.get) with a
compute-residency-only JWT and assert x-openai-internal-codex-residency and
ChatGPT-Account-ID reach the outgoing headers. Red with either call site
reverted to base; green at head.
2026-09-19 09:28:13 -07:00
teknium1
566513911b test: trim the residency header tests to two invariants
Five change-detector tests collapse into one positive invariant (data
residency wins, compute residency is the fallback, the sibling /usage
header builder derives the same value) and one control (no claim or a
malformed token never carries the header). The residency value must
otherwise be probed live; see the PR body.
2026-09-19 09:28:13 -07:00
teknium1
7471586ad0 fix: send the Codex residency header on every JWT-derived Codex request
The Codex models catalog probe (agent/model_metadata.py), the picker catalog
fetch (hermes_cli/codex_models.py), the quota-restored probe
(hermes_cli/auth_codex.py) and the /usage dashboard call
(agent/account_usage.py) each re-decoded the OAuth JWT for
ChatGPT-Account-ID and would 401 on residency-enforced workspaces exactly
like the chat client did (#23896). They now share
agent.codex_headers.codex_account_headers, which emits the account and
x-openai-internal-codex-residency headers from one decode; the three
duplicate `_extract_chatgpt_account_id` decoders are gone.

account_usage keeps auth.json's account_id as the winner over the JWT claim
(pool-only credentials still omit it); header casing unifies on the
codex-rs canonical `ChatGPT-Account-ID` (HTTP header names are
case-insensitive on the wire; the two tests asserting the old casing follow).
2026-09-19 09:28:13 -07:00
liuhao1024
0bbf7b7997 fix(agent): extract residency claims from Codex OAuth JWT for workspace auth
Extract chatgpt_data_residency (fallback: chatgpt_compute_residency) from
the OAuth JWT and send as x-openai-internal-codex-residency header on
requests to chatgpt.com/backend-api/codex. Residency-enforced workspaces
return 401 without this header.

Fixes #23896
2026-09-19 09:28:13 -07:00
teknium1
e132e69374 fix(codex): match the primary error on '401 unauthorized', not any bare 401 token
The source-aware classifier replaced base's '401 unauthorized' needle with a
\b401\b regex on the primary error, so any unrelated primary message
carrying a bare 401 token (e.g. 'request body exceeded limit by 401 bytes')
produced the re-login hint and retired the session. Restore the base needle
alongside 'unauthorized' (which already covers the issue's case) and drop
the regex and its 're' import; pin the negative case.
2026-09-19 09:27:40 -07:00
teknium1
3e7a8636e8 test(codex): pin source-aware OAuth classification for plugin 401 stderr
Two invariants for #75167, replacing the seven change-detector tests from
#75182: the classifier table (issue's four cases, parametrized) and the
session path (turn/start and thread/compact/start RPC errors keep the
primary error and stderr tail visible, no retire, no `codex login` hint).
The existing empty-input test moves to the keyword-only `stderr=` form.
2026-09-19 09:27:40 -07:00
cosin2077
48d59f151f fix(codex): plugin 401 noise in app-server stderr no longer masks the real turn error
`_classify_oauth_failure` joined the primary JSON-RPC error and the codex
stderr tail into one haystack and matched broad tokens ("unauthorized",
"401 unauthorized", "oauth"). codex writes independent ChatGPT plugin
prewarm failures ("HTTP 401 Unauthorized") to stderr while the core
JSON-RPC server keeps working, so any unrelated RPC error, timeout or
subprocess exit was rewritten into the `codex login` hint and the real
error plus stderr tail disappeared.

Classify by source: generic 401/unauthorized/oauth text is authoritative
only in the operation's own error; ambient stderr needs a strong
credential signal (invalid_grant, refresh/expired token, no auth profile).
Every call site (turn error, request timeout, dead subprocess, and the
compaction paths that share them) passes stderr by keyword.

Salvaged from #75182 by @cosin2077, hand-applied onto the refactored
session module (call sites collapsed into `_set_classified_error` /
`_request_for` / `_subprocess_died`).

Fixes #75167
2026-09-19 09:27:40 -07:00
teknium1
8f1542422b fix: fallback entries inherit a named provider's declared transport
A fallback_providers entry naming a `providers.<name>` block (or
`custom:<name>`) without its own api_mode was re-detected from the
resolved host: an Anthropic-Messages proxy on a plain host or a
Responses-only relay behind a generic gateway landed on chat_completions
while resolve_provider_client had already built the declared client
(#33062 bottom thread, #81932 provider-level `transport:`). The hint pass
now reads the named block's api_mode/transport as explicit, and entry-level
`transport:` is accepted as an alias of `api_mode` with the same
canonicalisation the providers block uses (`responses` -> codex_responses).

Fixes #33062
Fixes #81932
2026-09-19 09:27:08 -07:00
teknium1
45356370b0 fix(tests): pin Responses tool strictness at both production entry points
The direct-call helper test stayed green when convert_tools / the auxiliary
adapter re-flattened strict to False. Drive ResponsesApiTransport.build_kwargs
and _CodexCompletionsAdapter._build_responses_kwargs instead and assert an
explicit strict: True reaches kwargs['tools'] on both routes (#105401 parity).
2026-09-19 09:26:35 -07:00
Andrew Johnson
908fb96fb6 fix(agent): preserve explicit Responses tool strictness 2026-09-19 09:26:35 -07:00
teknium1
f75263f811 fix: render dict-form reasoning_effort in the TUI config.get and setup wizard readers
`agent.reasoning_effort: {enabled, effort}` now reads back as its tier name
in the TUI `config.get reasoning` result and in the setup wizard's
"currently in use" lookup instead of `str(dict)` (a disabling dict reads
as `none`). Both route the dict through `parse_reasoning_effort` so the
dict semantics live in one place. Docs note the dict form is config.yaml-only.
2026-09-19 09:25:29 -07:00
teknium1
5c113bb07c feat: accept dict-form reasoning_effort for providers with bespoke thinking tiers
`agent.reasoning_effort` and `agent.reasoning_overrides` values now accept
`{enabled: true, effort: <level>}`; the level is passed through verbatim to
the wire (CustomProfile / clamp_effort already forward unknown names), so a
relay exposing `fast`/`thinking` can be asked for its real tier instead of
silently running at the default `medium`.

Bare strings stay strict: a non-ladder string is still rejected with the
existing warning, so a typo like `hgih` never reaches a request. Only the
config parser (`hermes_constants.parse_reasoning_effort`) rejected custom
names — the transport layer was already designed to pass them through.

Slim redo of the change proposed in PR #93239 (the base function had since
been compacted, so the hunk no longer applied); docs and example config
updated in the same change.

Co-authored-by: HermesDev-Bot <309177324+HermesDev-Bot@users.noreply.github.com>
2026-09-19 09:25:29 -07:00
teknium1
4d1d3d05a3 fix(codex): drop the message id of Azure-trimmed reasoning turns too
_newest_reasoning_only pops older turns' codex_reasoning_items before the
converter runs, so those turns looked reasoning-free and replayed their
msg_* id with neither the reasoning item nor its rs_* id on the wire — the
exact orphan shape #97427 rejects. The trimmed row now carries a transient
codex_reasoning_trimmed marker and _replay_message_items honours it; the
single-use _turn_has_encrypted_reasoning wrapper is inlined into that
predicate. Test drives ResponsesApiTransport.build_kwargs on an Azure host.
2026-09-19 09:24:57 -07:00
teknium1
3e80025708 chore: map contributor email for salehelsayed (PR #97445 co-author) 2026-09-19 09:24:57 -07:00
teknium1
801590af2d fix(codex): drop a replayed message id when its reasoning id was stripped
Stateless Responses replay (store=False) strips every reasoning item's rs_*
id, but the assistant message minted in the same response kept its msg_* id.
GPT-5.6-family endpoints validate that link and reject the continuation with
HTTP 400 "Item 'msg_…' of type 'message' was provided without its required
'reasoning' item: 'rs_…'" on every post-tool turn, deterministically.

The converter now drops the message id whenever the stored turn carried
encrypted reasoning (replayed, suppressed by recovery, or dropped as a
foreign-issuer blob), keeping content/status/phase. Reasoning-free turns keep
their id for prefix-cache affinity. Done in _replay_message_items so the main
transport, preflight and the auxiliary Codex adapter all emit the same shape.

Co-authored-by: salehelsayed <saleh.fekry@gmail.com>
2026-09-19 09:24:57 -07:00
teknium1
29bc6343d3 fix(auth): read-only Codex reads take no store lock; /model picker never refreshes
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.

`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.

`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
2026-09-19 09:24:25 -07:00
teknium1
d43c9622f3 fix(auth): status and doctor Codex reads never adopt, refresh or persist credentials
`hermes status` / `hermes doctor` / the dashboard cards and the `/model` picker
call `get_codex_auth_status()`, whose singleton fallback ran
`resolve_codex_runtime_credentials()` with runtime defaults: a store missing
its refresh_token imported the Codex CLI's single-use token family, and an
expiring token was refreshed and written back. A diagnostic that spends or
borrows a rotating refresh token logs the other program out (#68004).

`resolve_codex_runtime_credentials(read_only=True)` reports the stored state
as-is (no CLI adoption, no refresh, no pool forced refresh, no write) and wins
over `force_refresh`; the Codex status snapshot uses it, and the xAI OAuth
snapshot passes `refresh_if_expiring=False` for the same reason. The pool side
was already an observation (`peek`, 964fbaae2c).

Superseded #68224 (@GauravPatil2515) — same mechanism, re-done on the current
`auth_codex` layout.

Co-authored-by: Gaurav Patil <gauravpatil2516@gmail.com>
2026-09-19 09:24:25 -07:00
teknium1
e62c5e3360 fix(classifier): pin the body-less 'Unsupported content type' envelope in the codex replay test
The wrapped-message-only branch of _is_codex_masked_replay_rejection (SDK paths
that surface no parsed body) had no test; it joins the existing parametrize as
codex-message-only. Neutering the branch turns that id red.
2026-09-19 09:23:54 -07:00
teknium1
3665a535bb fix(codex): send string role-message text as typed parts on the ChatGPT Codex backend
The second failure mode in #51512: with no reasoning replay at all, a single
``{"role": "user", "content": "<text>"}`` item still 400s on the ChatGPT Codex
backend with ``{"detail": "Unsupported content type"}``. The classifier mapping
from the previous commit cannot recover that turn (turn_recovery's strip needs
cached codex_reasoning_items), so the wire shape has to be right up front.

``_chat_messages_to_responses_input`` now wraps string user/assistant text as
``input_text`` / ``output_text`` parts when the issuer is ``codex_backend``; every
other Responses route keeps the string shorthand it has always received. The
preflight already validates typed parts, so the real call path (build_kwargs ->
preflight_kwargs) needs no separate change.

Test drives ResponsesApiTransport.build_kwargs + preflight_kwargs, red on the
old head for the codex case. Three test_native_compaction asserts pinned the
assistant string incidentally (they check history survives, not its shape) and
now expect the typed part on the codex route.
2026-09-19 09:23:54 -07:00
teknium1
c7cf3f8af1 fix(classifier): route two more rejected-reasoning-replay 400s to the replay strip
Two Responses-API 400s meant "the replayed encrypted reasoning was rejected"
but never reached the one-shot recovery in turn_recovery (disable replay,
strip codex_reasoning_items, retry):

- OpenAI's ``thinking_signature_invalid`` code contains "thinking" and
  "signature", so the Anthropic thinking-block heuristic claimed it first;
  that recovery strips Anthropic fields and resends the same stale encrypted
  item on every retry (#70595).
- The ChatGPT Codex backend returns a bare ``{"detail": "Unsupported content
  type"}`` for the same rejection, which matched nothing and aborted the turn
  as a non-retryable format_error (#51512). It joins the #92353 exact-envelope,
  provider-gated codex mapping, so a no-replay 400 keeps aborting as before
  (the recovery still requires cached codex_reasoning_items).

Co-authored-by: ooiuuii <al3060388206@gmail.com>
2026-09-19 09:23:54 -07:00
luyifan
43127a86ea fix(cli): bound Azure detect response reads 2026-09-19 09:23:22 -07:00
luyifan
0acd96a439 fix(agent): classify Codex account token failures 2026-09-19 09:22:50 -07:00
teknium1
a79d1d3a71 feat: one-shot runs drop the self-improvement footprint (no skill authoring, fewer process skills, delegation cap)
A finite `hermes chat -q` / `--oneshot` run has no later session in its HERMES_HOME
to learn for, yet it ran the full interactive self-improvement loop. Measured over
21 one-shot benchmark trajectories: 7 skills created and a bundled one patched
mid-task, 37 of ~215 tool calls on skill_view/skill_manage, skill text = 34% of all
tool-result bytes fed back into context, plus reviewer subagents spawned on the
agent's own diff (one task: 5 delegations, 62 subagent API calls, each re-paying a
cold system prompt).

Keyed on the existing HERMES_SINGLE_QUERY_SESSION marker (approval gate, delegation
dispatcher), so interactive and gateway sessions are byte-identical:

* agent/oneshot_footprint.py (new sibling): skill_manage is pruned from the tool
  set; the ## Skills block keeps the index + skill_view but drops the record/patch/
  offer-to-save coaching and the "load process skills for work you already know"
  push (SKILLS_GUIDANCE follows because it is gated on skill_manage).
* delegation.oneshot_max_children (default 2, 0 = unlimited): total children a
  one-shot run may spawn; past it delegate_task returns a tool error telling the
  model to finish inline.
* requesting-code-review skill: reviewer/fixer subagents (Steps 5 and 7) are
  interactive-only; one-shot applies the checklist inline.

Live: `hermes chat -q "list tools starting with skill_"` on the portal — base
"skill_manage, skill_view, skills_list", fix "skill_view, skills_list"; a real
AIAgent under a temp HERMES_HOME shows the interactive prompt unchanged (6643 chars
both) and the one-shot prompt without skill_manage / offer-to-save.
2026-09-19 09:21:18 -07:00
teknium1
fa0aade11b test(serve): idle-proof live probe waits for the verdict to settle between cron ticks
The three-child live test sampled /api/health/idle once, the instant each
child printed READY. With HERMES_DESKTOP=1 the child runs the in-process cron
ticker, and every tick holds retirement admission for its whole scan (#98745,
on purpose), so a sample inside a tick says `retirement_admission` for an idle
resident, or names the admission instead of `cron:<job>` for the busy child.
~1 in 10 runs locally, twice in a row on PR CI at 96 workers.

The contract is "stable between ticks", so poll until the verdict names the
stable state (idle for residents, the cron ledger for the busy child); the last
verdict is still asserted, so a genuinely wrong state fails. 20/20 green after.
2026-09-19 09:20:54 -07:00
teknium1
f64ac2fbdc fix(curator): never-used built-ins get their inactivity clock anchored at first curator sight
Telemetry writes a bundled skill's usage record the moment it is seeded, so by
the curator's first pass with prune_builtins on, created_at can be months old and
every never-used built-in is marked stale at once (71 on the reporting install,
#79295). The first-sight seed only covered records that did not exist yet.

apply_automatic_transitions now re-anchors a bundled, never-used record without
first_seen_at to now (skill_usage.reanchor_clock) and defers it one pass; the
stamp makes it one-shot so the skill still ages out after a full window. A record
the bug already marked stale is reactivated in the same step (the reviewer's
migration edge on #79311), instead of waiting a week for the next pass.

Design and first implementation by @webtecnica in #79311; re-applied by hand
because that branch predates the skill_usage decomposition.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-19 09:20:54 -07:00
teknium1
4fc3e9be85 fix(curator): tarball skips regeneratable dirs; ledger --compact garbage-collects blobs
Second half of #107539 on top of @liuhao1024's snapshot_paths filter: the same
TRANSIENT_DIRS set (venv, node_modules, caches, .git) now gates the whole-tree
snapshot (directory parts only, so a file named `venv` is still skill content),
and rollback carries a live nested venv back the way it already carries `.git`.
Nothing ever pruned the blob store; `hermes curator ledger --compact` now
deletes every blob no ledger entry references (98.9% of 47k blobs on the
reporting install). A malformed ledger line aborts the sweep — an unreadable
entry may still hold references.

Supersedes the curator_backup half of #100545 / #81669.
2026-09-19 09:20:54 -07:00
liuhao1024
0ed64d4a37 fix(skill_ledger): stop snapshot_paths from sweeping transient dirs into the blob store
snapshot_paths() hashed every file under a skill dir with no exclusion
filter, so a stray venv/node_modules/__pycache__/.git under a skill was
copied content-addressed into ~/.hermes/.curator_backups/blobs/ on every
mutation — and nothing ever prunes that store, so the blobs grew without
bound (real install: 47k blobs / 1.3 GB in one day, 98.9% unreferenced).

Filter at capture: skip any file whose parent-chain component matches
_SNAPSHOT_EXCLUDE_DIRS (the same set proposed for the curator tarball in
transient dir is still captured — only files inside those dirs are
dropped.

The unreferenced-blob GC suggested in the issue is intentionally left
out; capture-side filtering stops the growth, and pruning existing
garbage is a separate, riskier change.
2026-09-19 09:20:54 -07:00
kshitijk4poor
d7ea741ae0 test(discord): cover the YAML bridge and the global auto_thread gate
- config.yaml -> `_apply_yaml_config` seeding for discord.free_response_auto_thread
  (the operator path, previously only proven by an ad-hoc probe)
- `auto_thread: false` still disables threading with the opt-in on
2026-09-19 21:43:59 +05:30
kshitijk4poor
1536e1d242 docs(discord): finish the free-response auto-thread surfaces
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
2026-09-19 21:43:59 +05:30
kshitijk4poor
8a2799276f test(discord): cover the free-response auto-thread opt-in's exclusions
- clear DISCORD_FREE_RESPONSE_AUTO_THREAD in the adapter fixture, so a contributor
  shell exporting it can no longer decide the default-path test
- no_thread_channels precedence now shows the flip inside one test rather than
  passing on the default path alone
- voice-linked channels stay inline with the opt-in on: once skip_thread is cleared
  that exclusion is the only thing holding them back
2026-09-19 21:43:59 +05:30
kshitijk4poor
0b42f031ab docs(discord): describe the free-response auto-thread opt-in where the inline default is stated
The quick-start table and the auto_thread section still promised inline replies
unconditionally, and configuration.md's discord block omitted the key. Adds a
discord.free_response_auto_thread section with the precedence rules (no_thread_channels
wins, voice-linked channels unaffected).
2026-09-19 21:43:59 +05:30
kshitijk4poor
0c7738dfd6 refactor(discord): read the free-response auto-thread flag through a getter
Mirrors the _discord_require_mention()/_discord_max_attachment_bytes() shape, makes the
lookup lazy (it only runs on the free-channel path now), and keeps the in-code key
manifests in _handle_message and register() listing the new key.
2026-09-19 21:43:59 +05:30
Xunjin ZHENG
774e070731 feat(discord): add free_response_auto_thread opt-in
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.

Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.

Default behavior is unchanged; all 291 existing discord tests pass.
2026-09-19 21:43:59 +05:30
kshitijk4poor
2c7c44ea0d chore: map contributor email for FunJim (#18455 salvage) 2026-09-19 21:43:59 +05:30
kshitij
469296c748 refactor(compression): tighten the oversized-turn exception after review
Review findings folded into the exception added by the previous commit:

- The anchored-region check now reuses `_walk_tail_budget` instead of a
  second token sum with the default thought-charge rule. The walk charges
  thinking only on the newest assistant turn unless the route replays stale
  thinking (#73624/#84371), so the second rule could fire the exception on a
  region the walk itself considers inside the ceiling.
- Two conjuncts (`last_user_idx >= head_end`, `last_user_idx < cut_idx`) are
  implied by `user_anchored_cut < cut_idx`, which only changes when the
  anchor found a real user turn inside the compressible region; the check
  that depended on them is documented where it is read.
- The latest-assistant anchor is no longer computed and discarded under the
  exception (it also logged an anchor it never applied).
- `_find_last_user_message_idx` early-exits instead of materialising every
  actionable user index on the per-attempt boundary path.
- The exception logs at debug, like the sibling anchor decisions, rather
  than at info from both `_compress_window` and `has_content_to_compress`.

Adds the missing guard test for the new ceiling check: a transcript that
fits the tail budget must still anchor the active request rather than
triggering the exception. Red when the ceiling conjunct is removed.
2026-09-19 21:12:46 +05:30
kshitij
d32f162cad fix(compression): keep the tail anchors the split exception was bypassing
The oversized-turn exception (#80449) relaxed three tail anchors at once,
which voided two guarantees that hold on main:

- `compression.min_tail_user_messages` was skipped whenever the exception
  fired, so an N-user tail could come back with no user turn at all. The
  N-user anchor now always runs: the setting is a user-facing promise and
  outranks the budget.
- The exception fired even when the oversized weight was the active turn's
  own newest tool group. The pre-anchor cut retains that group anyway, so
  the split bought no reclaim while taking the active request out of the
  tail and losing the #10896 anchor. The exception now requires real turn
  body (a tool-call group) between the opening request and the cut.

Both failures are bound by tests, each red when its conjunct is removed:
`TestTailTokenBudgetCeiling::test_message_floor_does_not_unboundedly_override_soft_ceiling`
and `TestMinTailUserMessages::test_n_guarantee_wins_over_tail_token_budget_and_floor`
(each passes on main), plus `test_n_user_tail_guarantee_outranks_the_split`
for the N-user guarantee on a genuinely oversized turn.
2026-09-19 21:12:46 +05:30
embwl0x
cc4bb888a4 fix(compression): split oversized active turns 2026-09-19 21:12:46 +05:30
kshitijk4poor
39abfdf4ec fix(context-compressor): the marked-leaf guard must match the whole tail
Second review round on the previous commit, both findings reproduced:

- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
  about: a replayed leaf of head + marker followed by new content was never
  shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
  the marker to close the leaf, so that shape shrinks to head + marker with
  true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
  adds separator whitespace, so compact args with many keys could come back
  LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
  The helper now returns the caller's string unless the whole rewrite is a net
  reduction.

Tests cover both shapes plus a many-key compact payload.
2026-09-19 21:06:12 +05:30
kshitijk4poor
a5381a7dca fix(context-compressor): match the marker by position, and leave args byte-identical
Review findings on the previous commit, all reproduced:

- The "already marked" guard was a substring test, so a leaf that merely
  contains the marker — including one a model imitated into a new call, the
  #83714 failure mode itself — was exempt from shrinking forever. The marker
  is always written at `head_chars`, so the guard now tests that position: a
  1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
  JSON came back with inserted spaces and callers read it as a change
  (rewriting replayed history and counting a pressure hit for zero reclaim).
  Nothing replaced now returns the original string: a 547-char compact blob
  is byte-identical.
2026-09-19 21:06:12 +05:30
kshitijk4poor
1e216d1330 fix(context-compressor): only shrink args when it reclaims, never re-shrink
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:

- A leaf just over `head_chars` came back LONGER: a 628-char args blob
  shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
  compaction, so a leaf's marker was rewritten from the true count
  ("2,800 of 3,000 chars omitted") to a self-referential one
  ("223 of 423"), destroying the per-instance count property the marker
  relies on and churning those bytes on each pass.

A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
2026-09-19 21:06:12 +05:30
Daniel JB Clark
262a6436fa fix(context-compressor): stop injecting an imitable truncation marker into replayed tool_calls
Root cause for #83714 (write_file/patch_tool writing literal
"...[truncated]" into files, PR #83752's guard is the safety net, not
the fix): _truncate_tool_call_args_json() in the compression pass
shrinks long string values inside a PAST assistant message's
tool_calls[].function.arguments — the exact field that represents the
model's own prior generated output, replayed back to it verbatim on
every subsequent turn. The old marker, a bare "...[truncated]" suffix,
is indistinguishable from something the model itself could have
written (it's exactly the kind of terse ellipsis abbreviation models
already produce). A model conditioned on seeing itself "get away with"
that pattern in its own history imitates it in a new tool call,
writing the literal marker instead of real content.

This is the second bug from the same root text. The first (#11762,
MiniMax 400s from unterminated JSON) was fixed by shrinking inside the
parsed structure so the JSON stays valid, but kept the same visible
marker text — fixing the syntax problem while leaving the imitation
problem untouched.

Fix: replace the marker with one deliberately NOT shaped like prose a
model would write — distinctive non-ASCII delimiters, an explicit "not
part of the original tool call" disclaimer, and a per-instance
char-count that won't match the next omission point even if copied
verbatim. The shrunk value stays a plain string (not a nested object)
so the #11762 valid-JSON/matching-shape contract is unchanged — only
the marker text changed.

Checked context_compressor.py's other "...[truncated]" call sites
(_serialize_for_summary, _compact_fallback_turn, the user-message-only
one near _ACTIVE_TASK_MAX_CHARS) — none of them write into a value
that gets replayed as the main model's own assistant/tool_calls
history, so they don't share this priming risk and were left as-is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:06:12 +05:30
kshitijk4poor
6339f7a2bb chore: map contributor email for djbclark (#83843) 2026-09-19 21:06:12 +05:30
kshitijk4poor
d7b836ab1c refactor(tui_gateway): stack @method/@_profile_scoped on vault handlers like every sibling
The per-file method() wrapper only composed the two registry decorators;
methods_free_tier/_complete/_config_set/_session_control stack them directly.
Test: drop the manual multiplex save/restore (tests/conftest.py resets the
latch per test) and the inert OP_SERVICE_ACCOUNT_TOKEN setenv.
2026-09-19 19:32:53 +05:30
kshitijk4poor
ae7126353f fix(tui_gateway): vault.* RPCs bind the launch profile's secret scope once the process multiplexes
methods_vault.py carried its own scoping decorator that returned early for the
launch profile (`_profile_home()` → None). That is correct only while the
process serves one profile: the first secondary home served flips
`set_multiplex_active(True)` and unscoped `get_secret()` reads fail closed, so
every vault.sources / vault.list call for the launch profile (the Desktop sends
no `profile` for it) raised UnscopedSecretError inside
OnePasswordLoginBackend.__init__ and the Passwords & Logins panel showed
"Could not load vault items" until `hermes gateway restart`.

Route the handlers through server.py's `_profile_scoped` via
`HandlerRegistry.profile_scoped`, the path every other methods_*.py uses: it
binds `launch_secret_scope()` for the launch profile under multiplex and also
honours a session_id-only call. Unknown `profile` names now surface as the
dispatcher's 4064 instead of a vault 5095.

Regression test: vault.sources for the launch profile with multiplex active
(red on main with the production traceback). E2E A→B→A over
`handle_request` with two on-disk homes resolves ops_LAUNCH_A / ops_WORK_B /
ops_LAUNCH_A.
2026-09-19 19:32:53 +05:30
Siddharth Balyan
17b5df02f2 feat(mcp): connect to the official n8n server from the catalog (#116063)
* feat(mcp): add the official n8n server to the catalog

Connect to the user's instance over HTTP with browser OAuth. Keep a
separate n8n-official identifier so the retired n8n bridge is neither
relabeled nor overwritten and retains its credentials and tool filter.

Save ordinary catalog setup values in server config, retaining secret
references in .env. Pass field secrecy through the catalog API so URLs
and client IDs remain visible while credentials stay masked.

Keep the existing install-then-authorize lifecycle. Transactional setup
and cancellation changes are outside this catalog addition.

* refactor(mcp): limit n8n PR to catalog addition

Remove shared installer, storage, field-masking, and input-handler changes.
Those behaviors are being handled in a separate PR. Restore their tests
and Asana guidance to the base branch.

Keep only the official n8n manifest and setup documentation, using the
catalog's existing setup and persistence behavior.
2026-09-19 17:51:31 +05:30
Siddharth Balyan
8d4abc3ea6 fix(mcp): retire the n8n bridge catalog entry (#116048)
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.

Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
2026-09-19 17:51:30 +05:30
kshitijk4poor
7c6f21a5e1 docs(compression): examples and the delegation trigger note follow the default threshold_tokens cap
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
2026-09-19 16:28:57 +05:30
kshitijk4poor
af0be164c1 refactor(compression): one trigger derivation for update_model and the switch-guard preview
`preview_threshold_tokens` restated the resolve -> floor -> compute -> cap chain that `update_model`
runs; two copies of the trigger math drift the next time a step is added — the exact bug class
#83450 fixes (the guard quoting a number the compressor will not install). `_derive_trigger` is the
single pure derivation; the auxiliary-summariser ceiling stays in `_apply_threshold_tokens_cap`
because it is per-runtime, not per-model.

The startup banner names the cap only when it set the trigger; on windows where the ratio already
sits below it "(capped at 256,000)" was noise. Comments no longer repeat the default literal.
2026-09-19 16:28:57 +05:30