300 Commits

Author SHA1 Message Date
ppazosp
e81be5b66a fix(compression): preserve the active request after SQLite reload
(cherry picked from commit 6ee1d087caf3f158d7379c8e4ca09a67c9c0e684)
2026-09-27 00:48:22 +05:30
kshitijk4poor
e30e61f6ff refactor(compression): return the seeded prompt without re-normalizing it
The retain early return coerced agent._cached_system_prompt with `or ""` and
wrote it back. The only setter of _retain_seeded_system_prompt,
gateway.run._seed_hygiene_system_prompt, always stores a str ("" or the
stored prompt) in the same call, and nothing between the seed and the commit
boundary resets it on the detached hygiene / manual-compress agent (the only
None writers are invalidate_system_prompt, which this branch returns before,
and switch_model, which those agents never run). So the normalize/write-back
was dead defence; the empty-seed case still returns "".
2026-09-27 00:42:12 +05:30
kshitijk4poor
f82d1ee449 docs(compression): say why the seeded-prompt early return skips the tool refresh
The retain branch in _rebuild_system_prompt_at_boundary returns before
_refresh_agent_tool_definitions. That is intentional: refresh_agent_mcp_tools
-> persist_agent_tool_names would otherwise pin the live session row's
tools[] to the detached hygiene agent's memory-only toolset. Requested
inline by teknium on #122825 so a later refactor does not move the refresh
above the early return.
2026-09-27 00:42:12 +05:30
Regina
556b8427b7 fix(compression): hygiene and gateway /compress keep the seeded system prompt
Gateway hygiene and the gateway /compress handler compact with a detached
AIAgent(enabled_toolsets=["memory"]) seeded with the session's stored prompt
(_seed_hygiene_system_prompt, 76a17046e2 / 678916b427). Since #98426 the
commit boundary always rebuilds the prompt, so the seed was discarded and the
reduced-toolset build was persisted over the live session's snapshot: the
skills index (## Skills (mandatory) / <available_skills>) and Skill Safety
Rule vanish, and every later fresh agent for that session restores the
degraded bytes verbatim (_stored_prompt_matches_runtime does not compare
platform).

The seed now marks the agent (_retain_seeded_system_prompt) and the
commit-boundary rebuild keeps the seeded bytes for it. A live agent's own
compaction still rebuilds, so builder updates keep reaching long sessions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit dd5c8ef4e845b2d34e826e4d92afad73a12eefa6)
2026-09-27 00:42:12 +05:30
kshitijk4poor
62ceb7f4d7 docs(compression): note why adoption runs before the primary-runtime restore
At turn start the child prompt is checked against the (possibly fallback)
runtime before _restore_primary_runtime runs. A false reject only leaves the
slot None, and the normal restore re-reads the row and re-checks it, so the
order is intentional; say so, so nobody reorders it later.
2026-09-27 00:35:39 +05:30
kshitijk4poor
41115788f6 refactor(compression): collapse the adoption cache-slot seed into one assignment
The clear-then-conditionally-set sequence and its two overlapping comment
blocks said the same thing twice. One conditional assignment keeps the
behaviour identical (None when the child prompt is empty or fails the
runtime-identity check, the child prompt otherwise) and is easier to read.
2026-09-27 00:35:39 +05:30
Uttkarsh Tiwari
21b86ab13e fix(compression): drop the parent prompt cache when adopting a live child
_adopt_live_compression_child moves agent.session_id onto the live compression child,
but when the child's stored prompt fails _stored_prompt_matches_runtime the parent's
non-null _cached_system_prompt stayed in place. turn_context restores/rebuilds only
while that slot is None, so the child turn could send the parent session's prompt with
no validation. Clear the slot right after child confirmation, before the conditional
seed; a rejected tip then rebuilds through the canonical restore path and a matching
tip is still adopted verbatim.

The regression test now seeds a NON-NULL parent cache — the previous null-path-only
seed hid the leak — and asserts the rejected-tip agent carries session_id="child"
with the slot cleared (red on the previous head).

Review fix for #121843 (ehz0ah, agent/conversation_compression.py:1703).

(cherry picked from commit 531a5072ff4a2e74bd7003094786e32136d2545b)
2026-09-27 00:35:39 +05:30
teknium1
f807217ff1 fix(compression): apply the runtime-identity check when adopting a compression tip's prompt
_adopt_live_compression_child seeds agent._cached_system_prompt straight from the tip's row,
and turn_context only runs _restore_or_build_system_prompt while that slot is None. Until now
a route commit NULLed the row, so a tip whose model moved was never seeded; with the writers
preserving the snapshot, the tip carries the old Model:/Provider: trailer and would be served
to the new model with no check at all. Gate the seed on _stored_prompt_matches_runtime, the
same predicate the restore path applies; an unseeded slot rebuilds on the next restore.

Review fix for #121843 (found in review of the widened invariant; probe: parent ended by
compression, child /model-committed to model-b, adopting agent on model-b was seeded
"Model: model-a" on the PR head and nothing on main).

(cherry picked from commit d7143de9ce7f43201b17e8df5c0344aee777b7fc)
2026-09-27 00:35:39 +05:30
kshitijk4poor
045d93a908 refactor(compression): drop dead skills opt-out from _reset_read_dedup_caches
The Codex app-server compaction path was the only caller passing
skills=False; now that it re-arms skill_view dedup like every other
rewrite boundary, the kwarg and its early return are dead. Removing them
keeps the helper's contract simple: every boundary resets both caches,
so a future caller cannot reintroduce the #101518 asymmetry by accident.
2026-09-27 00:19:50 +05:30
Thomas Hudspith-Tatham
55a93d8f70 fix(compression): codex app-server compaction re-arms skill_view dedup
skill_view returns a stub on a repeat view of an unchanged file, pointing
the model at the earlier full result in its own context. That contract only
holds while the earlier result is still there, so a compaction boundary has
to advance the dedup generation -- which is what _reset_read_dedup_caches()
is for, and why its stub message promises "Re-issued after context
compression, this returns the full content again."

The Codex app-server path passed skills=False and reset only the file-read
half. A skill re-viewed after that compaction got a stub naming content the
compaction had just summarised away, so the model had nothing to copy from
and reconstructed it from memory instead. Observed with a delegating skill
whose child contract is a JSON schema: successive dispatches shipped
progressively degraded schemas -- a dropped allOf, then a bare
{"type": "array"} -- until the generation-time check no longer constrained
anything and the work was abandoned.

The opt-out carried the pre-refactor asymmetry forward mechanically
(0e9d46511a lifted these call sites without changing behaviour); nothing
depends on it. Drop it so both paths reset both caches. Tests cover the
helper and drive the codex path end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-authored-by: aurorabotticus-svg <aurorabotticus-svg@users.noreply.github.com>
(cherry picked from commit 22f36fe10ec27bf68e003058b3f1dda292e535e8)
2026-09-27 00:19:50 +05:30
Yuan Li
a5e8ad9f02 fix(agent): bound sustained summary-overload aborts so they cannot guarantee a session wipe
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.

After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.

Fixes #123167

(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
2026-09-26 22:42:38 +05:30
brooklyn!
1674499d00 fix(compression): archive the rows the compressor held, not every row up to the newest
A gap below the newest held id, and turns appended above an unpersisted current
turn, were summarized away without being read. Name those held ids and clone
the rest.
2026-09-24 11:57:56 -05:00
kshitijk4poor
c68091c043 fix(compression): a micro-compaction pass that loses the race is a true no-op
A compaction landing during the micro summary call made _sync_micro_compact_to_db
skip its write, but _micro_compact still returned the spliced list, advanced the
cursor and reported "absorbed". finalize_turn's _persist_session flush then
appended the unmarked summary row as live, beside the winning generation.

The commit-time re-check also never fired in the common case: held_archive_watermark
returned early when newest_held >= the start watermark, which is always true when
nothing was appended. A lease-less caller (stale_raises=True) now checks the newest
held row's liveness regardless; the in-place lease path is unchanged.

_sync_micro_compact_to_db now reports a stale abort, and _micro_compact then returns
the original list, restores the rolling summary/cursor (and a defragged marker), and
emits stale_generation telemetry. _micro_start_watermark returns
(skip_outcome, watermark) so the two identical pre-flight skip branches collapse.
2026-09-24 18:01:21 +05:30
kshitijk4poor
40051c2a13 fix(compression): a stale prune/micro-compaction generation aborts instead of publishing beside the winner
Prune and micro-compaction hold no compression lease. When the newest exact row of the history they
hold is no longer active, another compaction (a /compress on this or another surface, or an earlier
pass) has already committed. `held_archive_watermark` then falls back to the start watermark, which is
right for the in-place commit (its lease rules out overlap) but for a lease-less caller archives the
winner's rows and clones them back as a "concurrent tail": two summary generations live in one session.

`_archive_watermark_for` now asks `held_archive_watermark` to raise `StaleHeldHistory` on that branch.
The prune commit returns its input unchanged (`prune:stale_generation`), micro-compaction checks before
the slow summary call and skips the pass (telemetry `stale_generation`), and its commit re-checks so a
compaction that lands during the summary call is not overwritten either. The in-place commit keeps its
fallback.

Raised on #120634 by @ehz0ah (agent/conversation_compression.py:3623 thread) and independently by
@JoaoMarcos44 in #120821.

Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 18:01:21 +05:30
John Paul Soliva
3f7e0fd072 fix(compression): proactive prune and micro-compaction keep turns another surface appended
Both commits call archive_and_compact with no watermark, which archives
every active row. A turn another surface appended to the same session
after this process loaded it (a Desktop session continued from Telegram)
was archived with the rest as compacted (active=0, compacted=1): marked
summarized away though no summary holds it. The display, REST with
include_compacted and session_search still show it, but it is gone from
the model's history, so the agent forgets a turn the user can still see.
Micro-compaction also runs a slow aux summary call before its commit, so
a turn that arrived during it went the same way.

Both now cap the archive at the newest row the process held, the rule the
in-place compaction commit applies, so those rows take the concurrent-
append path and are cloned after the new set. Micro-compaction captures
the held history and the store's watermark before the summary call and
caps at commit. The cap moves into held_archive_watermark(session_db,
session_id, ...), which the compressor paths can call; _held_watermark
stays as the in-place commit's wrapper.

A store without get_active_message_watermark or get_message_role keeps
today's archive-everything commit, as a store without archive_and_compact
already skips the prune. Micro-compaction's DB sync runs under a broad
except, so an unguarded call there would have failed silently.

Held lists without row ids (the gateway's replay dicts) keep today's
behaviour, as in the in-place commit. Both features are opt-in
(compression.proactive_prune_tokens > 0, compression.micro_compact).

(cherry picked from commit 8d67bee760db0fc65d4330df15198ddfd5f24480)
2026-09-24 18:01:21 +05:30
kshitijk4poor
4d5899868a fix(compression): a here-N tail copy vouches for its row id only when its source is unchanged
`/compress here N` left two live copies of a resume-merged user row when
the merged row was the newest tail row. Resume repair merges consecutive
user rows into the first dict (keeps its _row_id, drops the persist
marker; the second row's id leaves the held set). _held_watermark
exempted a verbatim tail from the marker check because its copies are
marker-swept by construction, assuming their ids were exact and newest.
That was false for the merged case: the cap landed on the merged row's
id and the absorbed durable row above it was cloned as a concurrent
append beside the row that already carries its text (main archives it).

Put the exactness decision where the marker is still visible: compress_now
keeps a tail copy's _row_id only when the source dict still carries the
marker. _held_watermark then applies one rule to every row (an id is
exact when its dict is unchanged since load: marker on a held row, a tail
copy's id trusted as given) and the newest row must have one, with no
tail exemption and max(held) computed once. `here 2` joins the existing
merge test; it is red without this change.
2026-09-24 15:22:15 +05:30
John Paul Soliva
1a632c3ab3 fix(compression): a held row a merge rewrote does not bound the archive
Resume repair merges consecutive user (and assistant) rows into the first
one's dict: that dict keeps its _row_id, drops the persisted marker, and
the later row's id leaves the held history. _held_watermark capped the
archive at max(held), so the later row sat above the cap, was cloned as a
concurrent append, and stayed live beside the merged row that already
carries its content. A context engine that rewrites the held list in
place reaches the same state.

A dict loaded from the DB is born carrying both the id and the marker, so
a newest held row with an id but no marker is exactly the rewritten case:
keep the lease watermark there, as main does. A here-N tail is exempt; it
is copied verbatim without the marker, its ids are exact, and being
newest they lift the cap above anything a merge earlier in the history
absorbed.

The new test reproduces the merge through get_resume_conversations, the
real resume path, rather than a hand-built dict. The existing
exact-duplicate check cannot see this case (the merged and cloned rows
are different strings), so it asserts the absorbed prompt appears in one
live row. Red on the previous head, green here.

(cherry picked from commit c41b9df15d72765d7d4fbb690077de463d37beed)
2026-09-24 15:22:15 +05:30
teknium1
f53c0f6b20 fix(compression): a trailing held row without any stamp keeps the lease watermark
_held_watermark skipped trailing dicts that carried neither `_row_id` nor the persisted marker
when picking the "newest held row", assuming such a row is not durable. The TUI model-switch
marker is: server.py appends the bare dict to session["history"] and writes it with a plain
append_message, stamping nothing. With that row last in history a plain /compress capped the
archive at the previous stamped row, so the marker's durable row sat above the cap, was cloned
as a "concurrent append" AND inserted from the compacted set: two live copies (probe on the PR
head: rows 32 and 33 both the marker; origin/main keeps one).

Any trailing dict of unknown provenance now disables the cap (lease watermark, today's
behaviour) instead of being looked past. The regression case rides the existing
never-leaves-two-live-copies test as a third parameter; red on the previous predicate.

Review-fix on #120156.

(cherry picked from commit b114f564f0eddb3ccfe16123e364097747562928)
2026-09-24 15:22:15 +05:30
John Paul Soliva
3028995817 fix(compression): an in-place compaction never archives turns the compacting surface did not hold
The in-place commit archives every active row up to the lease watermark, the newest row in
state.db. But a surface compacts the history it holds: the Desktop/TUI session.compress RPC
compacts session["history"], the CLI's /compress its conversation_history. When another surface
appended turns to the same session since (a Desktop session continued from Telegram after
/handoff, #42962, or a CLI resumed elsewhere), those rows sat under the watermark, took the
positional rewind slots and became active=0, compacted=0: gone from every surface's model and
display history, REST, the gateway's next replay and session_search. The summarizer never saw
them, and nothing re-adopts them. `/compress here N` (#119962) loses them the same way.

The watermark is now capped at the newest durable row the compressor was handed (its messages
plus the kept here-N tail). Rows above it take the existing concurrent-append path, cloned after
the compacted set. That is the premise _adopt_out_of_band_turns already rests on, so the next
prompt adopts them as before. The cap applies only while the held history is a live prefix of
the session: its newest durable row must carry its _row_id and still be active. Otherwise (a
history another surface already compacted, or rows held without ids, as the gateway's replay
dicts are) the commit falls back to the lease watermark, unchanged.

(cherry picked from commit 930866a2bbc2ec460898a0174c3600d97d8ed088)
2026-09-24 15:22:15 +05:30
kshitijk4poor
3bd93daedc refactor(compression): sample the stall wait once, before the retry chain
_recover_from_stall now samples both the total wait and seconds-since-progress
itself, before the fallback retry, so the "reached its total ceiling after Ns"
warning reports the stall rather than stall plus retry time (as before the
ladder refactor) and the three callers stop passing the same value. The
snapshot check trusts the worker's Tuple[list, str] contract like the rest of
the module, and the teardown test goes back to a 0.3s idle budget: the 2.0s
idle made the 0.3s ceiling dead and stretched the join grace to 2s.
2026-09-24 15:21:54 +05:30
kshitijk4poor
f0672a0c43 refactor(compression): single stall-fallback ladder for every stall exit
The settled-at-deadline no-op path, the post-cancel unchanged-commit path
and the idle-timeout path each carried their own copy of the same
retry-chain -> on_timeout -> degraded-prompt ladder (three copies after
the deterministic-fallback fix). Fold them into one closure,
_recover_from_stall, plus _is_unchanged_snapshot for the identity check.

Behavioural alignment for the two newer paths: when the retry chain yields
None they now return (messages, fallback prompt) like the idle path did,
instead of the worker's own tuple with its unset prompt, and they log the
same "no progress" warning when no on_timeout callback is installed.

Co-authored-by: Mark DiPietro <mark@svavnc.com>
2026-09-24 15:21:54 +05:30
Mark DiPietro
aabb7a609d fix(agent): make compression stall fallback deterministic
(cherry picked from commit 4ac18e27445b1c667f7f9165113709e0de9fb2e3)
2026-09-24 15:21:54 +05:30
Eva
71cca15ada fix(compression): a compressor no attempt claimed is never judged stale 2026-09-24 00:18:32 -07:00
kshitijk4poor
fdfe61994d fix(compaction): reinsert the delivered reply only where it keeps the transcript valid
Builds on #118900's reinsertion so the just-delivered reply survives a
compaction commit without breaking the transcript:

- The anchor lives in its own sibling, agent/conversation_compression_reply_anchor.py,
  and runs before the todo snapshot folds, so the reply never lands after
  the restated question.
- A reply whose text is already in the kept tail (whitespace or
  list-content form) is a twin: nothing is inserted, so it no longer
  shows twice.
- The slot is found from the reply's real followers: tool-call assistants
  by their call ids and tool rows by tool_call_id, both read through
  coalesce_tool_call_id so Codex-shaped rows key the same. Ids the transcript reuses
  (llama.cpp emits constant ids) are not trusted as anchors.
- Placement is best effort; alternation and chronology are invariants.
  An assistant neighbour on either side, an ambiguous slot, /compress here
  (verbatim tail), or an insert that would make the compaction stop
  shrinking all skip the reinsertion and log why.
- The helper returns the reinserted row (or None); each skip reason is
  logged where it is decided.
- The reinserted reply's durable row is carried through the commit, so its
  original is rewound rather than archived; only the reply is carried.
2026-09-23 23:06:15 +05:30
finn763
946b68865b fix(compaction): keep just-delivered assistant reply live across commit (#118900)
(cherry picked from commit 6518c7a1a7c076c5d7d3a04a775a19cef61660f7)
2026-09-23 23:06:15 +05:30
John Paul Soliva
3591bf3bb8 fix(compression): turn-start in-place compaction no longer hides the newest summarized turn (#120187)
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.

tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.
2026-09-23 10:05:30 -04:00
John Paul Soliva
526d135a96 fix(compression): keep the /compress here N tail in state.db under in-place compaction
compress_now() hands only the head to _compress_context() and rejoins the
kept exchanges in memory afterwards. With compression.in_place (the
default), the commit's archive_and_compact() archives every active row at
or below the lease watermark, which includes the kept tail's rows, and
inserts only the compacted head. Its tail_count rewind then lands on the
newest rows under the watermark, which are the kept tail itself, so those
rows end up active=0, compacted=0 with no live copy. No surface writes them
back: the CLI re-flushes only after a rotation, the gateway skips in-place
on purpose, and the TUI only swaps its in-memory history. The exchanges the
user asked to keep verbatim are gone on resume, and on the gateway from the
very next message.

compress_now() now passes copies of the tail as verbatim_tail. The in-place
commit stores head + tail in the same archive_and_compact() transaction,
joined with the same seam rejoin_compressed_head_and_tail() builds in
memory, and adds the tail to tail_count. The rewind flags then land on the
kept tail's originals and on compress()'s own carried rows, which were
left compacted=1 and shown twice in the resumed display history. The
copies are stamped as persisted and compress_now() returns the stored list
instead of rejoining the tail a second time. Rotation, no-op and
rolled-back commits are unchanged: the copies stay unstamped and the
caller's tail is rejoined as before.
2026-09-23 01:11:58 -07:00
kshitijk4poor
68f7b90eeb fix: return a raced-successful compression from _await_worker_within_budget
The `future.done()` guard added for dead workers (#117261 / #63892)
unconditionally logged `future.exception()` and returned `(False, None)`.
If the worker finished SUCCESSFULLY in the window between
`result(timeout=)` expiring and the `done()` check, that discarded a
completed compression and sent the caller down the stall/fallback path —
`future.cancel()` becomes a no-op on a settled future and the fallback
route costs a second LLM call. Siblings `_await_in_flight_commit` and
tool_executor's `_poll_sequential_future` already re-read the result on
a settled future; do the same here: `exception() is None` → return
`(True, result)`, otherwise take the stall path as before.

The compression-seam test grows a settled-successful case through the
same helper (a Future subclass whose first timed `result()` still raises
TimeoutError, modelling the race); it fails on the previous guard and
passes with this one.
2026-09-21 21:34:44 +05:30
kshitijk4poor
a8a81f4cb9 refactor: log the dead compression worker's exception once; trim guard comments
_await_worker_within_budget now emits one INFO line with future.exception()
when it takes the stall path for a worker that already died, so the log shows
WHY the fallback chain was entered instead of silently returning (False, None)
— previously the only trace was the absence of "still streaming" lines.

The three 5-7-line comment blocks added by the #117261 pick restated the same
alias fact each time; each is now 2 lines citing #63892 and the 3.11 alias once.
No behaviour change beyond the log line.
2026-09-21 21:34:44 +05:30
kshitijk4poor
7740a4ac20 fix: report a worker that died with TimeoutError as exited in _join_cancelled_worker
_join_cancelled_worker returned False from its `except
concurrent.futures.TimeoutError:` arm. On 3.11+ that class IS the builtin
TimeoutError, so the arm also fires when the cancelled worker itself DIED
raising a timeout-class error (the aux client raises bare TimeoutError on a
stalled summary stream). The caller, _release_cancelled_worker, treats False
as "still running": it logs 'did not exit within grace' and skips
fence.allow_cancelled_lock_release(), so the session compression lease of a
provably-dead worker was retained/orphaned.

Return future.done() instead: a settled future never becomes unsettled, so a
done future means the thread exited and the lease can be released. A live
worker that merely outlasted the grace still yields False.

Same guard family as the sibling loops fixed in the preceding pick (#117261).
2026-09-21 21:34:44 +05:30
Trevor Nash-Keller
8305113328 fix(compression): don't mistake a worker's TimeoutError for a poll timeout
On Python 3.11+ concurrent.futures.TimeoutError IS the builtin TimeoutError
(asyncio.TimeoutError and socket.timeout alias it too). Poll loops shaped like

    try:
        return future.result(timeout=slice)
    except concurrent.futures.TimeoutError:
        ...keep waiting...

therefore cannot distinguish "the wait slice expired" (worker alive) from "the
worker raised TimeoutError" (worker dead). auxiliary_client raises a bare
TimeoutError when a summary stream stalls, so this is reachable in production.

When it happened the host re-waited on an already-settled future. result() then
returned instantly every iteration, spinning at ~2k iterations/sec and logging
"Context compression still streaming" about a dead worker, until the entire idle
budget elapsed. One session burned 535s and wrote ~90k duplicate log lines
(15.6MiB) before failing with context_compression_timeout, and every later turn
re-entered the same path.

Guard each loop with future.done(): a settled future never becomes unsettled.
- _await_worker_within_budget: take the stall path at once, so the configured
  fallback chain is actually reached instead of after a 120s false stall.
- _await_in_flight_commit: re-raise the worker's exception. This loop had no
  ceiling, so a dead worker spun forever.
- tool_executor._poll_sequential_future: same, and with deadline=None it also
  span indefinitely.

Non-timeout worker exceptions still propagate unchanged, and a live worker still
polls exactly as before.

Regression tests pin all three. Against unpatched code the two wait tests fail
and the commit-wait test hangs until the 300s harness SIGKILL, reproducing the
infinite loop directly.

(cherry picked from commit 274457304ce0393407574fe4e43c4a450f20bac3)
2026-09-21 21:34:44 +05:30
Chukuwebuka-2003
7d3c0b2f94 fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.

Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.

Part of #116472 (request 4).
2026-09-20 18:53:10 -07:00
teknium1
0765099ff4 fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
2026-09-20 15:50:49 -07:00
beardthelion
0e33dc9ebc fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.

Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.

Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.

(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
2026-09-20 15:50:49 -07:00
teknium1
d03d6c2b39 fix(compression): an over-window session that cannot shrink ends the turn with /new guidance and waits one idle budget, not the ceiling
A session far above the model window (~356k tokens on a 131k window in
#116472) re-ran context compression on every turn: a preflight pass that
reclaimed nothing still let the request go to the provider (400 -> overflow
handler -> another pass), and a summary stream that kept emitting tokens
while never committing held the pre-commit wait to the full 600s ceiling.
On the Desktop that blocked the gateway event loop for 10-20 minutes per
turn and the renderer was eventually killed.

- agent/turn_context.py::_fail_closed_on_insufficient_progress: when a
  preflight pass makes no (or sub-5%) progress and the request provably
  exceeds the model window, raise PreflightCompressionTimedOut with
  "start a new session (/new)" guidance so no provider call is sent. An
  unknown window or a fitting request keeps the send-as-is behaviour; a
  pass that no-op'd on a transient guard (summary-failure cooldown) keeps
  its typed cooldown result. Called from both insufficient-progress
  branches of turn_context_compaction._run_preflight_passes.
- agent/conversation_compression.py::run_compress_context_with_progress_timeout:
  an over-window request's pre-commit wait is bounded by one inactivity
  budget (compression.context_timeout_seconds) instead of
  context_total_ceiling_seconds; the existing first-stall deterministic
  fallback then carries the compaction. Config-derived, no new knob.

Slim slice of #116592's Python half.

Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
2026-09-20 12:52:07 -07:00
fangliquanflq
af93d57ca7 fix(agent): preserve context when summary provider overloads 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
96030a3536 fix(compression): scope the checkpoint remediation to capability refusals; trim tests
The remediation ("set compression.checkpoint_required: false, or switch
provider") is right only when the gate can never pass with the active provider
set. Attached to every _checkpoint_blocked() it also decorated the transient
"provider checkpoint API v2 failed" path, where the provider IS capable and the
correct move is a retry once the store recovers, and the codex_app_server
refusal, where switching memory provider changes nothing. A small
_checkpoint_incapable() wrapper carries it on the two capability branches only.

Tests: six near-duplicates collapse into three invariants (warn on mismatch,
parametrized over v1 provider / no manager; silent when off or capable; the
compress-time refusal names the flag). Proven red with the fix neutralized.

Consolidates #106879 (KoNit-K) and #106882 (kokhlo); fixes #106870.
2026-09-18 13:58:52 -07:00
KoNit-K
8e7120e2fd fix(compression): debug-log checkpoint capability probe failures
Keep init fail-open on a broken probe, but emit DEBUG so flaky probes are diagnosable.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
KoNit-K
ef7c7f9e08 fix(compression): warn and remediate opaque checkpoint_required blocks
When compression.checkpoint_required is set but the active memory provider
lacks checkpoint API v2, keep fail-closed compress and surface a startup
warning plus actionable remediation on the block error.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 13:58:52 -07:00
teknium1
2d333e68b2 fix(compression): read the aux compression budget unguarded in resolve_context_compression_timeouts
_effective_aux_timeout is a pure config read used unguarded elsewhere; a
swallowed failure would silently fall back to the 120s idle window this
change calls structurally wrong.
2026-09-18 10:36:11 -07:00
teknium1
303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00
joaomarcos
f2bd6b4431 fix(compression): host idle watchdog never undercuts the aux summary request budget
The owned compress_context wait judged silence at compression.context_timeout_seconds (120s)
while the auxiliary compression request itself is allowed _effective_aux_timeout("compression")
(floor 300s). A summariser that legitimately produces no streamed token for >120s (non-streaming
route, reasoning model thinking before its first delta, long prompt build over thousands of rows)
was always cancelled as "no progress" before the provider layer would have given up.

Raise-only: the idle window is floored at the aux compression budget (never above the ceiling,
which is itself raised when the aux budget exceeds it); an explicit larger value is kept.

Salvaged from PR #114635 (@JoaoMarcos44); the emergency-prune hunk of that PR is dropped because
prune_tool_results_only is gated on proactive_prune_tokens (0 = off by default).
2026-09-18 10:36:11 -07:00
teknium1
a37293a2ce fix(compression): probe aux feasibility before the first compaction; clear stale clamp notice on un-clamp
A fresh AIAgent whose main model already had the larger window still got
its first compaction at the main-window threshold, and only then did the
lazy probe in compress_context clamp the trigger to the aux window — so
the oversized compaction request was already sent (#114707 steps 1-6;
every instance recycle restarted the cycle).

ensure_compression_feasibility_checked(agent, estimated_tokens) runs the
probe once a request first reaches MINIMUM_CONTEXT_LENGTH — the smallest
window any summariser may have — from both compaction gates
(run_preflight_compression, compress_after_tool_results), so the clamp
lands before the main-window threshold fires. Requests below that size
stay probe-free, preserving the cold-start deferral of 6cb9917c73
(#28957); a failed probe leaves the latch unset for the lazy probe.

Symmetric un-clamp: when a re-probe finds the aux fits again,
_compression_warning / _last_feasibility_notice are cleared so
replay_compression_warning does not resend an "auto-lowered" notice for
a session that is no longer clamped.
2026-09-18 09:49:14 -07:00
teknium1
9f5b7ea02e fix(compression): keep the aux ceiling and re-probe feasibility on every main-runtime change
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).

- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
  _apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
  only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
  every runtime change: switch_model (outside the rollback guard), fallback activation
  and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
  fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
  rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:49:14 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
07c92d675a fix(compression): repeated summary stall escalates to the deterministic fallback summary
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).

WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.

Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.

Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
2026-09-17 08:59:50 -07:00
joaomarcos
f59d973af6 fix(compression): stall cooldown covers one idle window; drifted-prompt INFO once per session
After a no-progress stall the ladder's first rung (60s) was shorter than the
default idle stall window (120s), so the next oversized automatic turn
re-entered the same silent summary route about a minute after burning the
whole window (#112420: "made no progress ... continuing without
compression" followed by "context compression started" 62s later).
record_timeout_failure now floors every timeout-class cooldown at the
configured idle window; a window below the rung leaves the ladder untouched.

The "Compaction rebuilt a drifted system prompt" line fired on every compact
of a long session (19x/day observed). The rebuild stays mandatory; only the
first drift per session is INFO, later ones DEBUG.

Both hunks re-applied from PR #112504 (@JoaoMarcos44); its no-LLM prune on
stall is a product decision and was not ported.

Part of #112420
2026-09-16 17:12:33 -07:00
KoNit-K
8e64cc3347 fix(compression): retry stalled fallback in the same turn 2026-09-16 17:12:33 -07:00