The retain early return coerced agent._cached_system_prompt with `or ""` and
wrote it back. The only setter of _retain_seeded_system_prompt,
gateway.run._seed_hygiene_system_prompt, always stores a str ("" or the
stored prompt) in the same call, and nothing between the seed and the commit
boundary resets it on the detached hygiene / manual-compress agent (the only
None writers are invalidate_system_prompt, which this branch returns before,
and switch_model, which those agents never run). So the normalize/write-back
was dead defence; the empty-seed case still returns "".
The retain branch in _rebuild_system_prompt_at_boundary returns before
_refresh_agent_tool_definitions. That is intentional: refresh_agent_mcp_tools
-> persist_agent_tool_names would otherwise pin the live session row's
tools[] to the detached hygiene agent's memory-only toolset. Requested
inline by teknium on #122825 so a later refactor does not move the refresh
above the early return.
Gateway hygiene and the gateway /compress handler compact with a detached
AIAgent(enabled_toolsets=["memory"]) seeded with the session's stored prompt
(_seed_hygiene_system_prompt, 76a17046e2 / 678916b427). Since #98426 the
commit boundary always rebuilds the prompt, so the seed was discarded and the
reduced-toolset build was persisted over the live session's snapshot: the
skills index (## Skills (mandatory) / <available_skills>) and Skill Safety
Rule vanish, and every later fresh agent for that session restores the
degraded bytes verbatim (_stored_prompt_matches_runtime does not compare
platform).
The seed now marks the agent (_retain_seeded_system_prompt) and the
commit-boundary rebuild keeps the seeded bytes for it. A live agent's own
compaction still rebuilds, so builder updates keep reaching long sessions.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit dd5c8ef4e845b2d34e826e4d92afad73a12eefa6)
At turn start the child prompt is checked against the (possibly fallback)
runtime before _restore_primary_runtime runs. A false reject only leaves the
slot None, and the normal restore re-reads the row and re-checks it, so the
order is intentional; say so, so nobody reorders it later.
The clear-then-conditionally-set sequence and its two overlapping comment
blocks said the same thing twice. One conditional assignment keeps the
behaviour identical (None when the child prompt is empty or fails the
runtime-identity check, the child prompt otherwise) and is easier to read.
_adopt_live_compression_child moves agent.session_id onto the live compression child,
but when the child's stored prompt fails _stored_prompt_matches_runtime the parent's
non-null _cached_system_prompt stayed in place. turn_context restores/rebuilds only
while that slot is None, so the child turn could send the parent session's prompt with
no validation. Clear the slot right after child confirmation, before the conditional
seed; a rejected tip then rebuilds through the canonical restore path and a matching
tip is still adopted verbatim.
The regression test now seeds a NON-NULL parent cache — the previous null-path-only
seed hid the leak — and asserts the rejected-tip agent carries session_id="child"
with the slot cleared (red on the previous head).
Review fix for #121843 (ehz0ah, agent/conversation_compression.py:1703).
(cherry picked from commit 531a5072ff4a2e74bd7003094786e32136d2545b)
_adopt_live_compression_child seeds agent._cached_system_prompt straight from the tip's row,
and turn_context only runs _restore_or_build_system_prompt while that slot is None. Until now
a route commit NULLed the row, so a tip whose model moved was never seeded; with the writers
preserving the snapshot, the tip carries the old Model:/Provider: trailer and would be served
to the new model with no check at all. Gate the seed on _stored_prompt_matches_runtime, the
same predicate the restore path applies; an unseeded slot rebuilds on the next restore.
Review fix for #121843 (found in review of the widened invariant; probe: parent ended by
compression, child /model-committed to model-b, adopting agent on model-b was seeded
"Model: model-a" on the PR head and nothing on main).
(cherry picked from commit d7143de9ce7f43201b17e8df5c0344aee777b7fc)
The Codex app-server compaction path was the only caller passing
skills=False; now that it re-arms skill_view dedup like every other
rewrite boundary, the kwarg and its early return are dead. Removing them
keeps the helper's contract simple: every boundary resets both caches,
so a future caller cannot reintroduce the #101518 asymmetry by accident.
skill_view returns a stub on a repeat view of an unchanged file, pointing
the model at the earlier full result in its own context. That contract only
holds while the earlier result is still there, so a compaction boundary has
to advance the dedup generation -- which is what _reset_read_dedup_caches()
is for, and why its stub message promises "Re-issued after context
compression, this returns the full content again."
The Codex app-server path passed skills=False and reset only the file-read
half. A skill re-viewed after that compaction got a stub naming content the
compaction had just summarised away, so the model had nothing to copy from
and reconstructed it from memory instead. Observed with a delegating skill
whose child contract is a JSON schema: successive dispatches shipped
progressively degraded schemas -- a dropped allOf, then a bare
{"type": "array"} -- until the generation-time check no longer constrained
anything and the work was abandoned.
The opt-out carried the pre-refactor asymmetry forward mechanically
(0e9d46511a lifted these call sites without changing behaviour); nothing
depends on it. Drop it so both paths reset both caches. Tests cover the
helper and drive the codex path end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: aurorabotticus-svg <aurorabotticus-svg@users.noreply.github.com>
(cherry picked from commit 22f36fe10ec27bf68e003058b3f1dda292e535e8)
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.
After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.
Fixes#123167
(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
A gap below the newest held id, and turns appended above an unpersisted current
turn, were summarized away without being read. Name those held ids and clone
the rest.
A compaction landing during the micro summary call made _sync_micro_compact_to_db
skip its write, but _micro_compact still returned the spliced list, advanced the
cursor and reported "absorbed". finalize_turn's _persist_session flush then
appended the unmarked summary row as live, beside the winning generation.
The commit-time re-check also never fired in the common case: held_archive_watermark
returned early when newest_held >= the start watermark, which is always true when
nothing was appended. A lease-less caller (stale_raises=True) now checks the newest
held row's liveness regardless; the in-place lease path is unchanged.
_sync_micro_compact_to_db now reports a stale abort, and _micro_compact then returns
the original list, restores the rolling summary/cursor (and a defragged marker), and
emits stale_generation telemetry. _micro_start_watermark returns
(skip_outcome, watermark) so the two identical pre-flight skip branches collapse.
Prune and micro-compaction hold no compression lease. When the newest exact row of the history they
hold is no longer active, another compaction (a /compress on this or another surface, or an earlier
pass) has already committed. `held_archive_watermark` then falls back to the start watermark, which is
right for the in-place commit (its lease rules out overlap) but for a lease-less caller archives the
winner's rows and clones them back as a "concurrent tail": two summary generations live in one session.
`_archive_watermark_for` now asks `held_archive_watermark` to raise `StaleHeldHistory` on that branch.
The prune commit returns its input unchanged (`prune:stale_generation`), micro-compaction checks before
the slow summary call and skips the pass (telemetry `stale_generation`), and its commit re-checks so a
compaction that lands during the summary call is not overwritten either. The in-place commit keeps its
fallback.
Raised on #120634 by @ehz0ah (agent/conversation_compression.py:3623 thread) and independently by
@JoaoMarcos44 in #120821.
Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Both commits call archive_and_compact with no watermark, which archives
every active row. A turn another surface appended to the same session
after this process loaded it (a Desktop session continued from Telegram)
was archived with the rest as compacted (active=0, compacted=1): marked
summarized away though no summary holds it. The display, REST with
include_compacted and session_search still show it, but it is gone from
the model's history, so the agent forgets a turn the user can still see.
Micro-compaction also runs a slow aux summary call before its commit, so
a turn that arrived during it went the same way.
Both now cap the archive at the newest row the process held, the rule the
in-place compaction commit applies, so those rows take the concurrent-
append path and are cloned after the new set. Micro-compaction captures
the held history and the store's watermark before the summary call and
caps at commit. The cap moves into held_archive_watermark(session_db,
session_id, ...), which the compressor paths can call; _held_watermark
stays as the in-place commit's wrapper.
A store without get_active_message_watermark or get_message_role keeps
today's archive-everything commit, as a store without archive_and_compact
already skips the prune. Micro-compaction's DB sync runs under a broad
except, so an unguarded call there would have failed silently.
Held lists without row ids (the gateway's replay dicts) keep today's
behaviour, as in the in-place commit. Both features are opt-in
(compression.proactive_prune_tokens > 0, compression.micro_compact).
(cherry picked from commit 8d67bee760db0fc65d4330df15198ddfd5f24480)
`/compress here N` left two live copies of a resume-merged user row when
the merged row was the newest tail row. Resume repair merges consecutive
user rows into the first dict (keeps its _row_id, drops the persist
marker; the second row's id leaves the held set). _held_watermark
exempted a verbatim tail from the marker check because its copies are
marker-swept by construction, assuming their ids were exact and newest.
That was false for the merged case: the cap landed on the merged row's
id and the absorbed durable row above it was cloned as a concurrent
append beside the row that already carries its text (main archives it).
Put the exactness decision where the marker is still visible: compress_now
keeps a tail copy's _row_id only when the source dict still carries the
marker. _held_watermark then applies one rule to every row (an id is
exact when its dict is unchanged since load: marker on a held row, a tail
copy's id trusted as given) and the newest row must have one, with no
tail exemption and max(held) computed once. `here 2` joins the existing
merge test; it is red without this change.
Resume repair merges consecutive user (and assistant) rows into the first
one's dict: that dict keeps its _row_id, drops the persisted marker, and
the later row's id leaves the held history. _held_watermark capped the
archive at max(held), so the later row sat above the cap, was cloned as a
concurrent append, and stayed live beside the merged row that already
carries its content. A context engine that rewrites the held list in
place reaches the same state.
A dict loaded from the DB is born carrying both the id and the marker, so
a newest held row with an id but no marker is exactly the rewritten case:
keep the lease watermark there, as main does. A here-N tail is exempt; it
is copied verbatim without the marker, its ids are exact, and being
newest they lift the cap above anything a merge earlier in the history
absorbed.
The new test reproduces the merge through get_resume_conversations, the
real resume path, rather than a hand-built dict. The existing
exact-duplicate check cannot see this case (the merged and cloned rows
are different strings), so it asserts the absorbed prompt appears in one
live row. Red on the previous head, green here.
(cherry picked from commit c41b9df15d72765d7d4fbb690077de463d37beed)
_held_watermark skipped trailing dicts that carried neither `_row_id` nor the persisted marker
when picking the "newest held row", assuming such a row is not durable. The TUI model-switch
marker is: server.py appends the bare dict to session["history"] and writes it with a plain
append_message, stamping nothing. With that row last in history a plain /compress capped the
archive at the previous stamped row, so the marker's durable row sat above the cap, was cloned
as a "concurrent append" AND inserted from the compacted set: two live copies (probe on the PR
head: rows 32 and 33 both the marker; origin/main keeps one).
Any trailing dict of unknown provenance now disables the cap (lease watermark, today's
behaviour) instead of being looked past. The regression case rides the existing
never-leaves-two-live-copies test as a third parameter; red on the previous predicate.
Review-fix on #120156.
(cherry picked from commit b114f564f0eddb3ccfe16123e364097747562928)
The in-place commit archives every active row up to the lease watermark, the newest row in
state.db. But a surface compacts the history it holds: the Desktop/TUI session.compress RPC
compacts session["history"], the CLI's /compress its conversation_history. When another surface
appended turns to the same session since (a Desktop session continued from Telegram after
/handoff, #42962, or a CLI resumed elsewhere), those rows sat under the watermark, took the
positional rewind slots and became active=0, compacted=0: gone from every surface's model and
display history, REST, the gateway's next replay and session_search. The summarizer never saw
them, and nothing re-adopts them. `/compress here N` (#119962) loses them the same way.
The watermark is now capped at the newest durable row the compressor was handed (its messages
plus the kept here-N tail). Rows above it take the existing concurrent-append path, cloned after
the compacted set. That is the premise _adopt_out_of_band_turns already rests on, so the next
prompt adopts them as before. The cap applies only while the held history is a live prefix of
the session: its newest durable row must carry its _row_id and still be active. Otherwise (a
history another surface already compacted, or rows held without ids, as the gateway's replay
dicts are) the commit falls back to the lease watermark, unchanged.
(cherry picked from commit 930866a2bbc2ec460898a0174c3600d97d8ed088)
_recover_from_stall now samples both the total wait and seconds-since-progress
itself, before the fallback retry, so the "reached its total ceiling after Ns"
warning reports the stall rather than stall plus retry time (as before the
ladder refactor) and the three callers stop passing the same value. The
snapshot check trusts the worker's Tuple[list, str] contract like the rest of
the module, and the teardown test goes back to a 0.3s idle budget: the 2.0s
idle made the 0.3s ceiling dead and stretched the join grace to 2s.
The settled-at-deadline no-op path, the post-cancel unchanged-commit path
and the idle-timeout path each carried their own copy of the same
retry-chain -> on_timeout -> degraded-prompt ladder (three copies after
the deterministic-fallback fix). Fold them into one closure,
_recover_from_stall, plus _is_unchanged_snapshot for the identity check.
Behavioural alignment for the two newer paths: when the retry chain yields
None they now return (messages, fallback prompt) like the idle path did,
instead of the worker's own tuple with its unset prompt, and they log the
same "no progress" warning when no on_timeout callback is installed.
Co-authored-by: Mark DiPietro <mark@svavnc.com>
Builds on #118900's reinsertion so the just-delivered reply survives a
compaction commit without breaking the transcript:
- The anchor lives in its own sibling, agent/conversation_compression_reply_anchor.py,
and runs before the todo snapshot folds, so the reply never lands after
the restated question.
- A reply whose text is already in the kept tail (whitespace or
list-content form) is a twin: nothing is inserted, so it no longer
shows twice.
- The slot is found from the reply's real followers: tool-call assistants
by their call ids and tool rows by tool_call_id, both read through
coalesce_tool_call_id so Codex-shaped rows key the same. Ids the transcript reuses
(llama.cpp emits constant ids) are not trusted as anchors.
- Placement is best effort; alternation and chronology are invariants.
An assistant neighbour on either side, an ambiguous slot, /compress here
(verbatim tail), or an insert that would make the compaction stop
shrinking all skip the reinsertion and log why.
- The helper returns the reinserted row (or None); each skip reason is
logged where it is decided.
- The reinserted reply's durable row is carried through the commit, so its
original is rewound rather than archived; only the reply is carried.
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.
tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.
compress_now() hands only the head to _compress_context() and rejoins the
kept exchanges in memory afterwards. With compression.in_place (the
default), the commit's archive_and_compact() archives every active row at
or below the lease watermark, which includes the kept tail's rows, and
inserts only the compacted head. Its tail_count rewind then lands on the
newest rows under the watermark, which are the kept tail itself, so those
rows end up active=0, compacted=0 with no live copy. No surface writes them
back: the CLI re-flushes only after a rotation, the gateway skips in-place
on purpose, and the TUI only swaps its in-memory history. The exchanges the
user asked to keep verbatim are gone on resume, and on the gateway from the
very next message.
compress_now() now passes copies of the tail as verbatim_tail. The in-place
commit stores head + tail in the same archive_and_compact() transaction,
joined with the same seam rejoin_compressed_head_and_tail() builds in
memory, and adds the tail to tail_count. The rewind flags then land on the
kept tail's originals and on compress()'s own carried rows, which were
left compacted=1 and shown twice in the resumed display history. The
copies are stamped as persisted and compress_now() returns the stored list
instead of rejoining the tail a second time. Rotation, no-op and
rolled-back commits are unchanged: the copies stay unstamped and the
caller's tail is rejoined as before.
The `future.done()` guard added for dead workers (#117261 / #63892)
unconditionally logged `future.exception()` and returned `(False, None)`.
If the worker finished SUCCESSFULLY in the window between
`result(timeout=)` expiring and the `done()` check, that discarded a
completed compression and sent the caller down the stall/fallback path —
`future.cancel()` becomes a no-op on a settled future and the fallback
route costs a second LLM call. Siblings `_await_in_flight_commit` and
tool_executor's `_poll_sequential_future` already re-read the result on
a settled future; do the same here: `exception() is None` → return
`(True, result)`, otherwise take the stall path as before.
The compression-seam test grows a settled-successful case through the
same helper (a Future subclass whose first timed `result()` still raises
TimeoutError, modelling the race); it fails on the previous guard and
passes with this one.
_await_worker_within_budget now emits one INFO line with future.exception()
when it takes the stall path for a worker that already died, so the log shows
WHY the fallback chain was entered instead of silently returning (False, None)
— previously the only trace was the absence of "still streaming" lines.
The three 5-7-line comment blocks added by the #117261 pick restated the same
alias fact each time; each is now 2 lines citing #63892 and the 3.11 alias once.
No behaviour change beyond the log line.
_join_cancelled_worker returned False from its `except
concurrent.futures.TimeoutError:` arm. On 3.11+ that class IS the builtin
TimeoutError, so the arm also fires when the cancelled worker itself DIED
raising a timeout-class error (the aux client raises bare TimeoutError on a
stalled summary stream). The caller, _release_cancelled_worker, treats False
as "still running": it logs 'did not exit within grace' and skips
fence.allow_cancelled_lock_release(), so the session compression lease of a
provably-dead worker was retained/orphaned.
Return future.done() instead: a settled future never becomes unsettled, so a
done future means the thread exited and the lease can be released. A live
worker that merely outlasted the grace still yields False.
Same guard family as the sibling loops fixed in the preceding pick (#117261).
On Python 3.11+ concurrent.futures.TimeoutError IS the builtin TimeoutError
(asyncio.TimeoutError and socket.timeout alias it too). Poll loops shaped like
try:
return future.result(timeout=slice)
except concurrent.futures.TimeoutError:
...keep waiting...
therefore cannot distinguish "the wait slice expired" (worker alive) from "the
worker raised TimeoutError" (worker dead). auxiliary_client raises a bare
TimeoutError when a summary stream stalls, so this is reachable in production.
When it happened the host re-waited on an already-settled future. result() then
returned instantly every iteration, spinning at ~2k iterations/sec and logging
"Context compression still streaming" about a dead worker, until the entire idle
budget elapsed. One session burned 535s and wrote ~90k duplicate log lines
(15.6MiB) before failing with context_compression_timeout, and every later turn
re-entered the same path.
Guard each loop with future.done(): a settled future never becomes unsettled.
- _await_worker_within_budget: take the stall path at once, so the configured
fallback chain is actually reached instead of after a 120s false stall.
- _await_in_flight_commit: re-raise the worker's exception. This loop had no
ceiling, so a dead worker spun forever.
- tool_executor._poll_sequential_future: same, and with deadline=None it also
span indefinitely.
Non-timeout worker exceptions still propagate unchanged, and a live worker still
polls exactly as before.
Regression tests pin all three. Against unpatched code the two wait tests fail
and the commit-wait test hangs until the 300s harness SIGKILL, reproducing the
infinite loop directly.
(cherry picked from commit 274457304ce0393407574fe4e43c4a450f20bac3)
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.
Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.
Part of #116472 (request 4).
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.
Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.
Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.
(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
A session far above the model window (~356k tokens on a 131k window in
#116472) re-ran context compression on every turn: a preflight pass that
reclaimed nothing still let the request go to the provider (400 -> overflow
handler -> another pass), and a summary stream that kept emitting tokens
while never committing held the pre-commit wait to the full 600s ceiling.
On the Desktop that blocked the gateway event loop for 10-20 minutes per
turn and the renderer was eventually killed.
- agent/turn_context.py::_fail_closed_on_insufficient_progress: when a
preflight pass makes no (or sub-5%) progress and the request provably
exceeds the model window, raise PreflightCompressionTimedOut with
"start a new session (/new)" guidance so no provider call is sent. An
unknown window or a fitting request keeps the send-as-is behaviour; a
pass that no-op'd on a transient guard (summary-failure cooldown) keeps
its typed cooldown result. Called from both insufficient-progress
branches of turn_context_compaction._run_preflight_passes.
- agent/conversation_compression.py::run_compress_context_with_progress_timeout:
an over-window request's pre-commit wait is bounded by one inactivity
budget (compression.context_timeout_seconds) instead of
context_total_ceiling_seconds; the existing first-stall deterministic
fallback then carries the compaction. Config-derived, no new knob.
Slim slice of #116592's Python half.
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).
Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The remediation ("set compression.checkpoint_required: false, or switch
provider") is right only when the gate can never pass with the active provider
set. Attached to every _checkpoint_blocked() it also decorated the transient
"provider checkpoint API v2 failed" path, where the provider IS capable and the
correct move is a retry once the store recovers, and the codex_app_server
refusal, where switching memory provider changes nothing. A small
_checkpoint_incapable() wrapper carries it on the two capability branches only.
Tests: six near-duplicates collapse into three invariants (warn on mismatch,
parametrized over v1 provider / no manager; silent when off or capable; the
compress-time refusal names the flag). Proven red with the fix neutralized.
Consolidates #106879 (KoNit-K) and #106882 (kokhlo); fixes#106870.
When compression.checkpoint_required is set but the active memory provider
lacks checkpoint API v2, keep fail-closed compress and surface a startup
warning plus actionable remediation on the block error.
Co-authored-by: Cursor <cursoragent@cursor.com>
_effective_aux_timeout is a pure config read used unguarded elsewhere; a
swallowed failure would silently fall back to the 120s idle window this
change calls structurally wrong.
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
The owned compress_context wait judged silence at compression.context_timeout_seconds (120s)
while the auxiliary compression request itself is allowed _effective_aux_timeout("compression")
(floor 300s). A summariser that legitimately produces no streamed token for >120s (non-streaming
route, reasoning model thinking before its first delta, long prompt build over thousands of rows)
was always cancelled as "no progress" before the provider layer would have given up.
Raise-only: the idle window is floored at the aux compression budget (never above the ceiling,
which is itself raised when the aux budget exceeds it); an explicit larger value is kept.
Salvaged from PR #114635 (@JoaoMarcos44); the emergency-prune hunk of that PR is dropped because
prune_tool_results_only is gated on proactive_prune_tokens (0 = off by default).
A fresh AIAgent whose main model already had the larger window still got
its first compaction at the main-window threshold, and only then did the
lazy probe in compress_context clamp the trigger to the aux window — so
the oversized compaction request was already sent (#114707 steps 1-6;
every instance recycle restarted the cycle).
ensure_compression_feasibility_checked(agent, estimated_tokens) runs the
probe once a request first reaches MINIMUM_CONTEXT_LENGTH — the smallest
window any summariser may have — from both compaction gates
(run_preflight_compression, compress_after_tool_results), so the clamp
lands before the main-window threshold fires. Requests below that size
stay probe-free, preserving the cold-start deferral of 6cb9917c73
(#28957); a failed probe leaves the latch unset for the lazy probe.
Symmetric un-clamp: when a re-probe finds the aux fits again,
_compression_warning / _last_feasibility_notice are cleared so
replay_compression_warning does not resend an "auto-lowered" notice for
a session that is no longer clamped.
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).
- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
_apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
every runtime change: switch_model (outside the rollback guard), fallback activation
and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).
- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
inherited UI transport label without an explicit --source) resolves to `oneshot`; the
platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
(HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
`sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
`--source` flag reference (explicit flag always stored as given).
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
After a no-progress stall the ladder's first rung (60s) was shorter than the
default idle stall window (120s), so the next oversized automatic turn
re-entered the same silent summary route about a minute after burning the
whole window (#112420: "made no progress ... continuing without
compression" followed by "context compression started" 62s later).
record_timeout_failure now floors every timeout-class cooldown at the
configured idle window; a window below the rung leaves the ladder untouched.
The "Compaction rebuilt a drifted system prompt" line fired on every compact
of a long session (19x/day observed). The rebuild stays mandatory; only the
first drift per session is INFO, later ones DEBUG.
Both hunks re-applied from PR #112504 (@JoaoMarcos44); its no-LLM prune on
stall is a product decision and was not ported.
Part of #112420