Three read_text() calls in test_entry_stores_only_changed_paths_and_rollback_still_restores
gained encoding="utf-8" alongside the gc_blobs fix. They are unrelated
Windows hygiene (the footgun checker does not flag them); keep the PR
scoped to the blob-GC read-failure bug.
compact_ledger() caught OSError on the read but decoded the bytes outside
the guard, so a ledger with invalid UTF-8 raised UnicodeDecodeError out of
`hermes curator ledger --compact`. Same escape class the gc_blobs() fix
closes; treat it identically (no-op, ledger left untouched).
The new read-failure branch in gc_blobs() returned (0, 0) silently, so
`hermes curator ledger --compact` printed "0 unreferenced blob(s) removed"
as if the sweep had run. Log a warning with the underlying error so an
operator can tell "nothing to reap" from "could not look".
The kept invariant test now also asserts the warning via caplog.
Remove `monkeypatch.setattr(audio.shutil, "which", lambda _name: None)`
from test_caf_conversion_preserves_neighbors_and_removes_owned_output
(tests/tools/test_transcription_tools.py:1360).
Why it is dead: `_convert_caf_to_wav` (tools/transcription_audio.py:168-177)
tries candidates in order ffmpeg-then-afconvert and returns on the first
success. The test already patches `_find_ffmpeg_binary -> "ffmpeg"` and
`_run_quiet -> encode`, which never raises, so the ffmpeg candidate always
succeeds and the afconvert candidate is never executed. Whether
`shutil.which("afconvert")` returns a path or None cannot change the
outcome. The patch also mutated the process-global `shutil.which` for the
test's duration for no benefit.
Proof: `-k Caf` selection is 4 passed before and after.
Pass ignore_cleanup_errors=True to the TemporaryDirectory that holds the
converted CAF->WAV file (tools/transcription_tools.py:420).
WHY: the sibling trim-dir cleanup a few lines below already swallows
errors (shutil.rmtree(..., ignore_errors=True)), but the CAF work dir did
not. On Windows an AV scanner or the search indexer briefly holding the
freshly written WAV makes rmtree raise out of TemporaryDirectory.__exit__,
turning an otherwise successful transcription into an exception. The
kwarg is available since 3.10; pyproject requires-python >= 3.11, and
plugins/teams_pipeline/meetings.py:199 already uses it.
No new test: the behaviour is stdlib.
The invariant (source .caf untouched, sibling .wav not overwritten, owned
output removed) belongs beside the existing CAF tests rather than in a
new file. The conversion-error case exercised the same ExitStack
return-inside-with path as provider-error, so it is dropped to stay at
two cases. Also reverts the unrelated encoding="utf-8" drive-by on the
.env fixture in TestTranscribeCredentialReadGuard.
`_finish` (agent/context_compressor.py:3586-3588) only wrapped
`_coverage` and re-derived `shown` from the final `selected`; it was
called once, at the return. With `_merged` hoisted above the slice loop
the classmethod now has two closures (`_merged`, `_render`) plus
`_coverage` instead of four. Same expressions, same evaluation order
(`_render` first, then the coverage counters), so the output and
coverage dict are byte-identical — verified on the 47-shape capture and
probe_C2 (0 violations, base identity x4 True).
tests/agent/test_context_compressor.py:3397 asserted
`sampled.count("chars elided") < 7`, hardcoding the n-1 gap count that
the 8-slice constant implies and that the comment explained in prose.
Use `ContextCompressor._SAMPLED_INPUT_SLICES - 1` so the invariant
("the extension pass closed at least one initial gap") survives a
change to the slice count instead of silently becoming a change
detector. `_SAMPLED_INPUT_SLICES - 1 == 7` today, so the assertion
value is unchanged.
`_bound_oversized_record` (agent/context_compressor.py:3485-3488) split
`remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded
the tail slice with `if tail_len else ""`. `remaining = limit -
marker_reserve` is >= 1 because the line above already returned when
`limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1`
and `tail_len = remaining - head_len >= 1` always: the guard can never
fire. Use `remaining // 2` and drop the dead guard.
Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative
int (checked exhaustively for r in 0..10**6), and the guard was dead.
Capture of 800 bounded-record cases (4 record shapes x remaining 1..200)
plus the 47-shape sampling capture is identical before/after.
The `if cursor < len(records)` block at the end of `_render`
(agent/context_compressor.py:3580-3583) duplicated the gap-marker
arithmetic of the in-loop branch but had no path to it.
Invariant: `_render` is only reached with non-empty `records`
(empty returns early), and n >= 1, so the slice-build loop always runs
its last iteration, which anchors `end = len(records)` and
`start = end - 1 >= 0`, hence `end > start` and the final slice ends at
`len(records)`. `_merged` keeps `max(end)` for the last interval, and the
extension pass only grows the newest slice backward (`(s - 1, e)`) or
older slices forward but never past the next slice's start, so the last
slice's end stays `len(records)`. Therefore `cursor == len(records)`
after the loop for every input.
Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge
shapes incl. single-record and empty); probe_C2 -> 0 violations;
under-cap base identity x4 True.
The slice-build loop in _sample_summary_records hand-rolled the same
"overlapping/touching slice -> extend previous" fold that the `_merged`
closure 25 lines below implements (agent/context_compressor.py:3563-3568
vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end)
pairs, and merge once before the extension pass; the first `_render`
no longer wraps `selected` in a redundant `_merged` since it is already
merged. One interval-merge implementation instead of two.
Output is byte-identical: the merge is applied once to the same ordered
slice list, so the extension pass starts from the same `selected`.
Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes)
diffed before/after: identical sampled text and coverage; probe_C2
40 shapes -> 0 violations; under-cap base identity x4 True.
No fixture exercised the final `selected = _merged(selected)` in
_sample_summary_records: the re-gate mutant that dropped it survived
all 8 bounding tests and 40 random probe shapes. Add 20 x 8.5K records
(170K > 160K cap, headroom > one 8.5K gap) so the round-robin extension
closes the gaps between the newer slices. Assert consecutive shown
records are joined by exactly "\n\n" (an unmerged adjacent pair renders
as "...xxx[USER]: record-..." with no separator), remaining gaps carry a
correctly numbered marker, fewer than 7 markers survive, and
sampled_record_count equals the whole records shown.
RED with the final `_merged` call removed: AssertionError (14, 15);
GREEN at head.
Finding: $D/regate/C.md suggestion 2 (M-c, agent/context_compressor.py:3623).
The comment claimed "marker widths shrink as records leave a gap" — but
shrinkage can only make the render smaller, which the pre-check already
tolerates. The exact re-render exists for the opposite case: a gap's
first index can gain a digit or thousands separator (999 -> 1,000, +2
chars) when the added record moves a marker, and that growth can push a
render sitting at cap over it. Comment only; no behaviour change.
Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619).
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.
Before -> after (records per slice, fill unchanged):
10000x100 [1,1,1,1,1,1,1,8] -> [1,2,2,2,2,2,2,2] fill 0.944
8000x100 [2,2,2,2,2,2,2,5] -> [2,2,2,2,2,3,3,3] fill 0.957
4000x300 [4,4,4,4,4,4,4,11] -> [4,5,5,5,5,5,5,5] fill 0.987
2000x600 [9,9,9,9,9,9,9,15] -> [9,9,10,10,10,10,10,10] fill 0.995
500x2000 [37,...,37,39] -> [37,...,37,38,38] fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.
Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).
Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
tests/agent/test_lean_single_aux_call.py:163 accepted either `<seg09>`
or a `content[-500:]` substring. The last slice is anchored to the
newest record, so the fallback can never be what passes; assert
`"<seg09>" in out` outright and drop the now-unused `content`.
Red when the tail anchor is defeated (end = len(records) - 1): 1 failed;
green at head (gate 2c suggestion).
The coverage docstring (agent/context_compressor.py:3510-3511) said the
counters count "record content only", but `sampled_chars` sums the
*display* records (post `_bound_oversized_record` truncation) while
`input_chars` sums the raw records. Spell that out so telemetry
consumers do not compute `omitted` two different ways (gate 2c
suggestion, L3512-3513/3524-3525).
`_begin_compression_telemetry` (agent/context_compressor.py:1947-1960)
did not seed the six `summary_input_*` counters that
`_record_summary_input_coverage` adds on the lean path, so the emitted
`compression_attempt` line (conversation_compression.py:1451-1470) had a
different key set for lean and legacy attempts. Seed them as None so the
schema is stable; probe: lean and legacy attempts now emit identical key
sets (gate 2c suggestion, L1947-1960).
The whole-record greedy fill in `_sample_summary_records`
(agent/context_compressor.py:3541,3555-3563) stops a slice as soon as the
next record would cross `target` (~20K), so each slice can be short by up
to one record. With mid-sized records that wastes a large share of the
160K cap: probe worst case (10K records) filled 90,660/160,000 (57%),
8K records 129,021 (81%); main filled 159,988 by splitting records.
Add a bounded extension pass after `selected` is built: newest slice
grows backward, older slices grow forward, one whole record at a time,
only while the exact rendered length stays <= cap and the slice does not
run into its neighbour (adjacent slices are merged so no separator is
lost). Records are never split and the cap is never exceeded.
After: worst random shape 150,581 (94%, remaining headroom < one 10K
record), 8K records 153,117 (96%), 1551-record region 159,996 with
968 whole records, 0 partial, 8 buckets within +-10%, telemetry
input_chars == sampled_chars + omitted_chars (probe: $D/fold/probe_C_fold.py).
The existing keeps_record_boundaries fixture (1.2K records) fills 99.1%
without this pass, so no fill assert was added there — it would be toothless.
Delete the post-render overflow-trim loop (agent/context_compressor.py
:3591-3607) and state the invariant in its place.
Why it is unreachable: every display record is pre-bounded to `target`
via _bound_oversized_record; each slice's greedy fill stops before
`size + sep + next > target`, so a slice renders to <= target chars and
the n slices to <= n*target (a merge of overlapping/adjacent slices is
their union: overlap only removes chars, adjacency adds one 2-char
separator but drops one marker of >= 2 chars). There are at most n-1
markers, each <= marker_len because marker_len is formatted with
first=last=len(records) and elided=total_len, the maxima. Since
target = (cap - (n-1)*marker_len) // n, the rendered length is
<= n*target + (n-1)*marker_len <= cap. Gate mutation M5 (loop disabled)
left all 7 sampler tests green; 30 random shapes never entered it.
Delete the `if len(res) > limit` re-trim at the end of
`_bound_oversized_record` (agent/context_compressor.py:3488-3489).
Why it can never fire: `marker_reserve` is the marker width formatted
with `elided=len(record)`, the largest value `elided` can take, so the
real marker is <= marker_reserve; `head` is `record[:head_len].rstrip`
(<= head_len) and `tail` is `record[-tail_len:].lstrip` (<= tail_len);
head_len + tail_len == limit - marker_reserve. Hence
len(head + marker + tail) <= limit always. The branch was untested and
unreachable (gate 2c W2).
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Replace the eight sampler tests with two invariants. The boundary test
now drives _generate_summary so it fails on main for the actual defect
(a retained section starting mid-record) instead of on a helper name,
and folds in the multi-paragraph case since blank lines inside a message
were what broke the string-split approach. Dropped: covers_multiple_
regions and generate_summary_lean_samples_structural_records (green on
main, no teeth), omission_markers_account_separators_exactly (asserted
only that a marker exists), newest_message_ending_in_blank_lines (moot
once records are structural), and preserves_oversized_last_record (a
record above the whole cap is unreachable through _serialize_for_summary
because _CONTENT_MAX bounds message bodies). The remaining oversized
test asserts behaviour — tail kept, oversized record bounded — not the
marker wording.
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
The record split added .rstrip("\n") to every serialized record. That
changes the bytes _serialize_for_summary returns for legacy tail_mode
and for agent/micro_compaction.py's _micro_summarize_one, which never
sample and never needed it: a message with trailing newlines was no
longer reproduced verbatim in the summarizer prompt. Sampling receives
records structurally, so an empty pseudo-record can no longer become the
tail anchor and the strip has no remaining purpose.
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.
- agent/context_compressor.py:
* Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
* Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
* Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
* Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
* Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
* Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
* Test newest message ending in blank lines does not produce empty tail anchor.
* Test exact omission marker separator accounting.
* Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.
(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)
Add test_probe_exit_after_mark_disconnected_is_logged_at_info: connect with
_live_bot_factory(), wait for one healthy sample, flip _running through the
production setter _mark_disconnected(), wait for the probe task to finish and
assert exactly one INFO record containing "probe exiting (running=False".
WHY: 95d574949a collapsed the WARNING branch and deleted the test that poked
_client = None, so nothing pinned the exit log at
plugins/platforms/discord/adapter.py:1761-1764 any more (regate/A2.md W1:
mutation M5 "delete the logger.info call" survived). With this test the
mutation fails again (0 == 1), and the path is reached the way production
reaches it instead of via adapter internals.
Add a module-level `_live_bot_factory()` next to `_connect` and use it in
test_disconnect_cancels_liveness_task,
test_closed_transport_first_strike_forces_reconnect and
test_recovery_after_unhealthy_streak_is_logged, which each carried a
verbatim copy of the same 4-line `factory` closure
(tests/gateway/test_discord_liveness.py:346-349, :414-417, :455-458).
WHY: the branch took the file from one copy to four; the next liveness
test would have copied it again. Behaviour-equivalent (same kwargs, same
`fetch_user` stub). The `bots` sink the reviewers sketched is omitted —
its only consumer was the WARNING-exit test removed in the previous
commit, so no caller needs it.
Declare `_TERMINAL_HEALTH_REASONS = frozenset({"socket_closed",
"client_closed"})` beside `_read_websocket_health` (the producer of those
literals at adapter.py:1705/:1707/:1715) and use it at the escalation
site in `_liveness_loop` (was `reason in ("socket_closed",
"client_closed")` at :1791).
WHY: the first-strike classification was a magic-string match against
literals produced 80 lines away with no shared definition. A rename in
the producer, or a new hard-closed reason, would silently demote the
escalation back to the threshold path with nothing failing. One named
set next to the producer makes the coupling visible. The
`tuple[bool, str]` return shape is unchanged (existing tests monkeypatch
the sampler with 2-tuple lambdas).
Proof: mutating the set to {"socket_closed"} fails
test_closed_transport_first_strike_forces_reconnect[client_closed];
restored, the liveness file is green.
Collapse the INFO/WARNING level switch on the probe-exit log
(plugins/platforms/discord/adapter.py:1755-1766) to one unconditional
logger.info and delete `expected_exit`. Drop
test_unexpected_probe_exit_logs_warning, which was the only caller that
could reach the WARNING half.
WHY: the WARNING branch is unreachable in production. It needs
`_running=True and not _disconnecting and _client is None` at the guard,
and `_client` has exactly two production setters to None:
- adapter.py:1269 inside connect(): synchronously followed by
`self._client = commands.Bot(...)` at :1271 with no await between, so
the probe coroutine can never observe the None.
- adapter.py:1941 inside disconnect(): runs after `_disconnecting = True`
(:1915) and after `await self._cancel_liveness_task()` (:1917), so the
probe is already cancelled (exits via CancelledError at :1752) and the
flag would read as an expected exit anyway.
No other `_client = None` in plugins/platforms/discord/, gateway/platforms/
base.py or gateway/run.py. `_running=False` and `_disconnecting=True` come
only from teardown paths, all "expected". The deleted test pinned this by
poking `adapter._client = None` directly — a state the code never produces.
This drops the level switch borrowed from #118504 but keeps its intent:
the probe exit is always logged with its full state (#118487, the probe
must never disappear silently).
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
Commit b1530538ab raised the probe-exit log to WARNING when the adapter is
still running but has lost its client (plugins/platforms/discord/adapter.py:1761),
yet no test asserted the level, so a revert to a flat INFO would stay green.
Add one test: connect, null out `_client` while `_running` is True, wait for
the probe task to finish, and require exactly one "probe exiting" record at
WARNING containing `running=True, client=False`.
Red proof: switching the unexpected branch to logging.INFO fails with
`assert 'INFO' == 'WARNING'`; green at head (17 passed).
Commit 338553e7bb made `client_closed` terminal alongside `socket_closed`
(plugins/platforms/discord/adapter.py:1795) but only `socket_closed` had a
regression test, so dropping `client_closed` from the tuple would stay green.
Parametrize the test over both reasons (renamed to
test_closed_transport_first_strike_forces_reconnect) and loosen the exact
ERROR-text match to "forcing reconnect" plus the reason token, so a wording
tweak of the log line does not fail the behavioural test.
Red proof: with `"client_closed"` removed from the terminal tuple the
`[client_closed]` param fails (calls == 2, handler never called on strike 1);
green at head (16 passed).
Since the first-strike escalation, a closed transport (socket_closed /
client_closed) forces the reconnect on the first unhealthy sample and the
failure threshold only applies to soft signals (ack staleness, latency,
event silence). The user guide (website/docs/user-guide/messaging/discord.md:89)
and the env-var reference (website/docs/reference/environment-variables.md:780)
still described the threshold as gating every unhealthy sample, so an
operator reading `1/2` followed by a forced reconnect would think the
knob was ignored. One sentence each, citing #118487.
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
test_soft_unhealthy_signal_still_requires_threshold_strikes passes on
main unchanged: the threshold contract for soft reasons is already
covered by the pre-existing liveness tests, so it cannot catch a
regression of #118487 and only adds runtime. The two kept tests are
red on base for the right reasons (strike-1 escalation; reset logged).
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.
Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.
(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
The compound refusal check `_response_refusal_text(response) or
_is_summary_refusal(content)` was spelled out verbatim at both call sites
(agent/context_compressor.py:3639 in `_call_summary_llm` and
agent/micro_compaction.py:157 in `_micro_summarize_one`, the latter through
two separate `_cc()` facade lookups). `_classify_summary_failure` already
describes it as one concept ("refusal content (prose or provider `refusal`
field)"), and the ordering invariant — an explicit provider `refusal` field
wins over summary-shaped content — lived in two callers instead of one place.
Add `_is_refusal_response(response, content)` next to the detectors and call
it from both sites (micro: one `_cc()` lookup). `_is_summary_refusal` and
`_response_refusal_text` stay as-is for the unit tests. Behaviour unchanged.
Proof: guard + micro test files green (62 passed); mutating the new predicate
to `return False` fails exactly the batch refusal-body test, the provider
refusal-field test and the micro refusal test (3 failed, 59 passed).
`_response_refusal_text` (agent/context_compressor.py:199-201) returned ""
early when `_coerce_llm_message` yielded None or a str. That branch is dead:
`_message_field` (agent/auxiliary_client.py:8014) is
`msg.get(name) if isinstance(msg, dict) else getattr(msg, name, None)`, and
`getattr(None, "refusal", None)` / `getattr("x", "refusal", None)` are both
None, which the `isinstance(refusal, str)` tail already maps to "". So the
guard could never change the result — invariant: a non-dict, non-object
message has no `refusal` attribute, so the field lookup is None.
Probe (None, "x", bare dict, dict-with-choices, choices=None, dict refusal,
str message, refusal=None) returns ['', '', 'no', 'nope', '', 'r', '', '']
both before and after this change; the guard test file stays green.
Repo convention: no defensive guards around code that cannot fail.
agent/context_compressor.py:_classify_summary_failure (L697 marker list)
maps the "refusal content" RuntimeError onto empty_content without saying
so in the docstring, which makes the "LLM returned empty content" fallback
log line look wrong when the real cause was a refusal. Document the
intentional reuse (cooldown + main-model fallback + abort) in one place.
Gate finding: B.2ab (misleading fallback log for refusals).
tests/agent/test_compressor_truncated_summary_guard.py:197 asserted
`_previous_summary is None or refusal not in (_previous_summary or "")`. The
disjunction would still pass if the compressor committed ANY checkpoint
(e.g. a rewritten/derived one) after a refusal; the invariant this test
exists to protect is that a rejected refusal leaves no checkpoint at all,
exactly as the sibling provider-refusal-field test already asserts (L218).
Red with the refusal raise in _call_summary_llm defeated: compress() commits
the refusal as `_previous_summary` (is None -> False); green at head
(25 passed).
Gate finding: B.2c (weak disjunctive assertion).
agent/context_compressor.py:_is_summary_refusal (L177-184) matched on the
opener only, so a genuine templated summary that starts with a hedge such as
"I cannot see the earlier turns, but here is the summary:\n## Goal ..."
was classified as a refusal and compression aborted (gate probe: the
[preamble] compress() run aborted with "refusal content"). A refusal-only body
never carries the summary template's "## " section headings, while a real
summary always does, so exempt any body containing a "## " heading after the
prefix match.
Test: TestSummaryRefusalGuard::test_narrowing_and_heading_exemption. Red when
the heading exemption is removed (L91 assert True is False) and red when the
summary-term narrowing is dropped via an early `return True` (L85 assert True
is False, pins gate mutation M3); green at head (25 passed). Live probe:
[preamble] now compresses 25->7 with prev_summary_set=True.
Gate finding: B.2ab (false positive on prose-preambled summary).
agent/context_compressor.py:_SUMMARY_REFUSAL_PREFIX_RE (L170) only matched
"cannot"/"can't", so the two-word "I can not provide a summary" slipped
through, and _is_summary_refusal (L184) looked for the exact terms
summary/summarize/checkpoint, so British "I can't summarise this" was
accepted as a checkpoint. Allow optional whitespace in can(?:\s*not|['’]t)
and match the stems ("summar", "checkpoint") instead.
Test: two new params on TestSummaryRefusalGuard::test_rejects_refusal_body.
Red with the previous prod file (2 failed: assert False is True), green at
head (24 passed).
Gate finding: B.2c (detector misses "can not" / "summarise").
agent/context_compressor.py:_response_refusal_text (L187-201) re-implemented
the choices[0].message traversal that agent/auxiliary_client.py already
exposes as _coerce_llm_message (L7997) + _message_field (L8014) and that the
sibling extract_content_or_reasoning uses on the very same response. Use the
shared helpers so the two extractors cannot drift on response shape
(dict-shaped proxies vs ChatCompletion objects vs bare messages).
Behaviour kept: None / SDK object / dict-with-choices / missing attr /
refusal=None -> "" ; the dict-unwrap (message/reason/text) + isinstance(str)
tail is unchanged. Gate live probe ($D/gate/B_live_probe.py) output is
byte-identical before and after; the guard test file is green.
Gate finding: B.2c (duplicate response traversal).
OpenAI-style structured-output refusals arrive in
`choices[0].message.refusal` with `content` empty or filler, so the
prose detector never sees the refusal and the filler would be committed
as the compaction checkpoint. Read that field (str, or the dict shapes
some proxies emit) in both the batch (`_call_summary_llm`) and micro
(`_micro_summarize_one`) paths and route it through the same
"refusal content" failure so it can never become `_previous_summary`.
Also widen the prose opener with #118406's extra shapes — `we` as
subject, `refuse to`, `could not` / `couldn't`, `apologise`, and a bare
`I'm unable` / `I am not able` opener. #118406's `could(?:not|['’]t)`
missed "couldn't" (it only matched "couldnot"/"could't"); spelled out
here as `could\s*not|couldn['’]t`.
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>