docs(agent): say sampled_chars counts display chars in _sample_summary_records

The coverage docstring (agent/context_compressor.py:3510-3511) said the
counters count "record content only", but `sampled_chars` sums the
*display* records (post `_bound_oversized_record` truncation) while
`input_chars` sums the raw records. Spell that out so telemetry
consumers do not compute `omitted` two different ways (gate 2c
suggestion, L3512-3513/3524-3525).
This commit is contained in:
kshitijk4poor
2026-09-22 12:28:37 +05:30
committed by kshitij
parent eb05bde6df
commit c94fe7a259

View File

@@ -3566,8 +3566,11 @@ Summary generation was unavailable, so this is a best-effort deterministic fallb
def _sample_summary_records(cls, records: Sequence[str]) -> Tuple[str, Dict[str, int]]:
"""Sample complete serialized records while retaining the character bound.
Returns the bounded transcript and record-level coverage counters (chars count record
content only, not separators or elision markers) for compression telemetry.
Returns the bounded transcript and record-level coverage counters for compression
telemetry. `input_chars` counts raw serialized record content; `sampled_chars` counts the
*display* chars of retained records (after intra-record truncation by
`_bound_oversized_record`); neither includes separators or elision markers, so
`omitted_chars = input_chars - sampled_chars` also covers truncated-away bytes.
"""
input_chars = sum(len(r) for r in records)