docs(agent): say sampled_chars counts display chars in _sample_summary_records
The coverage docstring (agent/context_compressor.py:3510-3511) said the counters count "record content only", but `sampled_chars` sums the *display* records (post `_bound_oversized_record` truncation) while `input_chars` sums the raw records. Spell that out so telemetry consumers do not compute `omitted` two different ways (gate 2c suggestion, L3512-3513/3524-3525).
This commit is contained in:
@@ -3566,8 +3566,11 @@ Summary generation was unavailable, so this is a best-effort deterministic fallb
|
||||
def _sample_summary_records(cls, records: Sequence[str]) -> Tuple[str, Dict[str, int]]:
|
||||
"""Sample complete serialized records while retaining the character bound.
|
||||
|
||||
Returns the bounded transcript and record-level coverage counters (chars count record
|
||||
content only, not separators or elision markers) for compression telemetry.
|
||||
Returns the bounded transcript and record-level coverage counters for compression
|
||||
telemetry. `input_chars` counts raw serialized record content; `sampled_chars` counts the
|
||||
*display* chars of retained records (after intra-record truncation by
|
||||
`_bound_oversized_record`); neither includes separators or elision markers, so
|
||||
`omitted_chars = input_chars - sampled_chars` also covers truncated-away bytes.
|
||||
"""
|
||||
input_chars = sum(len(r) for r in records)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user