Commit Graph

45181 Commits

Author SHA1 Message Date
Benjamin PERRY
af595ff8a6 fix(files): preserve SQLite locks during previews
Co-Authored-By: Hermes Agent / OpenAI Codex / gpt-6-sol <noreply@agents.invalid>
(cherry picked from commit 60d879c125177cd4cacdbc840f16c2e048b4199c)
2026-09-27 20:45:56 +05:30
kshitijk4poor
fb7eda7416 refactor(persistence): drop a dead digest pop and scope the transcript-write docstring 2026-09-27 20:45:44 +05:30
kshitijk4poor
b1bb9031e3 test(compression): expect the rotation handoff digest on compressed dicts
The rotation handoff now stamps each child row's stored-row digest, so
the exact-dict comparison must ignore DB_ROW_SNAPSHOT. Assert every
compressed dict carries one: this pins the non-flush restamp, which
otherwise only probes covered.
2026-09-27 20:45:44 +05:30
kshitijk4poor
f337631f43 fix(persistence): restore caller row state when any transcript insert rolls back
_insert_message_rows stamps _row_id and the stored-row digest onto the
caller's dicts inside the write transaction. Only append_messages_batch
restored that state on rollback. archive_and_compact, replace_messages
and the rotation handoff left the rolled-back id + digest on the dicts;
SQLite reuses the id, so a later flush found a digest mismatch on the
foreign row, adopted it and silently dropped the user's message.

Move the capture/restore into _execute_transcript_write, used by every
caller that inserts caller-owned dicts: each attempt starts from the
caller's state and a final failure restores it before re-raising.
(Rewind replacement and import insert dicts built inside the txn.)

Also: bind _message_row_params directly on insert instead of the
serialized-dict round-trip, import the public DB_ROW_SNAPSHOT /
CANONICAL_ROW names, set adopt=False once, and reuse target_row
instead of re-SELECTing when nothing was written.
2026-09-27 20:45:44 +05:30
kshitijk4poor
1a95a75b1a fix(persistence): adopt content only on legacy rows and stamp every inserted row
A legacy (no-digest) dict over a non-blank assistant row adopted the whole
decoded DB row: tool_calls / reasoning* / codex_* were overwritten with the
stored JSON (which still holds the escaped lone surrogate the sanitizer just
fixed, re-injecting it into the provider payload) and live-only fields were
popped. Resumed dicts (_rows_to_conversation stamps _row_id without a
digest) and compaction clones hit this path. Adopt content only, as before
this stack, via a content-only canonical handled like the metadata-only one.

_insert_message_rows dropped a clone's parent digest but only the flush
path restamped it, so clones made by archive_and_compact / replace /
rotation handoff / import reached the legacy path and the first live edit
after a clone was not persisted. Stamp the stored-row digest inside
_insert_message_rows (one batched SELECT, cold paths only; the flush path
statement count is unchanged) and drop the duplicate call in
append_messages_batch.

Define the _db_row_snapshot / _canonical_row keys once in
agent/message_metadata.py and import them everywhere instead of repeating
the literals.
2026-09-27 20:45:44 +05:30
kshitijk4poor
6e8aa00626 fix(persistence): version only owned columns so metadata writes keep row ownership
The row digest hashed every repair column, so a same-process metadata write
(reaction, display-kind stamp, api_content / codex reasoning backfill,
platform message id) made our own row look like a foreign winner. The
re-flush then adopted the stale DB row: a later live edit (the non-ASCII
strip recovery) was reverted, and unsanitized tool_calls/reasoning were
copied back onto the live dict.

The digest now covers only the owned (non-metadata) columns: it means "the
row is still what we last committed". Match -> write the live owned values
and hand over only presentation metadata the live dict lacks; mismatch ->
genuine other writer, adopt as before. The r3 "stored content equals the
durable form of live" special case is subsumed and removed.

Also:
- _insert_message_rows drops a carried digest when it assigns a new row id
  (compaction/replace/import clones carried the parent's version).
- append_messages_batch restores each message's _row_id / digest /
  timestamp and pops the adopted row at the top of every _execute_write
  attempt, so a rolled-back attempt cannot resolve to a foreign row.
- message_id is no longer synced onto live (int -> str flip, spurious
  platform_message_id).
- The JSONL divert strips both bookkeeping keys via one frozenset.
2026-09-27 20:45:44 +05:30
kshitijk4poor
43b1dd7c5f test(persistence): pin metadata-only adoption and stored-row digests
Extend the kept active-row test: a reaction between flushes must not replace
the live multimodal user content with its text projection (red on the previous
tip at the image assertion), and the reaction metadata is synced. The user row
carries an int message_id so the stored-row (TEXT affinity) digest path is
exercised.
2026-09-27 20:45:44 +05:30
kshitijk4poor
ca7e4c1636 fix(persistence): keep live content on metadata-only row changes
Same-process writers (set_message_reaction, display-kind stamping, api_content
backfills, codex reasoning update) change stored columns after a flush without
refreshing the live row digest. The next sanitize + re-flush treated that as a
concurrent winner and copied the lossy durable projection over live multimodal
content, dropping image parts and shifting the prompt-cache prefix. On adoption
we now keep live content when the stored content is just the durable form of
it and sync metadata only; a real concurrent content winner is still adopted.

The legacy blank-assistant path (dict with _row_id but no digest, blank DB row)
again only fills the row from live content, as on main, instead of running the
full canonical sync that wiped live reasoning_content/finish_reason/tool_calls.
Adoption on the legacy path is limited to a non-blank row.

Also: transcript_row_snapshot returns str (the partial-row branch had no
caller), serialization only runs on the digest-match branch that reads it,
stamping reuses hermes_state_common._id_chunks, and _MESSAGE_WRITE_COLUMNS is a
plain top-level import (hermes_state_messages imports this module lazily, so
there is no cycle).
2026-09-27 20:45:44 +05:30
kshitijk4poor
1a43a4ef48 fix(persistence): strip the adopted canonical row from provider payloads
Gateway/TUI/CLI callers pass their live dicts straight to
append_messages_batch, so a concurrent-winner adoption leaves the
decoded durable row (_canonical_row) on a dict that may later be sent
to the model. Treat it as persistence-only like _row_id and the digest
so the outbound builder and token estimator both drop it.
2026-09-27 20:45:44 +05:30
kshitijk4poor
4f6ab19304 fix(persistence): keep live multimodal content and hash stored rows
The row-addressed repair stamped the decoded durable row on every
resolved message, so after our own rewrite the sync copied the lossy
durable projection (image parts -> "text\n[screenshot]") back onto the
live dict: multimodal user/tool messages lost their images and the
prompt-cache prefix changed. Adopt the DB row only when another writer
won (digest mismatch) or on the legacy assistant path, as BASE did.

The insert-time digest hashed Python bind values, but SQLite affinity
rewrites them on storage (int message_id -> TEXT, float token_count ->
INTEGER), so live and DB digests never matched and in-place edits were
silently dropped. Hash the stored rows instead, only on the
append_messages_batch flush path that reads the digest (one SELECT per
batch), incrementally (type tag + length prefix) instead of via JSON.
Also skip the no-op UPDATE, fix the _write_columns comment/spacing and
drop the duplicate top-level Optional import (F811).
2026-09-27 20:45:44 +05:30
kshitijk4poor
4854225903 fix(persistence): version transcript rows by digest, not a row copy
The CAS row snapshot was a full copy of each message's durable payload
riding on the live dict. The rough token estimator priced it (about 2x
estimates -> premature compaction) and it doubled transcript memory.

Replace it with a 16-byte blake2b digest of the repair columns. The
compare now runs in Python against the target row already read inside
the BEGIN IMMEDIATE transaction, followed by a plain UPDATE. Also:
- add _db_row_snapshot to PERSISTENCE_ONLY_MESSAGE_FIELDS so the
  estimator and the outbound request builder both drop it
- derive _REPAIR_COLUMNS/_SYNC_FIELDS from _MESSAGE_WRITE_COLUMNS
- use hermes_state_common._placeholders
- drop the dead resume-path stamp (the SELECT has no token_count, so it
  was always None) and the dead tool name assignment in
  _decoded_repair_row
- keep the digest out of divert JSONL

The kept active-row test now pins estimate stability across a flush and
the survival of a concurrent writer's row. It goes red on the old
prod files and red when the digest compare is removed.
2026-09-27 20:45:44 +05:30
kshitijk4poor
18ee4578b7 test(persistence): trim sanitized-row dedupe tests to two invariants
Keep one test per invariant: active user/tool rows whose _db_persisted marker
was popped by the outbound sanitizer keep their _row_id and are not re-inserted
(the #123462 desktop/serve path), and archived user/tool rows are repaired in
place instead of appended. Both fail on b4410b4bad; the other four PR tests
covered edge branches and are dropped per the <=2 invariant-test budget.
2026-09-27 20:45:44 +05:30
kshitijk4poor
53e49c49dd fix(persistence): import Optional where transcript_repair uses it
transcript_row_snapshot annotates Optional, which was only reachable via the
PLUGIN-COMPAT re-export block at the bottom of the module. Internal code must
not depend on that revert-scheduled block, so import it with the other typing names.
2026-09-27 20:45:44 +05:30
Nagisa-3000
02b8c4a655 fix(persistence): preserve transcript row identity safely
(cherry picked from commit a1f28d54482eab32bd119fbc856c6e910c595ba9)
2026-09-27 20:45:44 +05:30
Nagisa-3000
e4e7837623 fix(persistence): repair transcript rows for all roles
(cherry picked from commit 7cebc297b145c0e3db007c1f05ea348a626fd38d)
2026-09-27 20:45:44 +05:30
Nagisa-3000
7581677da8 fix(persistence): repair inactive transcript rows in place
(cherry picked from commit 1d33010a27fbd7f053e1037269b6059edaef77ef)
2026-09-27 20:45:44 +05:30
kshitijk4poor
f1826946ca chore: map Nagisa-3000 for salvage of #123849 2026-09-27 20:45:44 +05:30
kshitijk4poor
f39f76508e fix(gateway): tighten busy re-queue back-off identity check and test matrix
Review cleanups on the drain back-off:
- `_drain_after` now requires `guard`: a None default silently meant "release
  whatever guard is current", which is the guard-swap bug the parameter fixes.
- The identity rule is stated once (docstring, wrapped), and the redundant
  `pending_event is dispatched_event` disjunct is dropped: the same object
  always has an equal message_id, and when that id is empty its timestamp
  equals itself, so the remaining comparison already covers it.
- Tests drop the `_Adapter` alias and cut the hot-loop matrix from 6 to 4
  explicit cases (plain, rewrite with id, rewrite without id, steer). The
  demotion route does not interact with the identity axis, and each case
  spends a fixed 1s measuring window.
2026-09-27 20:44:38 +05:30
kshitijk4poor
eb32f07d0b test(gateway): deflake busy re-queue drain test and pin follow-up identity
Test 2 waited a fixed 0.6s before cancelling, which under load landed mid
first-handler and raised KeyError; poll for a real back-off instead. Route
chained follow-ups through the runner mid-turn and assert they dispatch
immediately, which fails (0.25s/0.5s gaps) with the identity check disabled.
Add an id-less rewrite-copy case to the hot-loop parametrization and drop the
typing assertion that could never fire in this harness.
2026-09-27 20:44:38 +05:30
kshitijk4poor
c0f8ac6645 fix(gateway): keep swapped command guards and id-less rewrites bounded in drain back-off
The back-off drain's slot-empty exit released whatever guard was current, so a
/stop, /new or /reset guard swapped in during the sleep was deleted, defeating
the #48300 guard-swap protection. Capture the guard owned by the drain at spawn
and release only that; also flush the text debounce buffer before popping the
slot, like every other task exit, so a debounced text isn't orphaned.

A pre_gateway_dispatch rewrite copy of an event with no message_id was never
matched to the dispatched event, so it still hot-looped (#123229). Fall back to
the copied timestamp when message_id is empty (a genuine new message gets a
fresh one); a plain message_id==/timestamp== form breaks the runner's re-queue
of a new object with the same id.

Cap the back-off at 1s: nothing wakes the sleep, so a genuine message merged
into the slot meanwhile waited up to 5s; 1 dispatch/s is still ~250x below the
unbounded loop and needs no new wake plumbing.
2026-09-27 20:44:38 +05:30
kshitijk4poor
1288744c79 test(gateway): harden busy re-queue hot-loop regression
Measure the 1s window after the first dispatch so a cold first handler call under load
cannot false-green the no-fix mutation; parametrize over interrupt demotion and steer
fallback; cover cancel during back-off keeping the event queued and the /stop tail
replaying it. Reuse RestartTestAdapter instead of a duplicate adapter.
2026-09-27 20:44:38 +05:30
kshitijk4poor
b6e6f00790 fix(gateway): back off only when the dispatched event bounces straight back
The _busy_requeued tag was reset only by an untagged drain, so chained queue/steer
follow-ups reaching the runner fast path (split-brain adapter, multiplex routes) backed off
exponentially again (0.05->2s gaps vs ~0.01 on main). Key the back-off on identity instead:
the drained event must be the one this task just dispatched (same object or same message_id
for rewrite-hook copies). That covers every demotion site (interrupt demotion, queue, steer
fallback, Telegram grace queue, agent-starting merge) with no per-site tagging, so the tag
and _hm_tag_busy_requeue are dropped.

A backed-off event now stays in _pending_messages during the sleep and is popped after it,
so a cancel needs no put-back and can no longer drop the older message when a newer one
took the slot. jittered_backoff import moved to module top (agent.retry_utils is stdlib-only).
2026-09-27 20:44:38 +05:30
kshitijk4poor
fc8539b7b7 fix(gateway): back off only runner-demoted events, not every None turn
The drain back-off keyed on `response is None`, but None is the normal
return for every streamed turn and for queue/steer busy modes, so
ordinary chained follow-ups were delayed 0.25s -> 5s forever. The busy
fast-path in run_inbound now tags the adapter's pending head
(`_busy_requeued`) where it demotes the same inbound event (or its
rewrite-hook copy) back into the queue (interrupt demotion, queue mode,
steer fallback); the drain backs off only for a tagged event and
otherwise resets the counter and dispatches at once — single reset owner.

Also: clear _requeue_counts on cancel_session_processing, stale-lock
heal, session end and shutdown; restore the pending event if cancelled
during the back-off sleep; reuse agent.retry_utils.jittered_backoff.
The regression test now chains 3 genuine follow-ups after a streamed
(None) turn and requires each to dispatch in <0.1s (red on pre-fold).
2026-09-27 20:44:38 +05:30
kshitijk4poor
506cc14def fix(gateway): back off busy re-queued events inside the drain task
The runner busy-demotion puts the event back into the adapter pending slot and
returns None; the in-band drain re-dispatched it at once, ~400 times/s for the
whole busy window (typing churn, log flood, platform connection storm).

Count consecutive unanswered re-queues per session on the adapter (not by event
identity, so a pre_gateway_dispatch rewrite via dataclasses.replace is still
caught), reset when a handler returns a response. First re-queue stays
immediate (restart auto-resume self-bounce); later ones back off 0.25s..5s.
The delay is slept inside the new drain task before its processing
try/finally, so cancelling during the back-off cannot reach
_finish_session_task late-arrival respawn (no concurrent handler, no
untracked task on shutdown).

Fixes #123229

Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
2026-09-27 20:44:38 +05:30
ahisblessed
a79d26da0d test(gateway): regression for busy re-queue hot loop in adapter drain
A busy-demoted event that the runner puts back into the adapter pending slot
is re-dispatched by the in-band drain with zero delay (~400/s). Salvaged from
PR #123259 (test only; the fix is rebuilt separately).

Refs #123229
(cherry picked from commit f83f4c958514c0be6e3f37af9d6ac5afc62fdf16, test file only)
2026-09-27 20:44:38 +05:30
hermes-seaeye[bot]
8c9fe96400 fmt(js): npm run fix on merge (#125322)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-27 14:20:28 +00:00
hermes-seaeye[bot]
e54d550521 fmt(js): npm run fix on merge (#125310)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-27 14:06:23 +00:00
flyer103
e5050a040a fix(desktop): don't announce a reconnect when the app is quitting
Quitting the Desktop app wrote "[boot] Restarting desktop connection" to
desktop.log and pushed the same message to the renderer over
hermes:boot-progress. Nothing was restarting: the quit coordinator called
teardownPrimaryBackendAndWait() with the default soft=false, and soft=false
is what makes resetHermesConnectionState() call
resetBootProgressForReconnect().

Name the two teardown intents in quit-teardown.ts and use them at every
deliberate primary teardown: a teardown that re-homes ('reconnect', update
hand-off and bundle swap) keeps the announcement; one that brings nothing
back ('quit', the quit coordinator and the uninstall teardown) stays silent.
Fixes #123437.
2026-09-27 08:53:09 -05:00
Brooklyn Nicholson
bbf670a98d style(kanban): satisfy the curly and blank-line lint rules in the alias-resolution change
Mechanical eslint --fix over the two files this branch touched; no behavior
change. 35 kanban notify tests still pass.
2026-09-27 08:52:49 -05:00
funky-xamarin
c4700d60de fix(desktop): preserve alias task cache invalidation 2026-09-27 08:52:49 -05:00
funky-xamarin
23d95743a0 fix(desktop): resolve default Kanban board before event subscription 2026-09-27 08:52:49 -05:00
Kevin Rajan
97065356c8 fix(desktop): keep model rows clickable before pointer movement 2026-09-27 08:51:26 -05:00
kshitijk4poor
2f14d5e6e4 docs(matrix): fix the parked-voice cap limit count 2026-09-27 19:06:05 +05:30
kshitijk4poor
67066dba29 docs(matrix): note the parked-voice cap limit
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
kshitijk4poor
35a24efaf3 fix(matrix): claim the newest voice sent before the bare mention, not after it
The r3 fold made park refuse an older voice once a newer one from the same
sender had parked. With one /sync batch holding [v1 (slow gate), m1 bare
mention, v2 (fast)], v2 parked first, v1 was then refused, and m1 claimed
v2 -- a voice sent after the mention. v1 was lost and v2's own mention was
answered as empty text.

The bare mention now takes an arrival limit (ParkedVoices.mark) before it
settles, and claim pops the newest parked voice that began before that
limit, dropping older ones. park no longer refuses by arrival; parked
voices are kept per sender ordered by seq (bounded to 4). A claim made
while gates are still in flight records a floor, so a late older voice
(seq <= claimed) cannot park and outlive the claim; the floor clears when
the sender's in-flight list empties. One answer per bare mention, newest
before the mention wins, and the r3 orphan case still leaves nothing
parked.

Also: pending() ignores entries past CLAIM_WINDOW_SECONDS so an expired
voice no longer sends every later text through the mention regexes; the
m.thread root lookup is one _thread_root helper used by both the park
decision and its pre-check; voice_gate is annotated.

The kept test gains a [voice slow, mention, voice2 fast] + mention2 row:
red on a932dc031c (['$voice2', '$text2']) and on f1ea71de7c
(['$voice', '$text2']).

Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
kshitijk4poor
df2f8b8afc fix(matrix): track every in-flight voice per sender and release the mark at the park decision
The same-sync-batch in-flight mark was one slot per (room, sender), so a
second voice's begin overwrote the first. When the older voice finished
gating after the newer one, the bare mention settled on (and claimed) the
newer voice only, and the older one parked afterwards as an orphan: the
sender's next unrelated bare mention within 120s re-dispatched it, so a
stale voice got downloaded, transcribed and answered.

The in-flight mark is now a list of gates per key. release removes only
its own gate (idempotent) and drops the key once it is empty; settle
waits on all current gates under the same 5s cap. Each gate carries an
arrival sequence and park refuses an older voice once a newer voice from
the same sender has parked, so a late older voice can neither replace
nor outlive the claim of the newer one (still one parked voice per
sender, newest wins).

The mark was also held across the whole _resolve_message_context,
including the display-name fetch and thread mark, and taken for voices
that can never park (the voice itself @mentions the bot, free rooms, bot
threads, non-allowlisted rooms). A bare mention racing such a voice
waited for the voice's display-name fetch (0.61s vs 0.31s on main with a
0.3s fetch). begin now runs only when a synchronous pre-check says the
voice can park, and the mark is released as soon as the park decision
is made (before the display-name fetch; DM voices release there too),
with the finally still covering exceptions and cancellation.

Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
kshitijk4poor
6016f30427 fix(matrix): let a same-sync-batch bare mention wait for its voice to park
mautrix dispatches every event of one /sync batch as its own task
(wait_sync=True). The voice only parks after awaiting the room identity,
which is a homeserver round-trip whenever the 60s identity cache is stale,
i.e. in any idle room. The bare-mention text checked the park without
awaiting, found nothing, dispatched an empty text, and the voice then
parked and expired unclaimed.

A parkable voice is now marked in-flight per (room, sender) right before
its first await and released in a finally once it is gated. A bare
mention that sees an in-flight voice waits for it (bounded, 5s) before
claiming; ordinary text only pays two dict lookups and never waits.

After a claim the bare-mention event also gets its read receipt, so the
read marker is not left one event short of the pre-claim behaviour.

Cleanups: ParkedVoices drops its unused window param, stores
(parked_at, voice) instead of a flat tuple, the voice_mention import
moves to the top import block, and the redundant require_mention check
on the claim path is dropped (parking/in-flight only happen under it).

Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
kshitijk4poor
ec504c1942 fix(matrix): pass the voice-mention claim explicitly instead of a claimed-id set
Review cleanups on the parked-voice fix:

- The claim bypass is now an explicit mention_claimed parameter threaded
  _handle_media_message -> _resolve_message_context. The old _claimed set was
  only drained inside the require_mention branch, so a voice whose thread became
  a bot thread while parked leaked its id for the adapter's lifetime.
- One MSC3245 predicate (has_voice_marker, `.get(...) is not None`) shared by
  is_voice_event and _classify_inbound_media; the two used to disagree on a
  null marker (parked as voice, classified as AUDIO).
- One mention helper (_content_mentions_bot) for both the gate and the
  bare-mention claim, instead of two copies of the m.mentions extraction.
- ParkedVoices.has() dict lookup gates the claim, so ordinary text messages
  no longer pay _strip_mention/regex work when nothing is parked.
- Test: a same-room mention WITH text stays a text (gives the bare-only guard
  teeth) and the parked voice is asserted never downloaded.

Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
kshitijk4poor
cb7bdd2177 fix(matrix): answer a voice message whose @mention was typed while recording
Element X ships an MSC3245 voice event with an empty m.mentions block and
sends the mention the user typed while recording as a separate m.text
event right after. With require_mention on, the gate dropped the voice and
the follow-up bare @bot was dispatched as an empty text, so the bot never
answered the voice.

Park an unmentioned voice keyed by (room_id, sender) instead of forgetting
it; a bare mention from the same sender in the SAME room within 120s claims
it and the voice is processed in place of the empty text. The claimed event
id passes the mention gate exactly once, so it is not re-parked; expired
entries are pruned on every access. Parked voices are never downloaded or
transcribed (no wake-word/STT of unmentioned audio), and a mention in
another room never pulls a voice across rooms.

The state lives in a small sibling module (voice_mention.py) because
adapter.py is already far past the file-size budget.

Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
2026-09-27 19:06:05 +05:30
Brooklyn Nicholson
8bcca88282 test(desktop): match the desktop.mjs compile step with native path separators
The --icons case found the compile command with endsWith('/scripts/build/desktop.mjs'),
which never matches the path.join-built argument on Windows, so the case failed with
TypeError instead of asserting. Match both separator styles.

Fixes https://github.com/NousResearch/hermes-agent/issues/125139
2026-09-27 08:34:17 -05:00
kshitijk4poor
062dc1e7f0 refactor(context): drop redundant or-empty before redaction 2026-09-27 18:38:58 +05:30
kshitijk4poor
ab84dc8d57 fix(context): drop stale _shrink reference in marker comment
_shrink was deleted along with _truncate_tool_call_args_json, so the
comment pointed at code that no longer exists.
2026-09-27 18:38:58 +05:30
kshitijk4poor
fb86bc708d fix(context): redact full tool-call args before the summarizer cut
62ceddd342 cut raw args to HEAD+4096 before redaction to save time. The
PEM redaction pattern only matches a complete BEGIN...END block. A long key
whose END fell past the cut stayed unredacted, and once an earlier key was
redacted and the text shrank, its body landed in the 1200-char head that
goes into the persisted summary. Go back to the BASE order: redact the full
args, then apply the MAX/HEAD cut. This is a cold path (once per summarized
call per compaction), and _SUMMARY_INPUT_MAX_CHARS still bounds the prompt.
Extend the kept canonical-args test with a two-PEM input that leaks on
62ceddd342 and passes now.
2026-09-27 18:38:58 +05:30
kshitijk4poor
be834681dc fix(context): measure pruned regions, bound arg redaction, drop dead counters
Follow-ups to making tool-call args byte-exact:
- _record_compression_regions measured canonical_messages slices while
  compress_start/compress_end are indices into the pruned copy that head/tail
  are assembled from; measure the pruned rows actually sent, as before. This
  also removes the only canonical slicing, so blank-echo classification drift
  between the two copies can no longer misalign anything.
- _render_tool_call_for_summary redacted the full (now unbounded) args before
  cutting to 1200 chars; cut to head+4096 first. Output unchanged for args
  within that window.
- pressure_hits always equalled demoted once arg truncation left; fold it.
- Drop the fixture-only tautological assert in the guardrail test helper.
- Reword stale compress()/compression_marker docstrings that still described
  canonical head/tail and compressor-written arg markers.
2026-09-27 18:38:58 +05:30
kshitijk4poor
66240ddc59 fix(context): restore marker import in tests, drop unused compressor imports
The arg-truncation removal deleted the only uses of the marker constants in
agent/context_compressor.py (ruff F401), and the test module had been
importing _COMPRESSION_MARKER_PREFIX through it, so test_context_compressor
line 137 raised NameError. Import it from its home, agent.compression_marker.
2026-09-27 18:38:58 +05:30
kshitijk4poor
697f07a56f test(context): pin byte-exact old tool-call args with one red-on-base test
The PR's two tests passed on the base code (the prune boundary never
reached their calls). Replace them with one invariant test that is red on
base: six old 3000-char write_file arguments must survive the prune
unchanged. Drop the PR's canonical-state test file (tail-canonical shape
contradicts #61932's pressure demotion).
2026-09-27 18:38:58 +05:30
kshitijk4poor
f6ce8bb23b fix(context): assemble compaction head/tail from the pruned copy (#61932)
The salvaged lossless-history change rebuilt the carried head/tail from
canonical history, which undid _pressure_demote_tail's tool-result
shrinking and re-broke #61932 (an all-oversized tail could no longer
compress). Pruning no longer rewrites tool_calls, so the pruned copy's
arguments are already byte-identical to canonical history: assemble the
head and tail from the pruned copy, keeping tool-result demotions and
exact tool-call arguments at once. Docs updated to match.
2026-09-27 18:38:58 +05:30
JoaoMarcos44
a7baa5f5eb fix(context): keep compaction history lossless
(cherry picked from commit d51c8f4f5096badfd0beddd78646617643f6028f)
2026-09-27 18:38:58 +05:30
kshitijk4poor
96cce6843d refactor(auth): drop impossible auth_store dict guard 2026-09-27 18:37:01 +05:30
kshitijk4poor
c8043a3630 fix(auth): drop redundant ownership auth.json reads in nous seeding and load_pool
_seed_nous_singleton re-read auth.json via _profile_owns_pool_provider even
though its only caller (_seed_from_singletons) just loaded the active store
and passed it in; check the passed auth_store through a shared
_store_owns_pool_provider predicate instead (same non-empty-list semantics,
also used by _profile_owns_pool_provider).

In load_pool, borrowing_root_grant repeated the guard that sets
owns_provider (non-None exactly when that guard holds), so test
`owns_provider is False` directly. The tail ownership re-read ran even
with no disk rows, where it could only assign set() -- the constructor
default -- so gate it on disk_ids and re-read only when _persist() ran.
Fix the stale "Computed once" comment.
2026-09-27 18:37:01 +05:30