Commit Graph

20351 Commits

Author SHA1 Message Date
briandevans
eb4f514b2d fix(tui_gateway): retain live-turn sessions unfinalized when the drain deadline expires
The drain reserves a slice of the shutdown budget so flush_all_sessions
still runs when in-flight turns outlast the window. But that flush was
unconditional: a session whose turn was still running got its one-shot
_finalize_session spent mid-turn, and the executor.shutdown(wait=False,
cancel_futures=True) immediately after does not join the turn. The
session was then permanently un-finalizable and its active-session lease
had been released out from under live work — the same persistence and
lifecycle race the drain exists to close, just relocated past the
deadline instead of removed.

Give _turn_futures a session association (Future -> sid, the same key
space as server._sessions) at both submit sites, and on deadline expiry
exclude the sids whose futures are still running from the flush. Those
sessions are retained unfinalized and therefore recoverable; sessions
with no live turn finalize exactly as before. The done-callback now pops
under the lock, since a bare dict.pop is not the drop-in set.discard was.

wait semantics, the reserve math and the bounded per-tick sleep are
unchanged, so this adds no shutdown latency. All three shutdown callers
(orphan, sigterm, and the tight stdin_closed wait=2.0 path) funnel
through this one function and are covered.
2026-08-03 09:59:58 +05:30
briandevans
e239184521 fix(tui_gateway): bound the shutdown drain sleep by the time left to it
The drain loop slept a flat 0.05s per tick, so it could overshoot its deadline
by up to one tick and spend part of the reserve withheld for
flush_all_sessions(). For a small `wait` the reserve is itself half the budget,
so a single overshoot can consume all of it: at wait=0.34 the drain budget is
0.17s but the loop requested 4 x 0.05 = 0.20s of sleep.

Clamp each tick to the remaining time. The new test asserts on the summed
*requested* sleep rather than wall-clock, which is deterministic: every sleep is
bounded by the strictly-decreasing remainder, so the total can never exceed the
drain budget regardless of how the scheduler interleaves.
2026-08-03 09:59:58 +05:30
briandevans
1fba2dbe85 fix(tui_gateway): drain in-flight turns before finalizing sessions on compute-host shutdown
ComputeHost.shutdown() called flush_all_sessions() before its own in-flight
turn drain loop. server._finalize_session latches on session["_finalized"]
and every later call returns immediately, so that one flush was spent while
turns were still producing output: the unflushed tail was never persisted,
commit_memory_session wrote long-term memory from a truncated transcript, the
session's DB row was marked ended while it was live, on_session_end fired with
completed=False/interrupted=True against a running session, and the
active-session lease was released out from under a turn. The drain loop exists
precisely so that mid-turn work survives a teardown; finalizing first defeated
it. Reachable from all three teardown paths: the parent/orphan guard (which
os._exit(0)s immediately after), the SIGTERM/SIGINT handler, and stdin close.

Drain first, then flush. A slice of the caller's budget (_FLUSH_RESERVE_SECS,
never more than half of it so a short explicit wait still gets a real drain) is
withheld from the drain so the flush still runs when turns outlast the window:
HostSupervisor SIGKILLs the host _SHUTDOWN_TIMEOUT_SECS after SIGTERM — 10.0s,
the same value as shutdown()'s default wait — so a drain allowed to consume the
whole budget would leave the durability write racing that kill. `wait` itself is
unchanged, so total shutdown latency and the SIGTERM->SIGKILL margin are
unchanged.
2026-08-03 09:59:58 +05:30
HexLab98
28d994f26c test(gateway): cover restart after-turn deferral (#77184) 2026-08-03 09:57:58 +05:30
HexLab98
db3f7e4eb9 fix(gateway): defer in-band restart until active turns finish (#77184)
request_restart was calling stop() immediately, so the requesting turn stayed
in the drain wait set and got force-killed at restart_drain_timeout. Wait for
active work to reach zero first, then stop against an idle gateway.
2026-08-03 09:57:58 +05:30
kshitij
f105db2136 Merge pull request #77335 from kshitijk4poor/chore/author-map-atran28
chore: add ATran28 to AUTHOR_MAP
2026-08-03 09:57:46 +05:30
kshitij
307ef0061c chore: add ATran28 to AUTHOR_MAP
at828@proton.me -> ATran28 (GH id=1445620)
Needed for PR #77270 whatsapp bridge reconnect fix.
2026-08-03 09:57:30 +05:30
kshitij
622c220ebb refactor(agent): rebuild hoisted patterns from shared tag-name tuples
Simplify-pass follow-up on the #69653 salvage: the original code built
these patterns from a name loop; hand-expanding them into 10 literals
lost that single source. Tag-name tuples restore it (adding a 6th
reasoning tag is now a one-place change), and the gnarly named-function
pattern regained a pointer to its step-1c rationale. Byte-equivalence
of every rebuilt pattern verified programmatically (alternation-order
neutrality probed: the \b and > anchors make order irrelevant).
2026-08-03 09:56:36 +05:30
Eugeniusz Gilewski
9421c5afdf perf(agent): precompile response and skill-scan regexes (#33208)
strip_think_blocks passed the same response-scrubbing strings through
re's pattern dispatcher on every response. Skills Guard repeated the
same work for 121 patterns against every scanned line.

Compile the existing expressions once and reuse Pattern.sub/search. Keep
each generic tool-call tag in its own paired expression so mismatched
openers retain their payload while existing stray-closer cleanup remains
unchanged.

Part of #33208
Salvaged from #32713 by @ErnestHysa.

Co-authored-by: ErnestHysa <takis312@hotmail.com>
2026-08-03 09:56:36 +05:30
joaomarcos
3b53a3c560 chore: drop stale diagnostic report from PR #77021
teknium1 flagged ISSUE_76870_RELATORIO_CAUSA_RAIZ.md as containing
stale metadata (references an unrelated local branch) and asked to
drop the standalone report, keeping only the focused server.py fix
and regression test.
2026-08-03 09:56:24 +05:30
joaomarcos
a93969dbd1 fix(tui): snapshot history after pending model switch applies (#76870)
Deferred model switches append a marker and bump history_version at
turn start; the dispatcher was snapshotting history before that
mutation, so the version-mismatch guard rejected the turn's own
result as a stale/concurrent write. Move the snapshot to after
_apply_pending_model_switch/_sync_agent_model_with_config, under
history_lock, so the turn's own preparatory mutation is included in
its baseline while the anti-stale guard still catches real external
writes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 09:56:24 +05:30
kshitij
8f61e61991 chore: add contributor email mapping for LFDMcore (core@lfdm.co) (#77331) 2026-08-03 09:55:44 +05:30
kshitij
1228bfb270 Merge pull request #77329 from kshitijk4poor/chore/attrib-ckaznocha
chore: AUTHOR_MAP for ckaznocha
2026-08-03 09:54:46 +05:30
kshitij
b6cdddd3df chore: AUTHOR_MAP for ckaznocha (matrix crypto PRs) 2026-08-03 09:54:28 +05:30
Gabriel Lesperance
68391930d0 perf(tui): bound scroll rendering and preserve anchors 2026-08-03 09:53:51 +05:30
kshitij
7598bc6c95 Merge pull request #77319 from kshitijk4poor/chore/attrib-egilewski-myk0la
chore: contributor email mappings for egilewski and myk0la-b
2026-08-03 09:39:20 +05:30
kshitij
af6efdcd15 chore: contributor email mappings for egilewski and myk0la-b 2026-08-03 09:39:00 +05:30
hermes-seaeye[bot]
d79a6f0e3e fmt(js): npm run fix on merge (#77300)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-03 03:44:44 +00:00
brooklyn!
aa8f0d4c85 Merge pull request #77292 from NousResearch/bb/composer-mixed-paste
Pasting a thread with images keeps the text and skips the blank thumbnails
2026-08-02 22:35:35 -05:00
Teknium
ef9f6effaf fix(cli): persist YOLO mode across --resume
A session's YOLO bypass lived only in the in-memory
tools.approval._session_yolo set (or the process-frozen --yolo env
var), so resuming a session in a fresh process silently reverted the
user's /yolo ON — dangerous commands started prompting again.

Persist a yolo_mode flag in the session row's model_config JSON and
restore it on every CLI resume path:

- SessionDB.set_session_yolo() merges the flag into model_config
  (same lineage-preserving merge as update_session_runtime_lock);
  SessionDB.session_yolo_enabled() reads it back, false on any parse
  failure.
- /yolo toggle persists ON and OFF through the new helper; the
  compression/branch session-id rotation carries the flag onto the
  continuation row.
- --yolo launches record the flag at session creation (agent_init),
  and a /yolo toggled before the lazily-created row exists is carried
  into the creation-time model_config (_ensure_db_session).
- HermesCLI._restore_session_yolo() re-enables the bypass on startup
  --resume/-c, the deferred init path, and mid-chat /resume, with a
  visible '⚡ YOLO mode restored from session' notice. No-op under a
  frozen process-wide --yolo and never enables on absent/garbage flags.
2026-08-02 20:30:06 -07:00
Brooklyn Nicholson
9e373b3828 fix(desktop): paste rich text with images as text, not blank attachments
Copying a Discord thread (or any rich-text selection with images) attached
one or more blank thumbnails and dropped the message text entirely.

Two causes. The clipboard's `text/html` was scraped for inline
`<img src="data:…">` regardless of whether the copy carried its own text —
and what Discord ships beside each image embed is a 32x5 blurhash
placeholder, which is exactly the blank attachment. Then, because any image
blob short-circuited the paste handler, the prose that came with it never
reached the composer.

Inline HTML images now only count for an image-only copy, and are ignored
below a thumbnail-sized floor so spacers and trackers don't attach either. A
mixed paste attaches its real images and still inserts its text.

Also registers pasteAndMatchStyle in the Edit menu — Cmd+Shift+V had no menu
entry, so the chord was never translated into an editor command anywhere in
the app.
2026-08-02 22:28:28 -05:00
Teknium
aac74be2f1 fix(approval): classify CLI/TUI approval timeouts separately from explicit denials
When an approval prompt expired without a response, every CLI-side path
collapsed the timeout into the same 'deny' choice as an explicit user
refusal, so the agent was told the user denied the action when the user
simply never answered. The gateway wait already distinguished the two
('timed out without user response... Silence is not consent.'); this
brings the CLI/TUI/ACP surfaces to parity.

- prompt_dangerous_approval(): input()-path expiry now returns a distinct
  'timeout' choice (still fail-closed).
- cli.py _approval_callback + hermes_cli/callbacks.py approval_callback:
  deadline expiry returns 'timeout' instead of 'deny'.
- check_all_command_guards / _run_approval_gate CLI tails: 'timeout' maps
  to outcome='timeout' with a 'timed out without user response... Silence
  is not consent.' BLOCKED message (matching the gateway wording);
  explicit deny keeps outcome='denied' and gains user_consent=False for
  shape parity.
- computer_use: 'timeout' verdict threads through the CLI adapter and
  yields a 'prompt timed out — the user did not respond' error instead of
  'denied by user'.
- ACP permissions bridge: FutureTimeout returns 'timeout' (other failures
  still 'deny'); elicitation maps 'timeout' to 'cancel' like the gateway's
  unresolved outcome; codex wire mapping documents deny/timeout→decline.
- write_approval already treats unknown choices as 'stage, not drop', so
  a timeout now stages the memory write instead of silently refusing it.

Every timeout path remains fail-closed — the action never runs; only the
classification reported to the agent changes.
2026-08-02 20:21:59 -07:00
brooklyn!
f4604f89de Merge pull request #77275 from NousResearch/bb/tab-reload
Reload a tab from its right-click menu
2026-08-02 22:14:27 -05:00
Brooklyn Nicholson
17bdc99db4 feat(desktop): reload a tab from its right-click menu
Right-click any tab and pick Reload: the pane's content remounts in
place — effects re-run, state resets, measurements are retaken — while
the tab keeps its slot and every other tab is untouched.

A per-pane epoch atom keys the contribution inside the zone body, so
reload never rewrites the layout tree. Both tab menus offer it: the
zone strip menu (tool panels, the file tree, a fresh draft's main tab)
and the session tab menu (tiles + the loaded main tab).
2026-08-02 22:08:09 -05:00
Teknium
a4a91610b0 test: cover gateway per-turn reload and flip dump terminal-backend pin
Follow-up for the salvaged #29239: regression test drives the real
gateway _reload_runtime_env_preserving_config_authority() path with a
stale .env TERMINAL_ENV=docker vs config.yaml terminal.backend=local,
and the hermes debug dump test that pinned the old stale-env-wins
symptom now pins the fixed contract (config wins, override line kept
as defense-in-depth for post-load env mutation).
2026-08-02 18:21:58 -07:00
Jiahui-Gu
e471c7165e fix(env): make config.yaml authoritative for terminal.backend (#29186)
A leftover TERMINAL_ENV in ~/.hermes/.env (written by `hermes setup` or
shell exports) was silently overriding terminal.backend in config.yaml,
so users switching from docker to local saw `hermes config show` agree
with their change while the gateway / cron / batch_runner still ran
against the old backend.

load_hermes_dotenv now re-applies config.yaml's terminal.* values on top
of whatever the .env files set, so the documented source of truth wins
for every entrypoint that goes through the loader.

Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
2026-08-02 18:21:58 -07:00
Teknium
a6defd4f15 fix(skills): match evidence quotes through markdown markup
Live-run findings from a real fact-checking task (ankylosing spondylitis
genetics, 7 authoritative sources) against the new mode:

- Verbatim check rejected a legitimate quote because web_extract returns
  markdown: the MedlinePlus sentence is "including _[ERAP1](https://...)_,
  _[IL1A](...)_" on the wire but plain prose to a reader. The agent was
  forced onto a weaker evidence fragment — the opposite of the point.
  Matching now canonicalizes inline links to their label and drops
  emphasis/code markers and backslash escapes on both sides, so quoting
  the sentence a reader sees works. Paraphrases are still rejected.
- Escaped asterisks (HLA-B\*27) no longer have to be reproduced in the
  quote, so extractor artifacts stop leaking into rendered evidence.
- New `render --replace-in <draft>`: rewrites a draft's Sources block in
  place, idempotently. Previously the only path was hand-slicing the
  file, which also tripped over the emitted heading being `## Sources`
  while the prose said "Sources:".
- verify stats: report the provenance total that the percentage is
  actually computed from (cited + [unverified], counted once), and print
  the line as `info:` instead of `warn:` when nothing is wrong. The old
  line printed 17 cited / 2 unverified next to 72%, which does not
  reconcile — a sentence can be both.
- SKILL.md documents the emitted heading, --replace-in, and exactly what
  counts as a prose sentence for --min-coverage.

7 new tests (47 total) using the real MedlinePlus/Frontiers markup;
6 sabotage runs, all red.
2026-08-02 16:18:28 -07:00
Teknium
4660673a3c feat(skills): add fact-checking mode to grounded-citations
Extends the citation ledger with evidence-backed fact-checking:

- New `quote` subcommand attaches verbatim supporting quotes to a
  source; the quote is rejected unless it appears verbatim
  (whitespace/case-insensitive) in the fetched page text, so a
  paraphrase or misremembered figure cannot masquerade as evidence.
- `verify --evidence` fails a draft whose cited sources carry no
  attached quote.
- `render --style evidence` prints each source's quotes beneath its
  URL, showing the claim -> source -> exact-text chain.
- `[unverified]` marker declares model-knowledge claims; counts toward
  --min-coverage so provenance is declared for every sentence without
  forcing fake citations.
- SKILL.md: new Fact-Checking Mode section + pitfalls; version 1.1.0.
- 10 new tests (40 total), all proven live by sabotage runs.

Covers the fact-checking/evidence-transparency half of #28289.
2026-08-02 16:18:28 -07:00
Teknium
8d755d0748 chore: contributor email mappings from open attribution PRs
Consolidates the surviving mappings from PRs #53509 (@xqdwww),
#29876 (@mihalyschroth), and #32842 (@vizicist, for @waefrebeorn's
bounty account) into contributors/emails/ files — the frozen
AUTHOR_MAP in release.py is not edited (the reason those PRs
couldn't merge as-is). Entries from those PRs already present on
main are skipped. All bare-noreply logins verified via the GitHub
users API.

Salvaged from @xqdwww (#53509), @mihalyschroth (#29876),
@vizicist (#32842).
2026-08-02 16:18:25 -07:00
Teknium
58e3dcf3d6 chore: round-2 review nits (re-review #9)
- tests/agent/test_session_activity.py asserts against
  ACTIVITY_DESCRIPTION_MAX instead of the literal 120.
- The session-stall WARNING log line names its config knob
  (agent.session_stall_timeout) so operators can find the setting.
- hermes_state.py: collapse the triple blank line near line 191.
- hermes_cli/status.py no longer imports the private
  hermes_cli.main._relative_time: the helper moved to a public home
  (hermes_cli.timefmt.relative_time); main._relative_time stays as a
  thin back-compat wrapper (sessions_cmd and external patchers keep
  working).
2026-08-02 16:16:36 -07:00
Teknium
44c362889a test(agent): deflake progress-extension timing per FLAKY policy (re-review #8)
test_progress_extends_idle_budget_until_success raced wall-clock: the
0.1s-idle/0.04s-tick shape left ~60ms of slack per tick, so one slow
scheduler pass on a loaded CI box lapsed the idle budget mid-loop.
Widened to 0.5s idle / 0.1s ticks (5x per-tick margin, total runtime
still <1s) per the FLAKY policy's minimum-margin guidance.
2026-08-02 16:16:36 -07:00
Teknium
b2963f8034 docs: document agent.session_stall_timeout + compression timeout keys (re-review #7)
- website/docs/user-guide/configuration.md (en) and the zh-Hans
  translation gain a 'Session Stall Watchdog' section: default 300,
  0=disabled, notify-only semantics (never kills the turn — contrast
  gateway_timeout), one notification per stall episode, and the exact
  stall message text so it is greppable.
- cli-config.yaml.example: the two in-agent compression timeout keys
  (compression.context_timeout_seconds /
  compression.context_total_ceiling_seconds) are shown as commented
  lines next to session_stall_timeout's example for discoverability.
2026-08-02 16:16:36 -07:00
Teknium
a0e700c4cf feat(agent): emit pool_saturated compression-attempt telemetry (re-review #6)
The fail-fast admission path (bounded compress pool, F6) only logged a
WARNING; in the compression-attempt telemetry stream a wedged pool
looked like compression simply stopped being attempted. Emit the
existing attempt telemetry with failure_class='pool_saturated'
(commit_status=aborted, split_status=aborted) on refusal, following
_emit_compression_attempt_telemetry's existing call shape. Regression
extends the F6 saturation test (sabotage-verified).
2026-08-02 16:16:36 -07:00
Teknium
06bdc48c1f perf(agent,gateway): back cancel-wait polls off from 1ms to 25ms (re-review #5)
The fence-cancel poll loops (sync host wait in conversation_compression,
async hygiene wait in gateway/run) spun at 1kHz while the worker held
the fence through its lock-setup window — which rides SessionDB write
patience and can last seconds. 25ms keeps sub-tick cancel latency
without the spin.
2026-08-02 16:16:36 -07:00
Teknium
0628e33347 fix(agent): clear archived parent's activity labels after rotation (re-review #4)
The compression heartbeat's terminal 'context compression completed'
stamp force-persists against the PARENT session id (agent.session_id at
stamp time). After the out-of-place rotation the parent is archived but
kept advertising a fresh last_activity_at + terminal label forever.
Clear the parent row's activity labels best-effort after a committed
rotation (keeps last_activity_at so idle clocks stay continuous; the
child carries live labels). Regression asserts the archived parent's
labels are cleared while the child's lineage is intact
(sabotage-verified).
2026-08-02 16:16:36 -07:00
Teknium
53a5983af0 chore(state): bump SCHEMA_VERSION to 24 for activity-tracking columns (re-review #3)
The last_activity_at/description/provenance columns already live in
SCHEMA_SQL and the column reconciler; existing DBs heal via the
reconciler, but the version stamp must advance so downgrade/upgrade
tooling sees the new layout. No version-literal test assertions exist
(tests compare against the imported constant).
2026-08-02 16:16:36 -07:00
Teknium
58f0fe305d fix(gateway): bound the stall-notify adapter.send (re-review #2)
A wedged adapter transport (network hang, dead websocket) previously
blocked _check_session_stalls forever: sibling candidates in the same
pass were never evaluated and the watcher stopped ticking. Wrap the
send in asyncio.wait_for (15s); on timeout log a WARNING and do NOT
latch, so the next tick retries. Regression uses a never-resolving fake
adapter and proves the pass completes, a healthy sibling candidate is
still notified in the same pass, and the watcher ticks again
(sabotage-verified against the unbounded send).
2026-08-02 16:16:36 -07:00
Teknium
0277cc48bd fix(agent): never release the durable compression lease mid-commit (re-review #1)
revoke_commit_admission() used to invoke the holder-qualified lease
release unconditionally — including while an admitted commit was still
mutating SessionDB — letting a second compressor acquire the durable
lock mid-commit and interleave with the first commit's writes.

The admission_revoked flag store stays lock-free, but the lease-release
decision now coordinates with the fence lock:
- revoke acquires the fence lock non-blocking; on success no commit can
  be in flight (an admitted commit retains the lock until finish_commit)
  and the release runs immediately, still under the lock so a racing
  begin_commit cannot slip between the check and the release.
- on failure the release is deferred: finish_commit() re-checks
  _admission_revoked and performs it AFTER the commit completes (prompt
  even if the worker thread is later parked), and the begin_commit
  refusal path does the same for a revoke that lost the race to a
  transient lock-setup/cancel boundary. All paths are idempotent with
  the worker's own outer cleanup (DB release is holder-qualified).

Invariant encoded + tested: no second compressor can acquire the durable
lock while an admitted commit is still mutating; after a post-revoke
commit finishes the lease is released promptly. Both regressions
(revoke-during-commit deferral, revoke-before-commit immediate release +
refused begin_commit) are sabotage-verified.
2026-08-02 16:16:36 -07:00
Teknium
15267a1d2d fix(agent): reconcile explicit hard-cancel (main d15b638a88) with pooled fence rework
Rebase onto origin/main brought in 'let explicit interrupts cancel safely',
which predates this branch's pooled progress-timeout + F1-F6 fence rework.
Reconcile the two:

- begin_commit(cancel_event) re-checks the hard-cancel Event under the
  fence lock again (lost in the mechanical rebase).
- compress_context: restore aux_interrupt_protection around the summary
  call, the post-return frozen-cause AuxiliaryExplicitCancellation check,
  and the full rollback/telemetry handler for explicit interrupts.
- run_agent._compress_context: recreate the per-attempt fence registration
  (_active_compression_commit_fence) that hard_interrupt() uses to
  serialize cancel admission, and thread that exact fence through both
  the direct and pooled paths (run_compress_context_with_progress_timeout
  now accepts an external fence).
- test: the pooled worker isolates the live transcript (F3), so the
  hard-interrupt rollback regression mutates the engine's input snapshot
  rather than reaching around it to the caller's list.
2026-08-02 16:16:36 -07:00
Teknium
fe2fc57240 test: skip live aux feasibility probe in worker-isolation regressions
Same hermetic-CI trap the concurrent-fork suite already guards against:
without credentials the one-time feasibility probe aborts compression
before the stubbed engine starts, so the F3/F4 blocked-state assertions
went vacuous-false in CI while passing on credentialed dev machines.
2026-08-02 16:16:36 -07:00
Teknium
038c1ad872 docs(gateway): precise watchdog scope, explicit import-resets-activity contract (review S4)
- Config docs now describe session_stall_timeout precisely: a RECOVERY
  notifier for an in-process AIAgent with an adapter-queued follow-up —
  not a general gateway/session stall detector — with a per-AIAgent scan
  cadence (not globally coordinated per durable session).
- import_sessions documents the deliberate export-includes /
  import-resets asymmetry for the activity fields (no resurrected
  'working' labels on machines where no agent runs), with a regression
  pinning both halves.
- Strip trailing whitespace in contributors/emails/fangliquan@qq.com
  (git diff --check housekeeping).

PR #76354 review, scope/contract items + housekeeping.
2026-08-02 16:16:36 -07:00
Teknium
89c4e26e23 fix(gateway): revalidate stall candidate immediately before /new delivery (review S2)
The stall watchdog gathered pending/activity candidates and later sent
the recovery notification from that aging snapshot — an agent that made
progress (or drained its queue) between the scan and the send received a
false stall notice mid-recovery.

Re-read the adapter pending slot, the overflow queue, and the live
activity snapshot immediately before delivery; abort the send and re-arm
the latch (pop it) when the candidate is no longer stale, so a future
genuine episode still notifies.

Race regressions: progress between scan and send aborts delivery;
pending-drained between scan and send aborts delivery; a genuinely
still-stale candidate is still delivered exactly once.

PR #76354 review, 'watchdog can send /new using a stale snapshot' /
merge gate 8.
2026-08-02 16:16:36 -07:00
Teknium
92c736919d fix(state): sub-second busy budget for observational activity writes (review S1)
Activity heartbeat writes and turn-end label clears ran synchronously on
the response-critical path with the full ~20s routine write-patience
budget — under contention an otherwise-finished reply could stall for
seconds just to update observation labels, mimicking the very stall the
watchdog detects.

touch_session_activity and clear_session_activity_labels now use a
dedicated 0.5s patience budget (they are observation-only; the next
heartbeat window retries naturally), and a no-op label clear (labels
already empty) skips the write transaction entirely.

Regressions: with another connection holding BEGIN IMMEDIATE, both writes
give up well under the routine budget; the no-op clear performs zero
write transactions.

PR #76354 review, 'activity writes are synchronous on critical paths' /
merge gate 9.
2026-08-02 16:16:36 -07:00
Teknium
99100843c6 fix(agent,gateway): charge the idle wait from the last progress event (review S3)
Both progress-aware waits (sync compress wrapper and gateway session
hygiene) slept a FULL idle interval and only then compared progress, so
progress early in an interval let silence approach 2x the configured
idle timeout before the waiter noticed. Compute each wait slice as
idle_timeout - elapsed_since_last_progress instead.

Regression: a worker that reports progress early and then goes silent is
timed out in ~1x the idle budget, not ~2x.

PR #76354 review, 'idle timeout can allow nearly twice that silence'.
2026-08-02 16:16:36 -07:00
Teknium
abc0db8cde fix(agent): bounded admission + stale-job cancellation for the compress pool (review F6)
The process-wide 4-worker pool retained the stdlib executor's unbounded
queue: four hung summaries wedged every slot, a fifth compression queued
silently, waited out its whole budget without starting, and remained
eligible to run later as an expensive stale job whose fence was already
cancelled (the first fence check used to sit AFTER the summary call).

- Bounded admission: submission fails fast (messages returned unchanged,
  loud warning) when all pool slots are occupied; slots are freed by a
  future done-callback. Recovery contract documented at the constant: new
  work fails fast while wedged, wedged workers are fence-cancelled and
  restore service when they return; a worker that never returns costs its
  slot — bounded, observable degradation instead of unbounded queueing.
- Not-yet-started futures are cancel()ed on timeout.
- The cancelled fence is checked BEFORE any expensive summary work, both
  in the pooled wrapper (stale queued job) and inside compress_context
  (pre-summary gate), so a stale job never burns an LLM call or acquires
  session state.

Saturation regression: 4 event-blocked summaries wedge the pool, a 5th
submission fails fast (asserted while the four are provably still
blocked), the refused job never runs after worker recovery, and a fresh
submission after recovery succeeds.

PR #76354 review, blocking finding 6 / merge gate 7.
2026-08-02 16:16:36 -07:00
Teknium
36a60c5edc fix(agent): rebind the caller's session ContextVar after out-of-place rotation (review F5)
Session rotation runs on the pooled worker thread, whose copied context
gets the child id — the CALLER's ContextVar still holds the parent, and
get_session_env() prefers a bound ContextVar over os.environ. Tools and
subprocesses invoked on the caller thread after a compression.in_place=false
rotation therefore saw the STALE parent HERMES_SESSION_ID.

After the pooled wrapper returns, rebind the session id in the caller's
own context (set_current_session_id) alongside the existing logging
repair; idempotent when no rotation happened.

Behavioral regression: with the gateway-style bound session context, a
post-compression get_session_env("HERMES_SESSION_ID") read on the caller
thread now returns the child id.

PR #76354 review, blocking finding 5 / merge gate 6.
2026-08-02 16:16:36 -07:00
Teknium
fdeb09a596 fix(agent): holder-qualified durable lease cancellation, cooldown ordering (review F4)
A host timeout previously left the timed-out worker holding the durable
per-session compression lock AND refreshing its lease indefinitely, so a
truly hung summary blocked every later compression attempt; and a LATE
successful summary could clear the failure cooldown the host had just
recorded.

Transplant the lease-cancellation invariants from PR #71569
(@ciabata-git): the worker publishes an idempotent, holder-scoped release
hook on the fence once it owns the durable lock (begin_lock_setup /
register_cancelled_lock_release close the acquire→publish race), the
refresher start is serialized against the release path, and the host
invokes the hook on idle timeout, hygiene timeout, and every unwind
(revoke_commit_admission now also releases). ABA safety: the SessionDB
release is holder-qualified (DELETE ... WHERE holder = ?), so a stale
release can never free a replacement holder's lease.

State ordering: the compressor consults a fence-cancellation check BEFORE
clearing the failure cooldown, so a late worker cannot undo the host's
timeout cooldown; the check is installed only for the fenced call and
removed in a finally.

Regression implements the reviewer's exact 5-step scenario: summary
blocked indefinitely → host timeout → a NEW compressor acquires the
durable lock while the old summary is STILL blocked → old worker released
→ it cannot clear cooldown, release the new holder's lease, or publish
stale state.

PR #76354 review, blocking finding 4 / merge gates 4 + 5.

Co-authored-by: ciabata-git <ciabata-git@users.noreply.github.com>
2026-08-02 16:16:36 -07:00
Teknium
971d81f892 fix(agent): isolate the pooled compression worker from the live transcript (review F3)
The pooled worker closure captured the caller's live `messages` list and
compress_context explicitly supports plugin/legacy context engines that
mutate that list in place — so after a host timeout, a late engine could
rewrite the live conversation (roles, ordering, persisted content)
concurrently with the resumed turn.

The worker now deep-snapshots the transcript on the worker thread before
any engine code runs; the caller's list object is never handed to pooled
code. Results reach caller-visible state only through the returned value
of an ADMITTED commit (the host discards results on timeout/cancel), and
durable SessionDB mutation was already gated behind the commit fence.
No-op passes map the unchanged snapshot back to the caller's original
list so identity-based no-op detection and flush dedup keep working.

Document the thread-safety contract for context-engine and
memory-provider extension points (they now run on pooled threads) in the
module docstring and the context-engine plugin guide.

Regression: an in-place-mutating engine plus host timeout proves the
caller's live transcript is byte-identical WHILE the worker is still
blocked inside the engine (released only after the assertions).

PR #76354 review, blocking finding 3 / merge gate 3.
2026-08-02 16:16:36 -07:00
Teknium
efdd229884 fix(agent): revoke commit admission on every host unwind (review F2)
The sync compress wrapper only handled concurrent.futures.TimeoutError;
KeyboardInterrupt, task cancellation, or any other exception while
waiting let the host unwind while the detached worker kept full commit
authority — it could later enter the commit fence and mutate durable
state (in-place archival, session rotation) behind the caller's back.

Wrap the whole host wait in try/finally: any exit that did not settle the
worker (returned result or won the fence race) revokes future commit
admission via a new lock-free CompressionCommitFence.revoke_commit_admission()
(begin_commit re-checks the flag under the fence lock, so no admitted
commit is ever abandoned mid-mutation). The gateway hygiene wait gets the
same guarantee via a BaseException handler that revokes admission and
defers helper cleanup until the worker actually returns.

Reconciliation with PR #74449 (suparious): that PR routes EXPLICIT host
interrupts into auxiliary-call cancellation; this change is the
complementary host-side guarantee that no unwind — explicit or not —
leaves an unfenced worker. The two compose (fence revocation here is the
outer safety net; #74449's aux cancellation remains the fast path) rather
than duplicating one another.

Regressions: KeyboardInterrupt and generic-exception unwinds assert the
fence is revoked WHILE the worker is still blocked pre-commit, then
release the worker and prove begin_commit() is refused.

PR #76354 review, blocking finding 2 / merge gate 2.
2026-08-02 16:16:36 -07:00
Teknium
980aea225e fix(agent): observe commit phase without the fence lock (review F1)
begin_commit() retains the fence lock until finish_commit(), so a hung
SessionDB commit made try_cancel_before_commit() return None forever and
the host spun ahead of the overrun-warning loop — a genuinely hung commit
stayed unbounded AND silent. Add a lock-free phase marker (threading.Event
set inside begin_commit while the lock is held, readable without it) and
break the host spin on commit_in_flight so the bounded overrun loop — and
its WARNING + on_commit_overrun surfacing — is reachable WHILE the commit
is still blocked. Applies to both the sync compress wrapper and the
gateway session-hygiene wait.

Regression asserts the warning and callback fire while the event-gated
fake commit is still blocked; the test releases the worker only after
those assertions (addresses helix4u's released-before-asserting callout).

PR #76354 review, blocking finding 1 / merge gate 1.
2026-08-02 16:16:36 -07:00