main
154 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b90b7ae7ed | fix(delegate): title subagent sessions after their goal without a model call | ||
|
|
444066b82c |
fix: exclude display-only message fields from token estimates
Edit tool rows now carry display_metadata.tool_result_metadata.inline_diff
(~9KB of ANSI per edit) on the live message dict. build_api_messages strips
display_kind/display_metadata/_row_id before the wire, but the local
estimator's wire shadow only dropped PERSISTENCE_ONLY_MESSAGE_FIELDS
({"timestamp"}), so estimate_messages_tokens_rough and
estimate_request_tokens_rough priced the diff: one edit row went from 32 to
2527 estimated tokens (100 edits: 3100 -> 252600). That inflates compaction
preflight, post-tool checks and turn-overflow scoring, compacts early and
breaks the prompt cache.
PERSISTENCE_ONLY_MESSAGE_FIELDS (agent/message_metadata.py) is now the single
set of local-only fields: timestamp, display_kind, display_metadata, _row_id.
build_api_messages pops exactly that set and the estimator shadow drops it,
so the two can no longer drift. The iteration-limit summary path
(_iteration_summary_api_messages) hand-builds its wire messages and stripped
timestamp but not display_*; its key set now unions the same constant, so
the inline diff never reaches the provider there either (strict gateways
reject unknown keys). The tail-budget walk (_estimate_msg_budget_tokens) is
allowlist-based and already ignored these fields.
|
||
|
|
afa0ddeaab |
fix(agent): tag background-review fork turns in logs and pin supersede interrupt scoping
Fork turns (background review, side questions) reuse the parent's session_id and, when unrouted, the model, so their `conversation turn:` and `Turn ended:` lines are indistinguishable from live turns. In #118693 this made the fork's own superseded exit read as a killed foreground stream. - build_cache_parity_fork stamps _turn_origin on fork agents - _log_turn_exit and the conversation-turn line append origin=<write_origin> for fork turns only; live turns are byte-identical - new invariant test: cancel_background_review_for_live_turn flags only the fork (parent instance flags and thread bit untouched) via the real InterruptControlMixin path - new regression test: fork turn-exit lines carry origin=background_review (red on base) Related to #118693 Co-authored-by: CommandCodeBot <noreply@commandcode.ai> |
||
|
|
bb87e6abce |
fix(agent): persist the multimodal pre_llm_call text part and stop double injection (#71998)
Builds on #72026 (@PRATHAMESH75): list content carries the turn's memory-prefetch / pre_llm_call context as a durable text part appended once in the prologue, in every api mode (MoA and codex_app_server included), so the request, the persisted row, compaction and a later resume all see the message the model saw. Persistence gap from the #72026 review: in-place preflight compaction (and a close/early flush that races the prologue) writes the current user row BEFORE the part exists and the crash persist identity-skips that dict, so a resumed session replayed the turn without the context. The list branch now pushes the appended part into that row via set_user_message_content under the same _row_id-under-lock protocol as the string sidecar backfill, keeping the writer's shape (compaction: raw parts; flush: text projection). Titling moves before the injection step so a list turn's title is derived from the user's ask, not the injected tail. Tests trimmed to one invariant per layer: hook edit reaches the wire on a list turn and replays after reload; in-place compaction + reload keeps the part (red without the backfill); memory query flattens parts. |
||
|
|
d1267d8045 |
fix(agent): run memory prefetch on multimodal turns by flattening the query (#71998)
The delivery half of this PR lets a multimodal (list) turn carry the memory /
pre_llm_call context via a durable text part. But the execution half never ran
on those turns: `_memory_turn_start_and_prefetch` keyed the query off
`isinstance(str)`, so a list turn collapsed to `_query = ""` — `on_turn_start`
saw an empty turn and `is_trivial_prompt("")` skipped `prefetch_all` entirely.
So on an image+text turn recall never fired and the sidecar had nothing to
deliver, silently (no `recall` audit rows, no `prefetch failed` warning).
Flatten str/list content to its text via `flatten_message_text` before building
the query. An image-only turn still flattens to "" and is correctly treated as
trivial (no semantic text to query on); a text+image turn now runs prefetch on
its text. Adds unit coverage for the flattening and the trivial-prompt gate.
Reported by @albert748 on #72026.
|
||
|
|
b8b519c398 |
fix(agent): deliver pre_llm_call/memory context on multimodal turns (#71998)
compose_user_api_content returns None for non-string (multimodal) user
content, so the pre_llm_call plugin context and memory-prefetch block —
which ride the string api_content sidecar — were silently dropped on
image-only/attachment turns. A profile plugin's pre_llm_call {"context":
...} return could route text turns but never image-only turns.
Deliver the composed context as a durable text part on multimodal turns
via the existing append_notes_to_multimodal_content channel (the same one
gateway must-deliver notes already use), so the injected context reaches
the model and the wire stays byte-identical to the persisted/replayed
transcript. String turns keep the api_content sidecar path unchanged; the
MoA/codex_app_server guards are preserved. Extract the shared injection
composition into _context_injection_parts so both paths inject byte-
identical context.
|
||
|
|
1ae6f650f9 | fix(compaction): invalidate pre-checkpoint usage anchors | ||
|
|
28a75285c4 |
refactor: drop the dead _vision_supported flag
Image rejections are now tracked per (provider, model) in agent._image_rejecting_models, and recover_before_classification gates on that set. Nothing reads agent._vision_supported any more, so the write in turn_recovery and the per-turn reset entry in turn_context are dead state. Remove both so the next reader doesn't assume a turn-global vision gate still exists; update the two tests that asserted or seeded the attribute. |
||
|
|
afc3b7c6f3 |
feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash merges ( |
||
|
|
efc947d72a |
fix: send the title model call after the turn on a shared custom endpoint
On a `custom` main route (llama.cpp, Ollama, vLLM, ...) whose
auxiliary.title_generation is not pinned elsewhere, the turn prologue fired the
`response_format: json_schema` title request on a daemon thread at the same
instant as the turn's own streaming request, against the same self-hosted
server. A single-slot server can decode the title grammar/completion into the
main reply: the user then receives `{"title": ...}` as the assistant turn, the
main loop persists it as a genuine assistant row, replays it, and the model
adopts the format (#117296). No Hermes writer routes the aux response into the
transcript; the leaked JSON is the main completion itself.
`maybe_auto_title` now returns the upgrade thread and leaves it UNSTARTED when
`title_upgrade_must_wait_for_turn(main_runtime)`; the prologue parks it on
`agent._deferred_title_upgrade` and `finalize_turn` starts it once the model
has answered. Hosted providers keep the turn-start timing. Usage accounting
(`task='title_generation'`) and `sessions.title` are unchanged.
|
||
|
|
d03d6c2b39 |
fix(compression): an over-window session that cannot shrink ends the turn with /new guidance and waits one idle budget, not the ceiling
A session far above the model window (~356k tokens on a 131k window in #116472) re-ran context compression on every turn: a preflight pass that reclaimed nothing still let the request go to the provider (400 -> overflow handler -> another pass), and a summary stream that kept emitting tokens while never committing held the pre-commit wait to the full 600s ceiling. On the Desktop that blocked the gateway event loop for 10-20 minutes per turn and the renderer was eventually killed. - agent/turn_context.py::_fail_closed_on_insufficient_progress: when a preflight pass makes no (or sub-5%) progress and the request provably exceeds the model window, raise PreflightCompressionTimedOut with "start a new session (/new)" guidance so no provider call is sent. An unknown window or a fitting request keeps the send-as-is behaviour; a pass that no-op'd on a transient guard (summary-failure cooldown) keeps its typed cooldown result. Called from both insufficient-progress branches of turn_context_compaction._run_preflight_passes. - agent/conversation_compression.py::run_compress_context_with_progress_timeout: an over-window request's pre-commit wait is bounded by one inactivity budget (compression.context_timeout_seconds) instead of context_total_ceiling_seconds; the existing first-stall deterministic fallback then carries the compaction. Config-derived, no new knob. Slim slice of #116592's Python half. Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com> |
||
|
|
e85cb94da1 | chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) | ||
|
|
c8ecc3db64 |
fix: drop replayed reasoning_details on every chat-completions route that does not read it
OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning continuity; every other OpenAI-compatible route either ignores the field or, when its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole request with 400/422 once an earlier reasoning turn is in history — wedging the session after an in-session model switch (#70233). Strip the field from the wire copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url), mirror it in the auxiliary wire boundary and the iteration-summary path; state.db history keeps the field so switching back to OpenRouter/Nous replays it again. |
||
|
|
f7567a62af |
fix: hand Codex reasoning-only stalls to the fallback provider instead of the incomplete sentinel
Three consecutive Codex Responses answers that carry only (encrypted) reasoning — no visible text, no tool call — used to exhaust the 3-continuation budget and end the turn on "Codex response remained incomplete after 3 continuation attempts", never touching configured fallback_providers (#67321). Encrypted reasoning items replay byte-for-byte, so a bare retry deterministically repeats the stall. - Track a per-turn `_codex_reasoning_only_streak` apart from the aggregate `_codex_incomplete_retries`: a visible partial resets the streak, so the mixed partial-then-stall variant still reaches its own recovery threshold while the turn-wide iteration budget stays the hard bound. - At streak 3, `continue_codex_incomplete` activates the next fallback with the semantic `FailoverReason.incomplete_response`, grants exactly one grace call when the trigger consumed the last iteration, and returns `CODEX_FALLBACK_ACTIVATED`; the intake re-syncs the Model:/Provider: identity on the system prompt. - Off the Codex wire the synthetic continuation nudge is stripped alongside the opaque replay state (`drop_nudge_marker`) so the Chat Completions payload keeps valid role ordering and no Codex-only control text. - No fallback configured: unchanged terminal sentinel, still bounded at 3 calls. Ported from PR #67336 by @PRATHAMESH75 onto the decomposed agent/turn_*.py siblings. |
||
|
|
8d25e69b6e | fix(desktop): use generated paste previews for titles | ||
|
|
303bcd804a |
fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request (even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the suggested recovery reproduced the same loop. - request_exceeds_model_window(agent, tokens): one predicate, two consumers. - Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does exactly this every turn); the summary-failure cooldown stops the retry from repeating. - Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline). compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript. - Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars, prompt build ms) so a stalled attempt is distinguishable from a slow prompt build. - Deterministic-rung wording no longer claims "again after a stall backoff". Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True, compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed (main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes. |
||
|
|
cd3de040ab |
feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression (PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge; the commits interleave with a cron delivery-ledger rework that the salvage removes in follow-up commits, so per-commit cherry-picks were not practical. Adds display.suppress_warning_notifications (global + per-platform, default false): one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning / emit_media_warning / warning_text, a notification_category classification carried through wakes, queues and persistence, and render/present boundaries for CLI/TUI. |
||
|
|
6f8389d67b |
fix(aux): keep the /btw and title snapshots duck-typed; make the out-of-turn header tests discriminating
Follow-up on the salvaged #112721 commits (@fangliquanflq): - agent/turn_context.py, tui_gateway/methods_prompt.py: add ``session_id`` to the explicit snapshot dicts instead of switching them to ``agent._current_main_runtime()``. The titling prologue is duck-typed (its tests drive a minimal stub) and the snapshot values feed the ``runtime_validator`` equality checks, where ``_current_main_runtime()``'s ``"" `` for a missing attribute would no longer match the live ``None``. Same outcome — the background request inherits the conversation's ``x-opencode-session`` — without changing what the callers read. - tests/agent/test_opencode_session_affinity.py: the salvaged title test passed on an unfixed tree because the titler thread republishes the conversation contextvar (``set_conversation_context``) and the affinity header falls back to it. Replace it with two invariant tests (sync + async ``call_llm(main_runtime=...)`` with EVERY ambient source unset via the ``out_of_turn`` fixture): red on origin/main, green here; both also pin that the explicit binding does not leak past the call. - website/docs/integrations/providers.md: name the background/out-of-turn auxiliary calls the header now covers. Fixes #112717 |
||
|
|
de03c34233 | fix(auxiliary): preserve session in out-of-turn snapshots | ||
|
|
24ae31f7e7 | fix(gateway): retain pre-admission interrupted input | ||
|
|
5c4c31cf4d |
refactor(agent): fail loudly on a missing turn clock; single lowercase in is_dangerous_confirmation
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the turn prologue now raises instead of silently falling back to per-request wall time, which would re-create the mid-turn drift the fix removes. Only production caller (assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the tripwire _inflight_turn_started so the two clocks are not "unified" by mistake. - is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path). - Tests: one _send(idx=) helper instead of three spellings of the builder call; the untrustworthy-stamp contract is its own test. |
||
|
|
820d3ca65d |
fix(agent): send-path canonicalization is prefix-only, clock is admission time, corrupt stamps fail closed
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the three blockers raised on that thread plus one regression the salvage found: - Prefix-only on the send path. build_api_messages now canonicalizes only messages[:current_turn_user_idx]; rows the current turn appended (its own tool calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block the previous iteration had already sent whenever a tool result matched the interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists to remove, and it also made the dangling-tail transform order-dependent on when the user row was appended. - Exact interrupt marker. is_interrupted_tool_result matched "exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary `grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal result to an orphan notice (or dropped a read-only block). That heuristic was tolerable at resume time only; it now runs per request. Match the executors' bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else. - Admission-time clock. The frozen expiry clock was the input's platform-event stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation live on the send path while replay expired it. _reset_per_turn_agent_state stamps time.time() once at admission; the three other writes (bind identity, stage message, build_api_messages write-back under suppress(Exception)) are gone. - Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`, `"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and treat an unknowable age as expired; missing stamps (legacy rows) are still left alone. - Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of _reset_per_turn_agent_state (only the test double needed it). - Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize → ChatCompletionsTransport bytes, equal to the send path with sidecars applied and the durable list untouched, live tail preserved; admission-clock freeze across iterations + corrupt-stamp fail-closed. Each is red under the matching mutation (send path unpatched, whole-list canonicalization, per-request clock, fail-open, loose heuristic). |
||
|
|
401fef6e6f | fix(agent): freeze turn confirmation expiry and verify wire parity | ||
|
|
e5ca5207de | fix(agent): unify replay history canonicalization | ||
|
|
55b3ea0b11 |
fix(gateway): count each bot message once in the loop guard and consume the author variable
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it again, and the busy path asks a third time. Each call counted one loop-guard event, so a Telegram bot tripped the budget after a third of the configured messages. The verdict now only refuses a chat that is cooling down. The ingress gate counts an admitted bot message once. `parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag, and returns None for an author with neither id nor name. Names keep format characters and non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive number. Issue numbers move out of code comments. |
||
|
|
8969511209 |
feat(memory): carry the turn's author to sync_turn
on_turn_start already received the author trio. sync_turn did not, so a provider that wanted to write the turn under its author had to stash state between the two hooks. sync_turn now takes turn_author as a keyword-only argument, and MemoryManager sends it only to providers whose signature accepts it, so existing providers keep working unchanged. build_turn_context resets the author on the agent at the start of every turn so a cached gateway agent never carries a bot author into the next human turn. agent/turn_author.py holds the parsing and the HERMES_TURN_AUTHOR carrier. MemoryProvider.identity_signature() is a new optional hook: the identity values a provider writes under, declared by the provider itself, for the gateway's agent cache to key on. |
||
|
|
70b1ff6930 |
feat(memory): carry the turn's author into the memory-provider contract
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include: remaining_tokens, model, platform, tool_count" — and `MemoryManager` forwards whatever it receives. Its only caller passed nothing, so a memory provider had no way to learn who wrote the turn it was being told about. Providers that key durable state on identity resolve one identity when the session is created. A shared session does not work that way: threads are shared by default (`thread_sessions_per_user` is False), so alice, bob, and another agent all write turns into a session whose peer is whoever spoke first. The gateway's answer today is the `[name]` prefix it prepends to the message text, which the model reads and a provider cannot. `turn_author` now travels from the gateway through `run_conversation` into `build_turn_context`, which forwards `author_id`, `author_name`, and `author_is_bot` to every provider. It stops there — the trio never reaches the model, and providers that ignore the kwargs are unaffected. The bot flag is sent on every transport, not only shared sessions: a provider deciding whether a turn may write to durable memory needs it in a DM too. `SessionSource.is_bot` is only as good as its producers. `build_source` defaults it to False and 3 of 32 adapter call sites pass it, so most platforms still report every author as human. Populating the rest is follow-up work; nothing here depends on the flag being right yet. |
||
|
|
feb03196a2 | fix(agent): isolate detached forks from lifecycle hooks | ||
|
|
91433c8466 |
fix(loop): the turn-boundary export skips preflight-timeout envelopes and stops re-anchoring the persist index
Follow-up to #106312. _preflight_timeout_result carries the prior history without this turn's user row (#7100); with a repeated prompt ("continue") the verbatim scan resolved to the historical copy and exported it as this turn's proven boundary — the exact relabeling the export exists to prevent. Nothing is exported for that envelope now. The trailing `agent._persist_user_message_idx = idx` ran after finalize_turn had already flushed the transcript, so it never influenced a persist and the next turn reset it: dead state, removed. |
||
|
|
37f42713ef |
feat(loop): export {turn_id, current_turn_user_idx} on every result envelope
Hosts that settle their own transcript by index (hermes-webui) cannot prove which row of result["messages"] is the current user turn once this loop rewrote history (alternation repair, compaction, post-turn micro-compaction): the instance-side _persist_user_message_idx predates those rewrites, and a text match relabels an identical historical prompt and claims its old answer. Only the producer can assert the coordinate against the exact list it returns. run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair through export_current_turn_boundary on every envelope that leaves the loop (success, partial/error, interrupt, retry-exhausted, tool-limit, preflight timeout, codex runtime), computed on the final messages after finalize_turn and micro-compaction. The pair is exported only when the addressed row is this turn's user message verbatim (reanchor's last-match rule); a rewritten row exports nothing so hosts fail closed. The final index is mirrored into _persist_user_message_idx for the persist override. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh |
||
|
|
e333113871 |
fix(review): keep /refine under the background_review origin; attendedness is its own flag
The salvaged commit forked an explicit /refine under a new "refine_review" origin so the memory delete gate would not treat it as unattended. But is_background_review() is the key for every other review guard — skill_manager_guards (curator-owned-only, read-before-write), skill_manager_tool (archive instead of rmtree), skill_ledger actor, write_approval staging, the [auto] tag — so a /refine fork silently escaped all of them. Carry attendedness separately: the fork keeps origin "background_review" and sets _review_attended; turn_context binds it beside the origin ContextVar; the memory gate keys on the new is_unattended_review(). Also run the gate AFTER _validate_single_op / the operations list check, as memory_tool's own docstring requires, so an invalid replace is rejected now rather than staged and failed at approve time. |
||
|
|
63c6f9bf14 |
simplify(agent): sidecar backfill — drop the hasattr guard and the duplicated row-id predicate; tests 7→6
_session_db is always a SessionDB (agent_init / delegate_tool), so the "fail closed on a store wrapper" hasattr was defense around code that cannot fail; the store's own guard binds the value into SQL, so the prologue only needs the sibling idiom isinstance(_row_id, int) that session_persistence and transcript_repair already use. The positional hazard is explained once, on set_latest_user_api_content. The in-place compaction test duplicated test_api_content_sidecar's test_inplace_compaction_backfills_sidecar_into_db verbatim (its row_id parameter was never varied); dropped, as was the positional-helper tail of test_older_identical_row_is_untouched already covered there. |
||
|
|
73e3547ffd |
refactor(agent): one durable-row rule for the flush and the sidecar stamp; trim tests
The turn-start stamp had grown its own copy of the "what does the current user row hold" rule (persist override = clean transcript, live content = wire bytes = sidecar when they differ) that _db_flush_row already implements. Two copies drift; extract durable_user_row_content() in session_persistence and call it from both. Also: reuse _persist_lock() instead of a third open-coded lock/nullcontext ladder; drop the hasattr guard on set_latest_user_api_content (it predates this fix and exists on every SessionDB); cut the comment to the WHY; trim the new test file from 18 cases to the 7 invariants (real close flush E2E, repeated-"ok" positional protection, API-only pre-flushed turn, normal path writes nothing, compaction keeps positional, store guards). Still 3 red / 4 green when agent/turn_context.py is swapped for main's copy. |
||
|
|
4126b144bb |
fix(agent): read the sidecar row id under the session persist lock
_stamp_api_content_sidecar read _row_id without holding _session_persist_lock. A close/early flush holds that lock while it commits the row and only afterwards writes _row_id back onto the live dict; a stamp that ran in between saw no id and skipped the backfill, the flush finished with api_content = NULL and marked the message persisted, and the turn-start persist skipped it — the row kept the wrong bytes with no writer left to fix it. Run the _row_id read and the DB backfill under the (re-entrant) lock, re-checking _row_id after acquiring it. Race reported by @ehz0ah on Co-authored-by: sal <141555468+salch-cred@users.noreply.github.com> #102411; same fix shape as @salch-cred's follow-up on #103721. |
||
|
|
bc16c32c05 |
fix(agent): row-addressed api_content backfill for pre-persisted user turns (#102194)
The api_content sidecar ('persist what you send') preserves prompt-cache
stability across turn boundaries by persisting the exact API-bound bytes
(including memory-manager prefetch, plugin injections, and API-only notes)
and substituting them on replay.
When a user turn was already materialized in the database before the
sidecar could be composed (in-place preflight compaction or a close/early
flush racing the prologue on the CLI path), the turn-start crash persist
marker-skips that message. Previously, the backfill was gated strictly on
in-place compaction (_preflight_compressed and _last_compaction_in_place),
so racing CLI flushes left api_content = NULL in SQLite and broke prompt
caching on subsequent turns (#102194).
Positional approaches (such as #102239 and #102286) using LIMIT 1 on the
newest active user row are unsafe: repeated common inputs ('ok', 'yes',
'continue') cause the backfill to match and overwrite the PREVIOUS turn's
row with the new turn's sidecar, corrupting history and breaking cache parity.
Resolve all landing blockers and review feedback from #102411:
1. Bounded state owner (Sahilvishnaliya):
Add SessionDB.set_message_api_content(session_id, row_id, content, api_content)
to SessionMessagesMixin in hermes_state_messages.py instead of growing
hermes_state.py. Update set_latest_user_api_content docstring with durable
warning on the positional hazard.
2. API-only turns & durable content selection (ehz0ah):
When a pre-flushed clean input has an API-only difference (e.g. voice
prefix or model-switch note):
- Retain the differing API-facing bytes as api_content even when no
new memory or plugin context was injected.
- Derive the durable content guard using _override_replaces_content so
the SQL 'content IS ?' guard matches the clean override text stored
in the DB row rather than the restored wire text.
3. Turn prologue gating (_row_id) & fail-closed store duck-typing (ehz0ah):
In agent/turn_context.py::_stamp_api_content_sidecar: check _row_id on
the live user dict (stamped by _insert_message_rows and synced by
sync_flushed_message_markers). If valid (positive int, not bool), address
by exact ID. Do NOT fall back to positional matching when a row ID is
present: if an external or custom wrapper lacks set_message_api_content,
fail closed and skip rather than corrupting a neighbouring row. If absent
but in-place compacted, fall back to positional update. On normal turns,
skip the backfill entirely (single atomic INSERT).
4. Real lifecycle test coverage (salch-cred, ehz0ah):
Comprehensive tests in tests/agent/test_api_content_row_addressed_backfill.py
covering store guards, surrogate scrubbing, gate non-arming, older identical
row protection, real close-flush row_id synchronization, API-only clean
override preservation with exact wire replay, and duck-typed store fail-closed
verification when set_message_api_content is absent.
Fixes #102194.
Closes #102411.
|
||
|
|
defdf64790 |
simplify(agent): surface switch — reuse flatten_message_text / agent_tool_names / one runtime-boundary split
- _transcript_row_texts re-implemented agent.message_content.flatten_message_text and the api_content sidecar rule; the note can only land on a user row, so the transcript scan now skips assistant/tool rows (the bulk of the bytes). - Three sites computed "names of agent.tools"; tools.mcp_tool_agent gains agent_tool_names() used by the switch note and conversation_loop, which also stops importing the private _def_name across modules. The name list is only captured when a switch was announced. - split_runtime_boundary() is the single owner of the runtime-block rpartition/END check for both identity_line_value and _stored_prompt_matches_runtime. - platform_surface_hint was a public alias of _platform_hint; the function is now platform_hint (its docstring pointed at the pre-move module). - consume_gateway_turn_context_notes and consume_surface_switch_note share _pop_turn_note so the two one-shot channels have identical semantics. - platform check hoisted above the transcript scan. |
||
|
|
6cc177a76c |
refactor(agent): surface-switch note lives in its own sibling; skip it where no sidecar exists
Move the six surface-switch helpers out of the conversation_loop facade into agent/surface_switch.py (AGENTS.md: new behaviour goes in a topical sibling), and fold the review findings on #104494: - MoA and codex_app_server turns never stamp the api_content sidecar, so the staged note could not be read back from the transcript and was re-sent on every turn after a switch. Those modes now skip the note (stored prompt still reused). - The announced surface was parsed with split(".") — a plugin platform with a dot in its name would never compare equal and re-stage the note every turn. The note now closes the name with a fixed terminator. - One identity-line parser (identity_line_value) shared by _stored_prompt_matches_runtime and the switch detector instead of two copies of the runtime-boundary/rpartition logic; tool names via the existing tools.mcp_tool_agent._def_name; the transcript scan is bounded to the last 200 rows (it ran every turn over the whole history). - consume_surface_switch_note reduced to a plain pop; developer-guide prompt-assembly.md updated (Platform is no longer an identity field); 17 new tests trimmed to 10 (same-shape pin/retire variants folded). Restoring Platform as an identity field still turns 5 tests red. |
||
|
|
80d6bda144 |
fix(agent): a surface switch must not re-prefill the whole request (#104414)
`_stored_prompt_matches_runtime` treated `Platform` as a runtime-identity field, so answering a live session from another surface — desktop -> TUI, or a resume after a dashboard restart whose chat is a PTY TUI child — declared the stored prompt stale and rebuilt it. The system prompt is the first thing in the request, so changing any byte of it moves the first divergent byte to the head of a 220K-token request and the entire conversation behind it re-prefills: a session that was hitting 240000/240287 came back at 1536/219861. The guard was not wrong about correctness — a desktop-built prompt on a terminal session advertises inline widgets and a MEDIA: channel the TUI does not have — but the surface is advisory metadata about the renderer, not a cache domain. Model/provider and cwd drift change what the prompt should SAY; the surface changes only one paragraph. Reuse the stored bytes across a surface switch and correct the paragraph where it costs nothing to cache: `_stage_surface_switch_note` stages a one-shot note carrying the CURRENT surface's guidance on the same per-turn user-message channel the gateway's must-deliver notes use. It lands after the cached prefix and is stamped into the byte-stable `api_content` sidecar, so later turns replay it instead of re-prefilling, and the prompt converges at the next compaction — a boundary that already breaks the cache. The saved tool_names prefix is not pinned across a switch: the tool registry is process-global, so `_merge_preserving_prefix` would carry a saved-but-unloaded tool forward (under `coding_context: focus` desktop gets a desktop_ui toolset the TUI cannot run). On the same surface the tools freeze is untouched. |
||
|
|
3114916ee4 |
fix(gateway): carry accepted-input ownership through persistence
Namespace delivery markers and assign fresh keyless turn identities instead of inferring ownership from IDs or process-local row baselines. Query only marker existence on the canonical live compression continuation and ancestors. Preserve raw reply IDs and exclude metadata from provider wire messages. Expand the two existing invariants with resumed cross-chat ID collisions, a real independent SQLite writer, reaped siblings, and archived-history allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass. |
||
|
|
93af3db01d |
fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the hard iteration cap, without accepting arbitrary diffs or weakening failure counting. Add opt-in per-turn iteration checkpoints for ordinary agents. Persist checkpoint text with the fresh tool result, never rewrite cached rows. Salvages the opt-in ratio and per-turn reset implementation from #104683; credits the earlier default-off signpost proposal in #92438. Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck workers still reach blocked/2 after two runs. Default-off control unchanged. Targeted and affected-directory suites queued behind campaign test lock. Co-authored-by: fangliquanflq <fangliquan@qq.com> Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com> |
||
|
|
dca7a90cf8 |
fix(agent): reclaim background processes by execution owner
Track raw task identities across an agent's turns and match them against process owner_task_id during close. Session IDs and shared terminal keys are not process ownership, so the old bulk cleanup missed delegated work. Preserve parent/sibling processes and consume teardown notifications. Move task-resource cleanup into the lifecycle mixin, add real-process isolation regressions, and document background process lifetime. |
||
|
|
be58c276ee |
feat(compression): per-image token cost learned from the provider's own usage (#70328, supersedes #70463)
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K, so compaction never fired and the provider rejected every request (#70328). The provider prices every image exactly on the request that carries it, so the cost is observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual between the next real prompt_tokens and anchor + text-only delta is the price of the N images that delta introduced. - agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in ~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar. - estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all read the same bound value, so trigger and walk agree; the per-message memo now caches text tokens and image COUNT so a recalibration re-prices cached rows. - One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone. evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn and the walk's error is +8.5%. Reporter and first-fix credit: @JonthanaHanh (#70328, #70463). |
||
|
|
d932fa5929 | fix(memory): spill oversized external prefetch | ||
|
|
0f4587e336 |
refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the compressor's rough/real projection (should_defer_preflight_to_real_usage with last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391). Now there is one authority: - Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw last_prompt_tokens ignored the tool results just appended), then real, then rough. - Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on the session row, else rough. - Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate over threshold waits ONE request for the provider's real count instead of compressing on a guess (first request, rewind/edit-resend, reloaded history without a persisted anchor). - The wait is one request, never a disable: a provider that omits usage (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure past the whole window, and provider-proven overflow all compress immediately; the post-compaction latch (#36718 / #104192) is unchanged. - Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures that scripted whole-history estimates now state the fact they relied on (provider omits usage). Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists). |
||
|
|
c0aaa238f6 |
feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since) identified the priced transcript by id() of the last message, so it was None on EVERY gateway turn (history is re-read from the DB each turn) and in every fresh process (--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4 estimate then fired local compression against payloads the provider priced far under threshold (#99421, #104462). - agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first resumed turn while the durable transcript still matches, cleared with the row on compaction / codex-native rewrite / session reset. - Callers repointed from model_metadata (the compat table follows). Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026 layout (the branch predates the model_metadata / agent_init split). |
||
|
|
cd71ee0708 |
fix(compression): defer local preflight after native checkpoint
A native Responses compaction checkpoint is opaque ciphertext; the rough preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against a 204K trigger) and fires local compression on a request whose real prompt is ~116K. Arm the existing one-response real-usage latch when a replayable checkpoint is captured (build_assistant_message) or restored into a fresh agent (_hydrate_from_history), honor it in the post-tool gate and idle compaction, and require non-empty encrypted_content for a checkpoint. Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365, bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by patch application. Source delta is byte-identical to the PR head d6ce3e236d. Fixes #100611 |
||
|
|
d63e380324 |
compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which
|
||
|
|
2776813df3 |
compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in the focused modules that define them. This commit is the ONLY thing keeping the old paths alive, so external plugins have time to update. It is deliberately a single, unsquashed commit: git revert <this sha> removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does. What it adds (see COMPAT_MANIFEST.md, compat_manifest.json): - 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file - 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import, so no import cycles; facades that already had `__getattr__` get a chained one - 592 third-party/stdlib names the old modules used to expose, with their original import statements - 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them) - 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/ relay_runtime, tools/environments/modal_utils) - private names (`_x`) get no pointer: they were never API (3,792 skipped) Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves; the lint reports zero in-tree uses; ruff clean; targeted suites unchanged. |
||
|
|
71e0f64679 | simplify(compat): tools/mcp_tool — repoint agent/turn_context between-turns refresh import (missed hunk) | ||
|
|
fcbe4acbef | simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings |