What: the metrics lineage no longer keys off the ambient Portal conversation id (it walks every
parent_session_id: gateway reset/idle expiry, /new, /branch). Compression calls
relay_shared_metrics.rotate_segment(old, new) from the committed rotation
(_notify_context_engine_compression_complete) and from a stale agent adopting the live tip
(_adopt_live_compression_child). The hand-off registers the new id in the lineage, moves the
route run and the spent-turn history into it, and closes the old segment now (or when its in-flight
compression turn ends). A rotated-to id closed before serving a turn, or still open at shutdown,
closes too; a 512-lineage backstop flushes conversations a surface never closed.
- m2: /undo or /retry on the new id reaches the turns earlier segments spent (tokens_bucket known).
- m11: /model right after a rotation finds the conversation's turn run.
- m8: background review forks (same session id, no pre_llm_call) add no turns/calls/messages.
Why: one conversation = one session / context_peak / tool_overhead / tool_enabled_unused row, on
time; N gateway conversations collapsed into one row emitted only at process exit, and rotated
segments stayed in memory until exit.
Tests: two existing tests simulated rotation by sharing a conversation context; their contract
changes to the explicit hand-off (test_a_compressed_conversation..., test_context_peak_is_one_row...).
New: reset/branch chain after a compression + tip-first close order, the compression seam, switch/undo
right after a rotation, review fork volume. RED on base (rotate_segment missing / switch_after [] /
turn bucket '2' for one user turn).
Probe (review/engagement/lineage probe_lineage.py -k c2, probe_branch.py; fix-lineage/probe_handoff.py):
before: c2 3 conversations -> session.count 0 until shutdown, then 1; branch -> 1; 50 compressed
conversations -> 0 rows, 50 lineages/sessions live.
after: c2 -> 3 rows before shutdown (turns 2,1,1), lineages=0; branch -> 2; 50 compressed -> 50 rows,
0 lineages, no sessions; tip-first close with the old turn running -> 0 then 1; shutdown with an
open tip -> 1 per conversation.
Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):
- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
A Hermes-known break suppresses the miss it causes. No prompt hashing.
Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):
- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
matcher reports the strategy that landed (or no_match / ambiguous) into a
context-local probe opened only around the patch/write_file handlers, so we
learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
the guardrail already warns/blocks/halts and at turn end for the iteration
budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
next_outcome. One row per failed tool call, resolved against the model's
next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
a table lookup on the first program word, never the text. Hermes' own
deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
backend result, so a command's own `exit 124` reads nonzero, not timeout.
Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
truncation only from the structured finish reason; empty/reasoning_only
from the normalized message; issue=none per response as the denominator
(model_route counts attempts and files empty replies as failures).
Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
Review findings M7, M8, m6, m7 and the browser half of m8 on the v4 loop metrics.
- M7: TUI/Desktop @-path completion on a non-local backend lists the directory through
terminal_tool, so every keystroke added a hermes.execution_backend.count row as if the user
ran a command. Hermes-owned calls now run inside shared_metrics_loop.unmetered_backend_calls()
(a contextvar checked in record_execution_backend) instead of a flag threaded through
terminal_tool. It was the only non-_host_local internal terminal_tool caller (bot DM runners
already use _host_local; the prompt-builder probe calls env.execute directly).
- M8: a scheduled curator tick that finds the run claim held recorded scheduled/skipped every
tick (5 ticks -> 5 rows for one pass). The holder's own pass is the one counted; the held
branch now records nothing. Real dry runs still report skipped.
- m6: a foreground timeout returns exit 124 with partial output and no error field, so it was
recorded as success. exit 124 now records outcome=failed, error_class=timeout.
- m7: vercel_sandbox (and managed_modal) are built-in backends in
tools/terminal_tool_config._BUILTIN_BACKENDS but collapsed to other. Added to
TERMINAL_BACKENDS and both schema enums; a test pins every built-in backend to its own bucket.
- m8 (browser): record_browser_call resolved the browser backend on every call even with
collection off. The resolver is now passed lazily and only runs inside the gated builder.
The existing test asserting a held claim records scheduled/skipped encoded M8 itself and is
replaced by one asserting held ticks record nothing.
A crash or SIGKILL cannot record itself. CLI, TUI, gateway, serve and cron-tick
drop a small marker under the profile's store dir at start (only when
collection is on), restamp it clean at exit, crash+exception family from the
excepthook, or watchdog before the gateway's os._exit. The next start in that
profile reports every marker whose owner fails the canonical start-time
liveness check (running => killed), on a daemon thread with rename-claims so
startup never waits and concurrent starts never double count.
Turns killed by the turn liveness watchdog or the gateway inactivity watchdog
count as exit_kind=watchdog once per turn, in the turn's own profile (home
captured at turn entry; off-thread watchdogs bind it explicitly).
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.
hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.
hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.
hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
- memory tool: one row per operation (a batch counts each op); gate/validation
refusals are "rejected", store-level non-application "failed". Memory provider
tools are counted in MemoryManager.handle_tool_call with the op read from the
action arg or tool-name verb.
- curator: run_curator_review records one row per pass (dry run = skipped) with
the before/after diff bucketed; a scheduled pass whose claim is held by another
process counts as skipped. The home is captured before the review thread starts.
- delegate_task: _run_batch opens the call, each joined unit folds its results in,
and the last unit emits the call's single row, so group-split background calls
are not counted per unit.
- terminal / execute_code / browser: counted per call that reached the backend;
the backend is resolved from the owning profile at call time (terminal plan
env_type, code local/remote, browser CDP > Camofox > cloud provider > engine, or
"extension" when the extension controller served the call). Guard refusals and
Hermes' own _host_local commands are not counted.
Tests read rows back from the real store; each call-site group is red with its
source reverted.
Independent review of the v3 metrics found:
- Provider fields were shape-checked only, so `custom:<config key>` (and
any unshipped provider id) reached setup.completed, install.snapshot,
model_switch, fallback, model_tokens and model_route. Providers now pass
only when Hermes ships them (auth registry, overlays, model catalog,
aliases, cached models.dev ids); everything else reads `custom`. A custom
or loopback provider's model id reads `custom`, as does any model id that
looks like a path or URL. Two existing tests asserted the old export of a
custom endpoint's model id and an unshipped provider; they now assert the
collapse, and the smoke run treats the model canary as prohibited.
- Install milestones read the install age on the Relay thread, which has no
profile binding, so a multiplexed profile got the launch profile's age.
The subscriber captures its home at construction.
- Gateway sessions retired by the store (auto-reset, /new, resume) now close
their metrics session at the route transition instead of at shutdown.
- The dashboard MCP add records its install after leaving the config lock.
- Compression and gateway slash counting helpers move into
shared_metrics_events so the >2k-line files barely grow.
The pending shared-metrics record lived on the agent, so a stalled attempt
whose worker unwound after its stall-fallback retry had begun could pop the
retry's record: the stall counted as failed and the retry's success was
lost. An attempt begins and emits on one thread (pool worker or caller), so
the record is now thread-local; overlapping attempts can no longer consume
each other's count.
try_activate_fallback is the one place a fallback provider is bound after a
classified provider error (every turn_* recovery phase calls it). Right after
the activation is logged it now records hermes.fallback.count with the
provider it left, the provider it bound and the classifier's FailoverReason.
Candidates skipped along the chain (unconfigured, raising) are not
activations and are not counted.
Every compression attempt already funnels its outcome through
_emit_compression_attempt_telemetry (committed, aborted, lock-contended,
pool-saturated). That choke point now also records hermes.compression.count
once per attempt: trigger (auto / overflow / manual), outcome
(success / skipped for lock contention / failed) and the context fill from
the pre-compaction token estimate against the model window.
The pending record is set on the agent in _begin_compression_attempt and
popped by the first emit, because abort paths restore the compressor
snapshot (telemetry seed included) before they emit. compress_context and
AIAgent._compress_context gain a trigger kwarg so provider-proven overflow
recovery (turn_overflow, long-context-tier recovery) is labelled overflow
instead of being inferred from bypass_cooldown, which the stall-fallback
retry also sets.
setup_logging(hermes_home=get_hermes_home()) is the whole fix; the FileNotFoundError fallback chain guarded a home that is being served (it exists), and the two extra tests re-tested it. Keep the A->B routing invariant and the no-bare-handler guard.
A Desktop serve backend builds agents for several profiles inside
set_hermes_home_override(); _setup_logging passed run_agent's
import-time _hermes_home freeze, so setup_logging saw a home it already
served, never adopted the profile, and every profile's records landed in
the launch profile's agent.log. Pass the active-scope home instead.
A home that cannot host logs (missing/tombstoned named profile —
mkdir_under_hermes_home refuses to materialize it on purpose) falls back
to the process launch home, then disables file logging with a stderr
notice: logging setup must never take agent construction down.
(cherry picked from commit 76bcb71bae400e764fb0e04c28e610a3d1b5d4ca)
A compression attempt that ends by host timeout, cooldown/backoff block, or gateway
hygiene turn-hold writes its terminal activity stamp through _touch_activity without
force_persist. The durable SessionDB activity projection is rate-limited to one write
per SESSION_ACTIVITY_HEARTBEAT_MIN_INTERVAL_SECONDS (60s), and the compression
heartbeat had just written 'context compression in progress' — so the terminal stamp
almost always lands inside the open window and is dropped.
Nothing writes after it: the turn is over, the worker is detached, the host may be
gone. sessions.last_activity_description stays 'context compression in progress' with
last_activity_provenance 'agent.compression' forever, next to a
compression_failure_error and a cooldown timestamp that prove the attempt terminated.
Every surface reading the durable row (chat lists, session listings, reconnect,
orphan reap, stall watchdog) shows a permanently stuck/compressing chat with no turn
running.
Fix at the single write path: _touch_activity bypasses the persist rate limit for any
terminal compression provenance, exactly as force_persist does. The rate limit itself
is unchanged for mid-compression heartbeats and every other writer, so no extra write
pressure is added to the contended SessionDB path — one extra write per terminated
compression attempt.
TERMINAL_COMPRESSION_PROVENANCES is deliberately separate from
conversation_compression._TERMINAL_COMPRESSION_PROVENANCES: that set answers 'may a
detached heartbeat overwrite this stamp?' and excludes TURNHOLD because the worker may
still be alive and adoptable. This set answers 'must this reach durable state now?',
which is true for turn-hold too.
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
The profile line names the home path, so every home/profile had a
different stable prefix and paid a full cache write (~20k tokens on
Opus) for tools + stable system on its first request. With the line in
the volatile tail the stable prefix is byte-identical across profiles
and homes on a host, so a multiplexed gateway serving N profiles shares
one cached prefix.
Codex webSearch items carry id/query/action; defaulting a missing status
to "completed" wrote a value codex never reported. Keep the provenance
marker, add status only when present, match sibling specs' ensure_ascii,
list webSearch in the module docstring, and pin args parity with the live
bridge (_codex_item_to_args) instead of a literal.
Codex app-server webSearch items fire the live tool bubble under
codex_web_search_<itemId>, but the projector had no entry for them, so the
stored transcript kept an opaque "[codex webSearch] {json}" note instead.
Any transcript refresh or reopen then replaced the search card with raw
JSON.
The projector now records webSearch as a web_search tool_call under the
same id and name the live bubble uses. The tool result is
{"provider": "codex", "status": ...}, so a later non-Codex turn can tell
Codex ran the search.
Salvaged from #66995: only the webSearch projection is kept, reworked
onto the _TOOL_PROJECTIONS table and the live web_search name. The MCP
elicitation trust change is not included.
_session_start_like localised the naive session-id stamp with
datetime.now().astimezone().tzinfo, which is a fixed offset for today.
A stamp from the other DST half got the wrong offset: under
Europe/London on a summer day, session 20260115_003000_* rendered as
"Conversation started: Wednesday, January 14, 2026". dt.astimezone()
applies the offset in force at the stamp itself, as
gateway/message_timestamps.py already does.
Also replace the two remaining datetime.utcfromtimestamp() calls
(deprecated since Python 3.12) with fromtimestamp(..., timezone.utc).
The output strings are unchanged. tui_gateway/methods_session.py runs
with tui_gateway/server.py's globals, so server.py now imports timezone.
Trim the salvaged detector/delivery pair to the salvage bar and fix the
delivery half. The PR resolved the refresh selection as
platform_toolsets.<surface> (desktop/tui), a key no config path writes,
so _get_platform_tools fell back to the constructed hermes-<surface>
composite that resolves to 0 tools (a cold resume went 30 -> 1 tool).
Delivery now goes through the builder new desktop/TUI sessions use
(tui_gateway.server._load_enabled_toolsets(platform) +
_load_disabled_toolsets) as an explicit refresh_agent_mcp_tools
override, so the refreshed set equals a fresh session's on that surface;
only desktop/tui agents are long-lived, every other surface builds a
fresh agent per process/turn and is left alone. Drops the
toolsets.py/tui_gateway policy relocation and the raw-value fallback in
the fingerprint; keeps the detector (platform_toolsets +
agent.disabled_toolsets replace the dead tools.enabled_toolsets key)
and the tools re-pin after the refresh.
Tests: 2 invariants in tests/agent/test_bot_chat_toolset_refresh.py
(desktop -> builder consulted, disabled override passed, re-pinned;
cli -> untouched) and the existing fingerprint axis test now edits
the key `hermes tools enable/disable` writes.
The capability epoch watched tools.enabled_toolsets, a key no surface
writes; real hermes tools enable/disable edits (platform_toolsets.* plus
agent.disabled_toolsets) never flipped it. Even when stale, the refresh
rebuilt only the prompt and never tools[]. Watch the real keys and
rebuild + re-pin the tool snapshot on Bot Chat capability refresh.
Closes#124211
ReviewIdleQueue stored (agent, session_key, kwargs, enqueued_at) with no
context, and the shared dispatcher thread called _still_enabled(item) and
item.agent._spawn_background_review_now(**item.kwargs) under its own
AMBIENT context. On a multiplexed gateway that meant the wrong profile's
background_review.enabled gate decided whether a queued review ran, and the
spawned worker inherited the ambient home / no secret scope (the
propagate_context_to_thread in _spawn_background_review_now copies a
context that no longer carries the originating profile).
Capture contextvars.copy_context() at enqueue() on _PendingReview and run
both the enabled re-check and the spawn via item.context.run(...).
Test: two profile homes (enabled / disabled) plus a disabled ambient home
under set_multiplex_active(True); only the enabled profile's item spawns,
and it observes its own home override + secret scope.
Salvaged from #108538 by @Liuzikaii (mechanism kept, test trimmed into the
existing test module).
Fixes#108537
(cherry picked from commit 4f8a64fd680e40de270a18b52c8f44a414334c56)
On a multiplexed gateway the curator tick runs inside profile_scoped_chore(),
which installs the profile home override and secret scope as contextvars.
run_curator_review() started _llm_pass on a bare threading.Thread, which
begins with an EMPTY context: _resolve_review_provider hit
UnscopedSecretError and every home lookup (skill snapshot, run.json /
REPORT.md, .curator_state) fell back to the process home, so the ROOT
home's skill library was read, reported and overwritten under another
profile's run.
Copy the caller's context into the thread (copy_context().run), the same
shape the gateway uses to carry profile scope into executor work.
Test: two temp homes under set_multiplex_active(True); the review thread
must observe the caller's home override + secret scope and write the
profile's .curator_state, never the root home's.
Salvaged from #125038 by @dskwe (mechanism kept, test trimmed).
Fixes#125032
(cherry picked from commit 3dbc99944d751ae0ab1c146e1fce68fefb48a22d)
After a contended wait with an unknown row state (get_session raised),
the carried input from an interrupted earlier wait was appended to the
reload before the "keep the caller's history when nothing reloaded"
check. An empty reload therefore looked non-empty, and the caller's
history was replaced by the carried message alone. The decision now
looks at the stored rows only, and the carried input is appended only
when the reload is used.
Salvage of #125125.
After the turn lease is admitted, a failed get_session was reported as 'row exists': the agent recorded the row as created and, after a wait, an empty reload replaced the caller's history. The read now has three answers. An unknown row leaves the row flag alone (the create is an upsert that never overwrites an existing row) and keeps the caller's history when the reload has nothing to offer.
(cherry picked from commit f06c354bc120d687df14cc48b6f0a136cf495f27)
admit_durable_turn_lease checked whether the session row exists before it
waited for the lease. The holder it waits for can create or delete that row,
so the answer could be stale in both directions:
- row created during the wait: the transcript was reloaded, but
_session_db_created stayed False and the first flush ran a redundant
create_session upsert;
- row deleted during the wait: the reload of the absent row replaced the
caller's history with an empty one.
The row is now read once, after admission. A missing row leaves
_session_db_created alone, so the flush-time heal for a row deleted under a
live agent (#123583) still replays the full in-memory transcript.
The threaded first-turn test no longer depends on scheduling: the second
writer's wait notice gates the first turn's finish, and a_finished is set
inside the turn, before the lease is released.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
(cherry picked from commit 29625ef05277be8eee6dcc8225e30c6b1ddb4459)
admit_durable_turn_lease skipped the lease when the session row did not exist
yet, assuming a fresh id is process-unique. Client-addressed ids are not: an
API-server X-Hermes-Session-Id, a /v1/runs session_id or a fingerprint-derived
chat id. The first turn writes its row mid-turn, so a second request on the
same id found the row, took the unheld lease at once and ran beside the first:
its model call merged the other turn's question into its own, and the stored
transcript ended user, user, assistant, assistant, with one answer dropped from
later context.
The lease is now taken whether or not the row exists. After a contended wait
the transcript is reloaded only when a row exists, so an in-memory seed on a
fresh id still survives.
(cherry picked from commit 80c2a4749c9ab0a3eedf729111406a5c94c6325b)
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".
- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
installed copy or None. Every caller already pm.ensure()s on None, so a
missing runtime is now provisioned instead of silently borrowing the
user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
another npm.
- source_build.source_product_current: run the freshness reader only with
PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
artifact; delete the "uv on PATH if new enough" developer shortcut.
Gate r2 Low cleanups (house rule: no aliases/shims):
- Drop the is_live_database_file alias; its point-in-time caveat now lives on
has_live_connection.
- _refuse_live_database reuses offline_file_access's message (via _serve_offline),
so a download 409 on state.db-shm names the main database like the read path;
the verb is "serve" so it fits read/download/stream.
- /api/files/read reads whole files in-process, so _read_base64_file now holds
offline_file_access through close (409 on a live DB; OSError stays 500). Only
the streamed FileResponse routes keep the point-in-time check.
- That check takes the global _live_lock, which other threads hold across
whole-file reads, so fs_download and the managed stream routes run it via
asyncio.to_thread instead of stalling the event loop.
- _managed_readable_file docstring no longer claims a size cap;
_read_file_reference returns (early, text) instead of a str|Expansion union
sniffed with isinstance.
Co-authored-by: Benjamin PERRY <benjaminperry6@yahoo.fr>
FileResponse opens and closes the file in the dashboard process, so
downloading a live state.db (or its -shm/-wal) via /api/fs/download or
the managed-file read/download/media routes still cancelled the
connection's POSIX locks. Both now return 409 via is_live_database_file;
the registry lock is not held across the streamed response.
The main-or-WAL-sidecar rule now lives in one _live_main_key helper used
by offline_file_access, has_live_connection and read_header_bytes_preopen,
and the sidecar refusal names the main database the connection is open on.
@file previews hold _live_lock only for the raw read; token counting and
formatting run after release. The Linux lock test gains requires_wal
(Hermes uses DELETE mode on WAL-reset-vulnerable SQLite), covers the
download refusal, and drops an ambiguous conditional assert.
Co-authored-by: Benjamin PERRY <benjaminperry6@yahoo.fr>
_insert_message_rows stamps _row_id and the stored-row digest onto the
caller's dicts inside the write transaction. Only append_messages_batch
restored that state on rollback. archive_and_compact, replace_messages
and the rotation handoff left the rolled-back id + digest on the dicts;
SQLite reuses the id, so a later flush found a digest mismatch on the
foreign row, adopted it and silently dropped the user's message.
Move the capture/restore into _execute_transcript_write, used by every
caller that inserts caller-owned dicts: each attempt starts from the
caller's state and a final failure restores it before re-raising.
(Rewind replacement and import insert dicts built inside the txn.)
Also: bind _message_row_params directly on insert instead of the
serialized-dict round-trip, import the public DB_ROW_SNAPSHOT /
CANONICAL_ROW names, set adopt=False once, and reuse target_row
instead of re-SELECTing when nothing was written.
A legacy (no-digest) dict over a non-blank assistant row adopted the whole
decoded DB row: tool_calls / reasoning* / codex_* were overwritten with the
stored JSON (which still holds the escaped lone surrogate the sanitizer just
fixed, re-injecting it into the provider payload) and live-only fields were
popped. Resumed dicts (_rows_to_conversation stamps _row_id without a
digest) and compaction clones hit this path. Adopt content only, as before
this stack, via a content-only canonical handled like the metadata-only one.
_insert_message_rows dropped a clone's parent digest but only the flush
path restamped it, so clones made by archive_and_compact / replace /
rotation handoff / import reached the legacy path and the first live edit
after a clone was not persisted. Stamp the stored-row digest inside
_insert_message_rows (one batched SELECT, cold paths only; the flush path
statement count is unchanged) and drop the duplicate call in
append_messages_batch.
Define the _db_row_snapshot / _canonical_row keys once in
agent/message_metadata.py and import them everywhere instead of repeating
the literals.
The row digest hashed every repair column, so a same-process metadata write
(reaction, display-kind stamp, api_content / codex reasoning backfill,
platform message id) made our own row look like a foreign winner. The
re-flush then adopted the stale DB row: a later live edit (the non-ASCII
strip recovery) was reverted, and unsanitized tool_calls/reasoning were
copied back onto the live dict.
The digest now covers only the owned (non-metadata) columns: it means "the
row is still what we last committed". Match -> write the live owned values
and hand over only presentation metadata the live dict lacks; mismatch ->
genuine other writer, adopt as before. The r3 "stored content equals the
durable form of live" special case is subsumed and removed.
Also:
- _insert_message_rows drops a carried digest when it assigns a new row id
(compaction/replace/import clones carried the parent's version).
- append_messages_batch restores each message's _row_id / digest /
timestamp and pops the adopted row at the top of every _execute_write
attempt, so a rolled-back attempt cannot resolve to a foreign row.
- message_id is no longer synced onto live (int -> str flip, spurious
platform_message_id).
- The JSONL divert strips both bookkeeping keys via one frozenset.
Same-process writers (set_message_reaction, display-kind stamping, api_content
backfills, codex reasoning update) change stored columns after a flush without
refreshing the live row digest. The next sanitize + re-flush treated that as a
concurrent winner and copied the lossy durable projection over live multimodal
content, dropping image parts and shifting the prompt-cache prefix. On adoption
we now keep live content when the stored content is just the durable form of
it and sync metadata only; a real concurrent content winner is still adopted.
The legacy blank-assistant path (dict with _row_id but no digest, blank DB row)
again only fills the row from live content, as on main, instead of running the
full canonical sync that wiped live reasoning_content/finish_reason/tool_calls.
Adoption on the legacy path is limited to a non-blank row.
Also: transcript_row_snapshot returns str (the partial-row branch had no
caller), serialization only runs on the digest-match branch that reads it,
stamping reuses hermes_state_common._id_chunks, and _MESSAGE_WRITE_COLUMNS is a
plain top-level import (hermes_state_messages imports this module lazily, so
there is no cycle).
Gateway/TUI/CLI callers pass their live dicts straight to
append_messages_batch, so a concurrent-winner adoption leaves the
decoded durable row (_canonical_row) on a dict that may later be sent
to the model. Treat it as persistence-only like _row_id and the digest
so the outbound builder and token estimator both drop it.
The row-addressed repair stamped the decoded durable row on every
resolved message, so after our own rewrite the sync copied the lossy
durable projection (image parts -> "text\n[screenshot]") back onto the
live dict: multimodal user/tool messages lost their images and the
prompt-cache prefix changed. Adopt the DB row only when another writer
won (digest mismatch) or on the legacy assistant path, as BASE did.
The insert-time digest hashed Python bind values, but SQLite affinity
rewrites them on storage (int message_id -> TEXT, float token_count ->
INTEGER), so live and DB digests never matched and in-place edits were
silently dropped. Hash the stored rows instead, only on the
append_messages_batch flush path that reads the digest (one SELECT per
batch), incrementally (type tag + length prefix) instead of via JSON.
Also skip the no-op UPDATE, fix the _write_columns comment/spacing and
drop the duplicate top-level Optional import (F811).
The CAS row snapshot was a full copy of each message's durable payload
riding on the live dict. The rough token estimator priced it (about 2x
estimates -> premature compaction) and it doubled transcript memory.
Replace it with a 16-byte blake2b digest of the repair columns. The
compare now runs in Python against the target row already read inside
the BEGIN IMMEDIATE transaction, followed by a plain UPDATE. Also:
- add _db_row_snapshot to PERSISTENCE_ONLY_MESSAGE_FIELDS so the
estimator and the outbound request builder both drop it
- derive _REPAIR_COLUMNS/_SYNC_FIELDS from _MESSAGE_WRITE_COLUMNS
- use hermes_state_common._placeholders
- drop the dead resume-path stamp (the SELECT has no token_count, so it
was always None) and the dead tool name assignment in
_decoded_repair_row
- keep the digest out of divert JSONL
The kept active-row test now pins estimate stability across a flush and
the survival of a concurrent writer's row. It goes red on the old
prod files and red when the digest compare is removed.
transcript_row_snapshot annotates Optional, which was only reachable via the
PLUGIN-COMPAT re-export block at the bottom of the module. Internal code must
not depend on that revert-scheduled block, so import it with the other typing names.
62ceddd342 cut raw args to HEAD+4096 before redaction to save time. The
PEM redaction pattern only matches a complete BEGIN...END block. A long key
whose END fell past the cut stayed unredacted, and once an earlier key was
redacted and the text shrank, its body landed in the 1200-char head that
goes into the persisted summary. Go back to the BASE order: redact the full
args, then apply the MAX/HEAD cut. This is a cold path (once per summarized
call per compaction), and _SUMMARY_INPUT_MAX_CHARS still bounds the prompt.
Extend the kept canonical-args test with a two-PEM input that leaks on
62ceddd342 and passes now.
Follow-ups to making tool-call args byte-exact:
- _record_compression_regions measured canonical_messages slices while
compress_start/compress_end are indices into the pruned copy that head/tail
are assembled from; measure the pruned rows actually sent, as before. This
also removes the only canonical slicing, so blank-echo classification drift
between the two copies can no longer misalign anything.
- _render_tool_call_for_summary redacted the full (now unbounded) args before
cutting to 1200 chars; cut to head+4096 first. Output unchanged for args
within that window.
- pressure_hits always equalled demoted once arg truncation left; fold it.
- Drop the fixture-only tautological assert in the guardrail test helper.
- Reword stale compress()/compression_marker docstrings that still described
canonical head/tail and compressor-written arg markers.
The arg-truncation removal deleted the only uses of the marker constants in
agent/context_compressor.py (ruff F401), and the test module had been
importing _COMPRESSION_MARKER_PREFIX through it, so test_context_compressor
line 137 raised NameError. Import it from its home, agent.compression_marker.