m22: migrations ran save_config -> record_config_saved inside surfaced
processes (`hermes config migrate`, console migrate, dashboard/TUI profile
create), so e.g. v45 appending `connections` to platform_toolsets read as a
user re-enabling a toolset. _persist_migration (the single migration write
path) now wraps save_config in hermes_applied_write(), a ContextVar the hook
honours.
m23: the old side is the raw file (`memory_enabled: ${MEM_ON}`) and the new
side the env-expanded config, so any unrelated save reported `disabled` once a
day forever. A setting whose value on either side is an unexpanded `${...}`
string is now skipped (its value is unknown to the diff).
m24: dashboard handlers call save_config inside config_write_scope
(_CONFIG_MUTATION_LOCK), and the diff (tools_config._get_platform_tools per
platform) + record ran there. record_config_saved now runs only the cheap
gate inline (surface, enabled, surface/home resolution) and hands deep copies
to a non-daemon thread bound to the owning profile (non-daemon so a CLI
command exiting right after its write still records).
Test contract change: test_record_config_saved_needs_a_surface_and_collection
and test_real_config_writes_report_the_move_away_from_default now join the
recording thread before asserting (the record is asynchronous by design).
Probes (review/signals/mine):
p2_migration.py config|dashboard
before: re_enabled toolset connections row; after: <no db>
p4_envref_configset.py
before: memory.memory_enabled disabled after an unrelated save;
after: no row (config set curator.enabled false still records)
p3_lock_v2.py (times config_transitions, try-acquires the web lock)
before (p3_lock.py): record_config_saved 0.047s/0.004s, lock held=True
after: diff 0.003-0.008s, lock held=False
Tests test_migrations_and_env_templates_are_not_user_disables and
test_diff_and_record_run_off_the_callers_lock_in_the_owning_profile are RED on
base.
- `curator` latched on the default-on scheduled pass (outcome=success, all
buckets 0), so it meant "idled through one curator interval". It now needs
trigger=manual (the user ran it).
- `skills_created` and the v4 milestone `first_skill_created` latched when
the background-review fork created a skill (provenance=agent_created).
Both now require provenance != agent_created (foreground creates are
stamped "learn" -> not agent_created, see tools/skill_usage.record_created).
- days_since_install_bucket caught only sqlite3.Error; a non-numeric
sessions.started_at raised ValueError into the subscriber on every event,
re-opened state.db each time and never latched. It now also catches
OSError/TypeError/ValueError -> `unknown`.
Probes:
probe_internal_adoption.py
before: feature_adoption curator + skills_created, milestone first_skill_created
after: no feature_adoption rows, no milestone
probe_statedb_errors.py textstamp
before: ValueError warnings, 0 rows, settled 0 of 1
after: 1 state.db read, feature_adoption {unknown, projects}, settled 1 of 1
Tests test_hermes_internal_work_never_latches_adoption and
test_unreadable_first_session_reads_unknown are RED on base.
tool_unavailable_fields compared against agent.valid_tool_names, the
model-visible array after tool-search assembly. With the default
tools.tool_search.enabled=auto, 14 enabled built-ins (session_search,
todo_list, process_manage, ...) live behind the tool_call bridge, so a direct
call to one was reported as "disabled in this session". Cron agents had
clarify stripped on purpose by the scheduler and still counted.
- Names in the session's tool-search scoped set
(agent.tool_executor._tool_search_scoped_names: the deferrable names its
enabled/disabled toolsets allow, cached on the agent) are enabled, not
unavailable. A session that turned the toolset off still reports it.
- platform == "cron" returns None, like delegated children and background
review.
Probes:
probe_deferred_tools.py (real AIAgent, default config)
before: rows for session_search and todo_list; after: no rows
probe_cron_narrowed.py
before: tool_name=clarify row; after: no rows
Test test_deferred_builtins_and_cron_narrowing_are_not_unavailable is RED on
base.
begin_process's collection-off branch purged only parked update receipts.
A provider_setup marker (pid, start_time, surface, provider) or a process
exit marker written while collection was on stayed on disk through the
opt-out and was reported (`abandoned` / `killed`) once collection came back
on, in a period the user had not consented to.
The same branch now removes process_markers/ and provider_setup_markers/.
Probe p2b_off_then_on.sh:
before: marker survives two OFF starts; the ON start emits
`abandoned openrouter`
after: markers {} after the first OFF start; the ON start emits only its
own started/completed rows
Test test_opted_out_start_purges_markers_left_while_opted_in is RED on base.
Every PUT /api/env write of a provider-shaped variable recorded
started+completed: an empty value (a clear), a same-value re-save,
GITHUB_TOKEN/GH_TOKEN (-> copilot), HF_TOKEN, and keys the TTS/STT/image
panels also save (GEMINI_API_KEY, XAI_API_KEY, DEEPINFRA_API_KEY). TUI
model.save_key counted re-saves; editing a custom endpoint counted as a new
setup.
- record_api_key_saved takes the value and the stored value read before the
save: only a non-empty, changed value counts.
- provider_for_api_key_env returns None for ecosystem tokens (GITHUB_TOKEN,
GH_TOKEN, HF_TOKEN) and any key a TOOL_CATEGORIES panel asks for: a bare
key save cannot tell a provider connection from a tool setting (the PUT
body has no provider flag; adding one is Desktop work, see NOT_COVERED).
Provider pickers (CLI, TUI model.save_key with its explicit slug) still
count those providers.
- model.save_key skips a same-key re-save.
- upsert_custom_endpoint records only when the endpoint did not exist.
Probe p4_web_double.py (real route functions):
before: openrouter started x3/completed x3 (save, re-save, clear),
copilot x1 from GITHUB_TOKEN, custom x3 (create + 2 edits)
after: openrouter x1, no copilot row, custom x1
Tests test_web_forms_count_only_a_new_provider_key_or_endpoint and the
extended test_api_key_env_maps_to_provider_only are RED on base.
cli_provider_setup treated the setup menus' BaseException control flow as
errors: Esc (_SetupCancelled) recorded failed/other, and Back (_SetupGoBack)
recorded failed/other and then a second `started` when the wizard replayed
the provider picker.
- setup_failure_class maps _SetupCancelled to `cancelled`.
- Back leaves the flow open on the thread-local; the replayed picker
re-entering the same surface+provider continues it (no new `started`).
Picking another provider, or leaving the entry point
(provider_setup_surface exit other than a Back), ends it failed/cancelled.
Probe p5_cli_nav.py (real cmd_model):
before: back_then_pick started x2, failed/other x1, completed x1;
escape started x1, failed/other x1
after: back_then_pick started x1, completed x1;
escape started x1, failed/cancelled x1;
back_then_esc_menu started x1, failed/cancelled x1;
back_then_other openrouter started+cancelled, anthropic started+completed
Test test_cli_setup_navigation_esc_cancels_and_back_resumes_one_flow is RED
on base (failed/other, extra started).
m13: the renderer stored and sent raw plugin palette/keybinding ids, rage-click
targets and config paths below user-named containers (providers.<name>.api_key);
the backend collapsed them, but only after truncating the action list to 200.
- recordAction collapses anything outside the built-in ids (KEYBIND_ACTIONS +
KEYBIND_READONLY + the named buttons, minus numbered slots — exactly the
backend's DESKTOP_ACTION_IDS) to `other` before it is counted or used as a
rage-click target.
- recordSettingsSaved sends only keys the config schema publishes (the Settings
page passes its schema: DEFAULT_CONFIG leaves), so nothing below a user-keyed
container leaves or is kept.
- Backend DAILY_ACTION_ROWS_MAX = |DESKTOP_ACTION_IDS| x |vias| (296): every
distinct collapsed (action, via) fits, so none is cut before collapsing.
m14: first-run steps happen before the consent question and collection defaults
off, so every step but consent/first_message was dropped. Chose the in-memory
buffer (smaller than removing the steps from contract, schema and TS): while the
answer is unknown/undecided, recordOnboarding/closeOnboardingStep calls are held
in memory (max 100, never persisted or sent) and replayed on the switch turning
on; a decided "off", a profile switch or quitting discards them.
setDesktopMetricsGate takes `decided` (hook passes consent.decided).
Tests (desktop-metrics.test.ts): recordSettingsSaved callers pass a schema
(contract change). New invariants, RED on base: plugin action id / rage target
stored and sent as `other`; a providers.<name> key is not sent; pre-consent
onboarding held, sent on yes (with closes applied), dropped after a decided no.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, raw ids):
before: localStorage actions {"acme-plugin:/home/alice/clients/bigco|palette":1, ...|click:3};
wire rage_click target "acme-plugin:/home/alice/clients/bigco", setting
"providers.acme-corp-internal.api_key"
after: localStorage actions {"other|palette":1,"other|click":3}; wire
[rage_click target "other"] only
M9 + the renderer half of M10, and m15.
M9: the renderer kept one install-wide localStorage record and bound the new
gateway before reading its consent, so a day accumulated under profile A was
flushed to B before B's switch was read (or dropped by B's off), and A's area
latch suppressed B's first use.
- The record is keyed per (connection, profile): `hermes.desktop.metrics.v1:
<FNV hash>` (no profile name in the key). bindDesktopMetrics takes the
profile scope; a different scope forgets in-memory state and nulls the gate,
so nothing is kept or sent until that profile's switch is read.
- The hook pins its requester to the focused (connection, profile) with
requestGatewayForAgent instead of the session-tile router, so the RPCs land
in the profile whose switch gates them. main receives the same scope hash.
M10 (renderer): each window cached the record in module memory and the last
debounced persist() won, so peer windows overwrote each other's counts. Every
change now reads the record fresh and writes it back (no debounce, no cache;
persistDesktopMetricsNow is gone). withState calls are no longer nested
(noteMessageSent / abandoned-step reporting read first, record after), since a
nested write would be overwritten by the outer one. A daily flush is pinned to
the scope and requester it started with. Both windows may still send the same
finished day; the backend's durable latch (earlier commit) records it once.
m15: consent was re-read only on attach or profile switch. The hook re-reads
it on window focus; Settings › Privacy applies its save to the gate when it is
scoped to the focused profile by name, not only when unscoped.
Tests (desktop-metrics.test.ts): `stored()` now finds the per-profile key and
setEnabled is expected with the scope argument (contract change). New
invariants, RED on base: profile switch A→B→A keeps/sends nothing of A under B
and A still reports its day; two module instances (windows) of one profile sum
into one day.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, adapted to scopes):
before: B received before its consent read: [["shared_metrics.desktop_daily",{...view.toggleSidebar 4...}]]
areas latch A: [terminal_pane] B: []
two windows stored: {"composer.send|shortcut":4} (W2's session.new lost), W2 daily sends
composer.send 3 + session.new 2
after: B received before its consent read: []; areas latch A: [terminal_pane] B: [terminal_pane]
stored: {"composer.send|shortcut":4,"session.new|shortcut":2}; W1 and W2 send the same
aggregate (the backend latch keeps one)
M11. RendererCrashRecorder held one install-wide `enabled` flag that every window
overwrote with its own profile's switch, so an opted-out window's crash was
written to disk and reported into another profile, and an opted-out window
wiped the opted-in profile's pending crash.
- Consent is kept per webContents id (the IPC sender), with a hash of the
window's focused profile key; set-enabled now carries that key.
- A crash is recorded only when the crashed window's own profile collects;
each pending entry is tagged with the profile hash (file v2: {entries:[{profile,
reason}]}; an untagged v1 file is ignored — whose it was is unknown).
- take() returns only the calling window's profile's reasons; one claim per
profile (the same window may re-claim after its renderer reloaded); ack drops
only the claimed count. Turning a profile off drops only that profile's entries.
- main.ts passes the window's webContents id to recordRendererGone.
Tests: renderer-crash-metrics.test.ts moves to the per-window API (contract
change: every call names its window, set-enabled names the profile); one new
invariant test (opted-out window records nothing, another profile neither
drains nor purges) — the base API cannot express it, so RED is shown by the
reviewer's probe instead.
Probe (review/desktop/jsrun/.../renderer-crash-metrics.probe.test.ts, adapted to
window ids):
before: after opted-out window crash, file exists: true {"v":1,"reasons":["crash"]}
take by window 2: {"reasons":["crash"]}; A crash record after B-window off: false
after: file exists: false; take by window 2: null; A crash record after B-window off: true
m16. shared_metrics.set(enabled=false) purged the renderer's onboarding latch
but left telemetry/shared_metrics/desktop_onboarding/<step>.<event> in the
profile, so a content-free per-profile record outlived the opt-out and the two
sides disagreed on re-opt-in. The Desktop's consent RPC now removes that
directory when it turns collection off (shared_metrics_desktop.
purge_onboarding_latches). Opting out from the CLI (config set / setup wizard)
leaves them until the Desktop's switch is next turned off — see NOT_COVERED.
Test: test_turning_collection_off_drops_the_onboarding_latches — RED on base
(latch file survives), green now.
Probe (review/desktop/probe_optout_set.py):
before: surviving desktop latches: [.../desktop_onboarding, .../intro.reached, ...]
after: surviving desktop latches: [] (only metrics.sqlite3 / outbox remain)
M10 + m17 (+ the backend half of M9). The backend's once-per-day latches for
hermes.desktop.{action_use,mode_use} (the finished-day report) and
hermes.desktop.feature_use lived in process memory, so a second window's report,
a backend restart or a pooled backend re-counted the same day/area. The daily
rows also went through the relay's current-day marks, so a day flushed a week
late was counted in the flush day's period.
- record_desktop_daily now writes through SharedMetricsStore.update_rollup_state:
one write transaction checks a `desktop_daily_reported` state row (the usage
days already reported, bounded to the last 8 days) and records the rows with
period_start = the usage day. Days outside [today-8, today] settle unrecorded.
- record_desktop_feature_use uses store.record_counter_once_per_day (the same
durable per-UTC-day-per-dimension latch feature_disabled uses).
- With collection off, desktop_daily answers recorded:false instead of true, so a
renderer that has not learnt the switch yet never treats that as "settled" and
drops the day (the client's own gate purges it when it reads off).
- Neither metric has feature/milestone side effects in the subscriber, so
writing to the store directly is equivalent apart from the latch and period.
Tests: the fixture mocked record_process_marks_saved; the daily/feature-use
assertions now read the profile's store (their contract changed: rows no longer
go through the relay). The disabled-profile test now expects recorded:false.
The store-refusal test raises from update_rollup_state instead of mocking the
relay's saved count. New invariant test (two windows / a restarted backend, all
in-process state reset between reports): RED on base, green now.
Probe (review/desktop/probe_latch.py, run twice on one home = backend restart):
before: run 2 → action_use composer.send 2, feature_use terminal_pane 2,
mode_use 2, all in period 2026-09-28 for usage day 2026-09-20
after: run 2 → action_use 1, feature_use 1, mode_use 1; action_use/mode_use in
period 2026-09-20. Opted-out home: daily -> {'recorded': False}.
What:
- M5: `relay` is a gateway surface alias, so an inbound the connector could not stamp is
entrypoint gateway_message / surface gateway (it counts as engagement) instead of other/none;
`relay` joins the task/session gateway platform set (it was collapsing to `plugin`).
- m9: RelayAdapter._metrics_platform labels a chat by its inbound platform, and falls back to the
connector's primary platform only when the socket fronts exactly one identity; _source_platform
uses it (egress routing keeps _chat_platform unchanged).
Why: the same turn was `other/none` in task/session while delivery/latency said `discord`; an
unknown chat on a multi-platform connector was misattributed to the primary platform.
Tests: the existing unresolvable-chat stub now implements _metrics_platform (renamed hook). New:
task_start_fields({"platform": "relay"}) -> gateway_message/gateway/relay; unstamped chat on a
two-identity connector -> delivery platform `relay`. Both RED on the previous commit
(other/other/none; discord).
Probe (review/engagement/relay/probe_d_coverage.py), before: source.platform='relay' ->
{'entrypoint': 'other', 'execution_surface': 'other', 'platform': 'none'}, engagement surface None.
After: the contract test above.
Not done: resolving the turn-start task platform to the connector's platform (needs the adapter at
agent construction) and M6 (delivery row profile under multiplex); see NOT_COVERED.md.
What:
- M1: the engagement turn mark (primary model) needs a user turn (pre_llm_call) on an interactive
surface or a gateway message; API-server, python embedding and curator turns no longer create
engagement.day rows or outvote the person's model. close_day emits nothing for a day with no
engaged surface (the root profile's host row with an active-profile count stays).
- M3 (owner ruling): unattended cron runs are dropped from engagement entirely. surface_day always
carries an active-minutes bucket, so it cannot express cron presence without minutes; hermes.cron.run
already counts runs. The `cron` enum value stays in the contract for already-stored rows.
- M7: background review forks (no pre_llm_call) no longer advance model_switch_after's turn run.
- M8: the run observes the route each turn was sent on (its first request), so a provider failover
neither restarts the run nor names the fallback as the model left.
- m6: documented that the root host row (surfaces_used_count 0) is excluded from days-active counts.
Why: Hermes-internal and unattended activity is not user engagement; switch_after measures how long
a person stayed on the model they chose.
Tests (RED on the previous commit): API/cron-only day -> no day row, mixed day -> primary = person's
model; failover + review fork -> switch_after names the selected model with 2_to_3.
Probe (review/engagement/probe_internal.py, switch/probe_switch.py -k "p3c or p7"), before:
API-only / curator-only day -> engagement.day {"primary_model": "openai/gpt-5", "surfaces_used_count": "0"}
cron all day -> day {"active_minutes_bucket": "gte_6h", "active_profile_count_bucket": "1"} + surface_day cron
2 CLI + 5 API turns -> primary openai/gpt-5
P3c (6 turns on A, 7th failed over to B) -> {"model": "gpt-5", "turns_before_switch_bucket": "1"}
P7 review fork -> 4_to_10 (routed=False) / gpt-5-mini "1" (routed=True)
after: the equivalent invariant tests pass (no day row, primary = PUBLIC, selected model 2_to_3).
What: the metrics lineage no longer keys off the ambient Portal conversation id (it walks every
parent_session_id: gateway reset/idle expiry, /new, /branch). Compression calls
relay_shared_metrics.rotate_segment(old, new) from the committed rotation
(_notify_context_engine_compression_complete) and from a stale agent adopting the live tip
(_adopt_live_compression_child). The hand-off registers the new id in the lineage, moves the
route run and the spent-turn history into it, and closes the old segment now (or when its in-flight
compression turn ends). A rotated-to id closed before serving a turn, or still open at shutdown,
closes too; a 512-lineage backstop flushes conversations a surface never closed.
- m2: /undo or /retry on the new id reaches the turns earlier segments spent (tokens_bucket known).
- m11: /model right after a rotation finds the conversation's turn run.
- m8: background review forks (same session id, no pre_llm_call) add no turns/calls/messages.
Why: one conversation = one session / context_peak / tool_overhead / tool_enabled_unused row, on
time; N gateway conversations collapsed into one row emitted only at process exit, and rotated
segments stayed in memory until exit.
Tests: two existing tests simulated rotation by sharing a conversation context; their contract
changes to the explicit hand-off (test_a_compressed_conversation..., test_context_peak_is_one_row...).
New: reset/branch chain after a compression + tip-first close order, the compression seam, switch/undo
right after a rotation, review fork volume. RED on base (rotate_segment missing / switch_after [] /
turn bucket '2' for one user turn).
Probe (review/engagement/lineage probe_lineage.py -k c2, probe_branch.py; fix-lineage/probe_handoff.py):
before: c2 3 conversations -> session.count 0 until shutdown, then 1; branch -> 1; 50 compressed
conversations -> 0 rows, 50 lineages/sessions live.
after: c2 -> 3 rows before shutdown (turns 2,1,1), lineages=0; branch -> 2; 50 compressed -> 50 rows,
0 lineages, no sessions; tip-first close with the old turn running -> 0 then 1; shutdown with an
open tip -> 1 per conversation.
What:
- shared_metrics_catalog.model_metric_name also collapses a model id that is
a network address: a loopback/IPv4 host or localhost at the start, or a
leading host:port with a 2-5 digit port (so Bedrock's "...-v1:0" and
OpenRouter's ":free" / Ollama's ":120b" tags stay readable). The existing
URL/path/weight-file rules move into the same compiled pattern.
- shared_metrics_contract._bounded_dimensions (every mark -> counter
projection) and model_token_counters re-run provider_metric_name /
model_metric_name on the mark's provider/model identifier fields and drop
the row when either would rewrite it ("none" = unset provider passes).
Record-time only: counter_dimensions_are_valid (also used at packaging) is
unchanged, so catalog drift can never raise out of _package_metric and
block a package of rows already stored.
Why: `127.0.0.1:8080/x` on a public provider passed verbatim through every
route metric, and the subscriber validated identifier shape only, so a mark
that skipped the producer's catalog pass stored custom:acme, ollama/<model>,
unknown vendors and even http:// URLs.
Probe (fix-catalog/probe_b1_m12.py):
before: route(openrouter, 127.0.0.1:8080/x) -> model '127.0.0.1:8080/x'
(also localhost/qwen, 10.0.0.5/x, gpu-box.lan:8000/qwen verbatim)
subscriber injected switch_after {127.0.0.1:8080/x, openrouter} -> stored
subscriber injected switch_after {acme-secret, acmecorp-internal} -> stored
after: all of the above -> model 'custom'; both injected marks -> None;
bedrock anthropic.claude-3-5-sonnet-20241022-v2:0, openai/gpt-4o:free,
gpt-oss:120b and the legit anthropic mark unchanged.
Test: test_network_address_model_ids_and_raw_injected_marks_never_reach_counters
(RED with the catalog reverted: model_route kept 127.0.0.1:8080/x; RED with the
contract reverted: the injected host:port mark was stored; GREEN with both).
What: shared_metrics_catalog.provider_names() no longer reads the live
PROVIDER_REGISTRY / models._KNOWN_PROVIDER_NAMES. It is now the built-in
auth rows (BUILTIN_PROVIDER_IDS), HERMES_OVERLAYS, the three static alias
tables, openrouter/custom, and the names+aliases of profiles registered
from the in-tree plugins/model-providers dir (providers._SOURCES ==
"bundled", process-wide layer, independent of the bound profile home).
Why: auth_plugin_providers mirrors every $HERMES_HOME (and pip) provider
plugin, plus its aliases, into PROVIDER_REGISTRY and the picker labels, so a
user-chosen provider name was treated as shipped and its model id was not
collapsed either. Every consumer goes through provider_metric_name, so the
one fix covers model_route and every per-model v4/v5 metric (tokens,
context_peak, friction, tool_quality, loop_guard, tool_recovery,
model_reply_issue, task_cost, wasted_tokens, cache_break, switch_after,
tool_unavailable, engagement.day), provider_setup (all surfaces, incl. the
PUT /api/env key path), the on-disk pending marker, setup.completed,
model_switch/fallback and install.snapshot main_provider. Pip-installed
provider plugins now read custom too (third-party, not shipped); doc says so.
Probe (fix-catalog/probe_b1_m12.py, user plugin acmecorp-internal alias
acme-llm):
before: provider_metric_name('acmecorp-internal') -> 'acmecorp-internal'
route: {'model': 'acme-secret-model-v2', 'provider': 'acmecorp-internal'}
marker on disk: {..., "provider": "acmecorp-internal"}
after: provider_metric_name('acmecorp-internal') -> 'custom' ('acme-llm' too)
route: {'model': 'custom', 'provider': 'custom'}; marker "provider": "custom"
install main_provider custom; switch/fallback/setup custom
shipped plugin deepinfra: {'model': 'meta-llama/...', 'provider': 'deepinfra'}
Test: test_user_provider_plugin_name_and_model_never_leave (RED on base:
marker carried "acme-llm"; GREEN after).
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in
the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality.
- hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface
(cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow;
a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern).
- hermes.feature_adoption.count: once per feature per install at first real use, derived from existing
counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the
owning profile's install age.
- hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the
re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day.
Surface comes from the user entry point; setup/migrations record nothing.
Desktop love/hate telemetry needs a backend half: six counters
(feature_use, action_use, mode_use, friction, dislike, onboarding) with
closed dimension sets, the v3 JSON schema entries, the public doc rows,
and profile-scoped gateway RPCs the app reports through.
Every value the app sends is collapsed server-side onto the closed set
(unknown -> other); feature use latches per area per UTC day, friction
and dislike are capped per day, the daily report (action counts + mode
time + bot count) latches per day and answers recorded=false on a store
failure so the app keeps the day for a retry. For setting changes the
app sends only the config key: the backend reads the saved value and
compares it to DEFAULT_CONFIG itself, so values never cross the wire;
keys outside DEFAULT_CONFIG read other. All helpers check enabled()
before doing any work.
Engagement (local daily rollup, emitted once the UTC day closes, exactly once per profile):
- hermes.engagement.surface_day.count {surface, active_minutes_bucket} per surface used
(cli/tui/desktop/gateway/acp/cron); active time = inter-interaction gaps capped at 5 min.
- hermes.engagement.day.count {active_minutes_bucket, surfaces_used_count, primary_provider,
primary_model, active_profile_count_bucket}; primary model = most attended turns that day.
Active profiles are folded into one host accumulator in the root profile's store (opaque local
hashes); only the root profile's row reports the count, other profiles report 0.
- Days-active-per-week / return-by-model are derived server-side; no weekly state, no new id.
Model satisfaction:
- hermes.model_switch_after.count {provider, model, turns_before_switch_bucket}: turns the model
being left served in the conversation, from every /model site (CLI, TUI, ACP, gateway).
Decisions:
- hermes.session.count counts once per conversation across compression rotation (merged over the
v4 conversation lineage), and gains message/model-call/tool-call count buckets with a longer
tail (251_to_1000, gte_1001); rows recorded before the upgrade still package.
- Relay-delivered hermes.gateway.reply_latency and hermes.platform.delivery report the originating
platform instead of `relay`.
Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):
- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
A Hermes-known break suppresses the miss it causes. No prompt hashing.
Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):
- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
matcher reports the strategy that landed (or no_match / ambiguous) into a
context-local probe opened only around the patch/write_file handlers, so we
learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
the guardrail already warns/blocks/halts and at turn end for the iteration
budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
next_outcome. One row per failed tool call, resolved against the model's
next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
a table lookup on the first program word, never the text. Hermes' own
deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
backend result, so a command's own `exit 124` reads nonzero, not timeout.
Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
truncation only from the structured finish reason; empty/reasoning_only
from the normalized message; issue=none per response as the denominator
(model_route counts attempts and files empty replies as failures).
Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat
call added a row. The TUI and Desktop now send an opaque per-launch id
(TUI: one per process; Desktop: minted when the once-per-launch main-process
claim succeeds) and the backend claims (surface, launch_id) once per
process: a reconnect re-sending the same launch counts once, a new Desktop
launch against a long-lived backend still counts. Legacy clients without an
id count once per backend process and surface. The id is only a local latch
key (bounded set, length-capped), never recorded. Only a usable measurement
spends the claim. Contract + apps/shared regenerated.
- m10: relaunch() execs in place, keeping the PID, so `sessions browse` ->
resume counted the picker time as CLI startup. relaunch() stamps its PID in
HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it
matches, while children (other PIDs) inheriting the env still count.
- m8 (psutil half): record_process_ready checks the opt-in gate before reading
the process start time (psutil) or starting a thread. The claim is taken
first so a later opt-in never records a stale "startup" mid-session.
Test changes to existing assertions: the TUI/Desktop param assertions now
expect launch_id (a new wire param); the RPC surface test sends distinct
launch ids so its three calls stay three launches under the new latch.
Review findings M7, M8, m6, m7 and the browser half of m8 on the v4 loop metrics.
- M7: TUI/Desktop @-path completion on a non-local backend lists the directory through
terminal_tool, so every keystroke added a hermes.execution_backend.count row as if the user
ran a command. Hermes-owned calls now run inside shared_metrics_loop.unmetered_backend_calls()
(a contextvar checked in record_execution_backend) instead of a flag threaded through
terminal_tool. It was the only non-_host_local internal terminal_tool caller (bot DM runners
already use _host_local; the prompt-builder probe calls env.execute directly).
- M8: a scheduled curator tick that finds the run claim held recorded scheduled/skipped every
tick (5 ticks -> 5 rows for one pass). The holder's own pass is the one counted; the held
branch now records nothing. Real dry runs still report skipped.
- m6: a foreground timeout returns exit 124 with partial output and no error field, so it was
recorded as success. exit 124 now records outcome=failed, error_class=timeout.
- m7: vercel_sandbox (and managed_modal) are built-in backends in
tools/terminal_tool_config._BUILTIN_BACKENDS but collapsed to other. Added to
TERMINAL_BACKENDS and both schema enums; a test pins every built-in backend to its own bucket.
- m8 (browser): record_browser_call resolved the browser backend on every call even with
collection off. The resolver is now passed lazily and only runs inside the gated builder.
The existing test asserting a held claim records scheduled/skipped encoded M8 itself and is
replaced by one asserting held ticks record nothing.
Three defects in how hermes.update.run/stage and hermes.process.exit rows
are recovered from disk:
- Opt-in (M1): the pre-pull park path only checked that the telemetry dir
existed, so it wrote the WHOLE receipt (argv, step detail) while the user
was opted out, and a later opt-in counted that run. It now reads consent
through the already-loaded hermes_cli.config (this interpreter predates
the checkout swap, so nothing may be imported), parks only when
collection is on, parks only the fields update_receipt_fields reads (no
argv, no step text), and purges pending_updates when collection is found
off. begin_process purges them too on a start with collection off.
- Double count (M2): the completion child records the receipt and the
parent's boundary finalize (run_completion returned receipt=None) parks
the same update_id again, so one run became (failed, deps) + (success,
none). record_update_receipt now claims the update_id with an O_EXCL
latch under shared_metrics/recorded_updates (last 64 kept); the second
finalize records nothing.
- Loss (m3): dead-process markers and parked receipts were unlinked before
their row reached the store; "database is locked" under concurrent starts
lost them (17/20 rows in the reviewer's race). The claim-by-rename stays,
but the file is deleted only after record_process_marks_saved confirms
the rows settled: rows carry a random commit ticket in event metadata
(allowlisted in the contract), the subscriber tallies tickets whose rows
persisted, and the reporter flushes the Relay subscribers and compares.
Otherwise the file is renamed back for the next start. A parked receipt
whose rows partly landed is not retried (that would double count).
Also: _epoch returns None for a timestamp float() cannot represent
(10**400 raised OverflowError).
Rotating compaction continues a conversation under a new session id, and
each metrics session emitted its own context_peak at close, so one
conversation produced a row per segment (the fuller one plus a low-fill
one), skewing the per-model distribution. The agent turn already
publishes the lineage root (the Portal conversation id); segments sharing
it join one lineage, contribute their peak when they close, and the
merged peak (fullest fill, limit hit anywhere) is emitted once when the
last open segment closes, whatever the close order. Delegated children
share their parent's root but stay their own conversation.
Review finding m2.
provider_names() accepts ollama/local/vllm/llamacpp/llama.cpp as published
names (they are aliases of the generic `custom` provider), so every
provider/model metric kept the raw config.yaml model id under them, and a
missing provider (`unknown`) kept it too. Map the aliases Hermes itself
routes to `custom` (derived from the providers/auth/models alias tables)
to `custom`, and report the model as `custom` whenever the provider is
unknown: nothing proves such an id is public. Shipped providers with
public ids are unchanged. Covers model_route, model_tokens (primary and
auxiliary), fallback, model_switch, setup.completed, install snapshot
main_provider, friction, tool_quality and context_peak, which all go
through this pair.
Gateway /model also labelled the configured model with switch_model's
`openrouter` default when config.yaml names no provider; it now reports
the provider actually configured/overridden (None -> `unknown`). And the
switch row + switch_away friction bind the routed profile's home on a
multiplexed runner (slash dispatch installs no profile scope), as the
/retry and /undo friction already do.
Review findings B2, M3.
The update pipeline already writes one machine-readable receipt per run, so the
run/stage counters are derived from it in the process that finalizes it instead
of instrumenting stages for metrics. The receipt gains minimal stage END marks
(plan, snapshot, apply[mode], deps, build, restart; verify is inferred from the
fleet matrix), an initiator fact (desktop when a live orchestrator claim
names another pid) and the pre-update commit date, which is all the derivation
needs for outcome, failed stage, duration and from-version-age buckets.
A pre-pull interpreter must never import pulled code: when it is the finalizer
(same pid, checkout sha moved) it parks the receipt (stdlib only) and the next
Hermes start records it. Collection off => nothing imported beyond a config read.
The completion-process test stub gains the new record_stage API the real
module now exposes (no assertion changed).
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.
hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.
The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.
hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.
hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.
hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
The gateway and cron ticker were blind spots in shared metrics: we could not
tell which messaging platforms fail to connect or drop, how often replies fail
to reach users, how long users wait for a first reply, or whether scheduled
jobs run, fail, get skipped or get missed while Hermes was down.
New opt-in counters (all through the existing enabled() / record_process_mark
gate, recorded fire-and-forget on one background worker that runs in a copy of
the caller's context so the owning profile gets the row):
- hermes.platform.health: platform, event (connect_ok/connect_failed/
reconnect/disconnect), error_class. One seam in the runner
(_connect_adapter_with_timeout, used by cold start, multiplex secondaries and
the reconnect watcher) plus the fatal-error handler; classified from
exception types, HTTP statuses and Hermes's own fatal codes, never text.
- hermes.platform.delivery: platform, outcome, failure_class. One row per
logical BasePlatformAdapter._send_with_retry call (retries and plain-text
fallback included), from SendResult's typed fields or the exception type.
- hermes.gateway.reply_latency: platform, first_response_bucket. Clock starts
when _handle_message accepts a non-internal turn; stops on the stream
consumer's first delivered text or send_final_ledgered. Busy acks and
command replies never stop it.
- hermes.cron.run: outcome (success/failed/missed/skipped), delivery_kind,
duration_bucket. Recorded at the write-once terminal execution row
(finish_execution) and where the due scan drops an occurrence (catch-up
disabled, expired one-shot). Job names, prompts and targets never leave.
Platform names: platforms Hermes ships under plugins/platforms/ now report by
name, and a plugin-catalog platform reports its catalog entry name only when
the installer-owned .install-metadata.json record proves a catalog install of
the plugin dir that defines the registered adapter factory (a URL install
cannot claim a catalog name via its tree or manifest). Everything else stays
"plugin". GATEWAY_PLATFORMS becomes catalog-backed (_CatalogValues), and the
v3 schema's platform fields accept published catalog identifiers.
Adds four opt-in shared-metrics counters so we can see whether the learning loop
actually runs and pays off, how wide delegate_task fan-outs go, and which
sandboxes carry real work:
- hermes.memory.op.count: op, provider (builtin / bundled plugin / plugin),
origin (foreground / background_review), outcome (success / failed / rejected)
- hermes.curator.run.count: trigger, outcome, archived/created/merged/patched buckets
- hermes.delegation.run.count: subagent_count_bucket, depth, mode, outcome
- hermes.execution_backend.count: kind (terminal/browser/code), backend, outcome,
error_class
Every dimension is a closed enum or COUNT_BUCKETS value; unknown providers and
backends collapse to plugin / other. Builders live in the new
shared_metrics_loop.py sibling (stdlib-only imports so tool hot paths can use it)
and go through the existing enabled() gate. Off-turn producers (curator thread,
async delegation units) record in the profile captured when the work started.
Schema and the "what is collected" doc gain matching entries.
The first-run consent dialog blocked the composer on every undecided
profile, so a healthy install no longer opened straight to chat (caught by
the Desktop first-run E2E). The question now lives in the composer status
stack beside the free-tier strip: Send to Nous / Local only / No thanks /
Details, no focus steal, one owner across split composers, one offer at a
time. Details opens the full explainer on request; closing it decides
nothing. The backend's `decided` stays the only latch.
- Docs: a Decision-data metrics section mapping each metric to the product question it answers;
skill-load and snapshot paragraphs describe the new public-name and install-age fields.
- setup consent copy names the new data classes (session length, token totals, command and
catalog names, bucketed setup counts).
- The real-CLI smoke now asserts the session summary, token sums, TTFT bucket and milestones in
both the SQLite store and the exported, schema-validated package.
- A profile with no sessions yet is a brand-new install (lt_1h), not unknown: the first task's
milestone and the first snapshot fire before state.db has a session row.
- Two invariant tests: session summaries + once-per-install milestones; token sums per model and
auxiliary task with private task names collapsed.
The fleet rollups could say that tasks and tools failed, but not why, on which
model, on which messaging platform, or which features installs actually use.
- model_route rows gain call_role, outcome and error_class (the classifier's
FailoverReason for the last failed attempt; a success row keeps the error it
recovered from).
- task_run rows gain platform (built-in gateway platform, plugin, or none);
task_run.finished gains failure_class from a closed set.
- new hermes.tool.usage.count: tool_name (only names from the static
toolsets.TOOLSETS snapshot; mcp/plugin otherwise), outcome, error_class.
- new hermes.install.snapshot, latched to once per rolling 24h like
client.active: memory provider plus bucketed counts of MCP servers,
plugins, skills, cron jobs and profiles; no names.
- package schema v3; v2-shaped pending counters still validate and drain.
- consent copy and developer docs list the new categories; the relay smoke
asserts v3 shapes and disables title generation (its third model request
broke the smoke's request-count check on main).
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.
Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
* fix(compression): compaction follows the ratio again; no default token cap
compression.threshold_tokens defaulted to 256000 since #115986, so every
window above ~341K compacted at 256K regardless of compression.threshold
or model_thresholds: a 1M-window model compacted at 25% of its window,
and threshold: 0.8 still compacted at 256K. No single token count suits
windows from 64K to 1M+, so the cap goes back to opt-in (null) and the
ratio decides, as it did before #115986.
The template seeder and `hermes doctor --fix` copied the 256000 default
into config.yaml, where it reads as a user choice and would keep capping
those installs. Config v47 drops threshold_tokens only when it equals
256000; any other explicit cap and an explicit null are preserved.
The rest of #115986 stays: the model-switch warning quoting the real
trigger and the shared _derive_trigger are correct with any default.
Tests that encoded 256K now derive expectations from DEFAULT_CONFIG.
* docs(compression): threshold_tokens is an optional cap, default null
User guide, developer guide and delegation page described the 256K
default cap; they now describe the ratio trigger as the default and
threshold_tokens as an opt-in cost ceiling.
A named import of a missing SDK export fails to link through the runtime
shim, so the plugin never loads; it is not undefined. Point authors at a
namespace-import feature-detect or the raw Contribute form.
Refs #123597
Originally authored by Justin Haynes (@jhaynes).
A Kanban board opened in a split route tile rendered no board switcher:
the board contributed it to WORKSPACE_PAGE_HEADER_AREA unconditionally,
and only the workspace pane paints that area. The tile's contribution
also leaked into another page's header and shared its id with the full
page's, so closing the tile removed the page's switcher.
Add WorkspacePageHeaderControl (exported via the plugin SDK). The
workspace pane's render provides a private host context; inside it the
control projects into the page header, anywhere else it renders inline.
The board mounts BoardSwitcher once, through it, in its own header row.
Fixes#123597
Originally authored by Justin Haynes (@jhaynes).
model.info and saved profiles report a user-defined provider as custom:<key>, but the catalog row uses the bare key as its slug. Settings > Model and the Bot Mode picker compared the two with ===, so a saved custom provider never found its row. Settings showed a duplicate custom:<key> entry and a Set up provider button, and the bot editor fell back to the manual form.
Both now match rows with catalogProviderMatches, like the composer picker already does. Settings uses a small findCatalogProvider helper for every row lookup, including the aux and MoA slots and the endpoint passed on Set to main. catalogProviderMatches is now exported through the plugin SDK so the bot picker can use it.
A deterministic fallback summary replaced the older handoff in the transcript but never updated _previous_summary. The next compaction kept the stale in-memory summary and dropped the fallback row from its window, so the fallback's user asks, files and last dropped turns never reached the summarizer. Store the fallback body in _previous_summary the same way a normal summary is stored.
(cherry picked from commit 35417d2e1ffbb775c3eaff17b26623896afa56c1)
Compaction spared a pending tool round's text results but not its
images. The image-retirement pass kept only the newest three image
results across the transcript, so four parallel vision calls in one
unread round lost the oldest image before the model saw any of them.
On a single-prompt run the final media pass lost three of the four:
compaction re-appends the task as a user row after the round, and the
pending round was looked up after that row was added, so nothing was
spared.
Both image passes now skip a pending round that fits the hard share,
and the pending round is found before the task row is re-appended.
Completed rounds and a pending round over the hard share keep the
existing policy.
(cherry picked from commit eb7661f4365f009d5f9ef85e26f2f4b6ac83632e)
The pending-round check skipped only one trailing /steer row. Two can
land in the same iteration: one is appended when the tool batch ends,
and a steer sent after that is drained before the next request and
inserted right after the newest tool result. The transcript then ends
tool -> steer -> steer, so the check stopped on the second steer row,
found no pending round, and preflight compaction replaced the unread
tool output with the one-line pressure stub.
The check now skips every contiguous trailing steer row, identified the
same way as before, and still stops on any other user row. The existing
regression gains a case with two steer rows after the round.
(cherry picked from commit 455a1a3868b3843d15401781363ce796bae96b5b)
Preflight compaction can fire right after a tool round, before the model
has read its results. The protected-tail passes then treated that round
like any old output: pass 2 and the pressure pass (#61932) replaced its
results with one-line stubs, and when the newest row alone exceeded the
tail ceiling the cut landed at the end of the transcript and summarised
the round with its whole turn. The model then re-ran the call, side
effects included, or answered without the output.
_pending_tool_round finds the tool results the transcript ends with,
skipping one trailing /steer row (a steer is delivered after the newest
result, before the next API call). Both passes spare that round, and the
tail cut aligns from the row before the end so the group stays whole.
The one exception is a round that alone exceeds the tail's hard share
(20%) of the input budget, the window minus the output reservation: it
still gives way, as #61932 requires. _effective_input_window is
extracted from _compute_threshold_tokens so both use the same budget;
thresholds are unchanged.
(cherry picked from commit 9d74e22379cd7dc39636c522175772d0165db4ac)
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.
After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.
Fixes#123167
(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.
Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
Every gate and every candidate now needs only admit, so the signed macOS
and Windows builds and the docker image stop waiting behind the full CI
run. acceptance stays the one join: it still needs ci, docker and all six
candidates, so what can be promoted is unchanged.
The four gates that skipped under --skip-tests only because ci skipped
(nix, termux-checks, windows-live, install-e2e) and pm-bundle needed their
own condition. SKIPPED_BY requires the gate to observe 'skipped', so
dropping the edge alone would have left them running and blocked the
release instead of failing it.
`release.py release` gains two flags. They can be used together.
--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.
--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.
The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.
The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.
A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:
- The next version was derived from it, so the next cut would reuse the
version. It now takes the newer of the R2 head and the newest
published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
their publication pass, so a bundle-less release would re-advance
every 15 minutes. The head is now the newer of the R2 head and the
published release whose final tag binds the Docker stable alias
digest.
`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.
Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:
- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]
Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
Stable Release Publication ran every 15 minutes (96 runs a day, each
checking out full history, setting up node and buildx, logging into
Docker Hub, and taking the release-signing environment) only because the
sequencer held a failed run for a 15-minute backoff that the failure
event could never satisfy, so the cron was what actually retried.
Drop the backoff: the reconcile pass started by a failed Stable Release
reruns its failed jobs right away. MAX_ATTEMPTS burning, oldest-first
retry ordering, the attempt-entry check, and the needs_retarget repair
stay. The schedule trigger goes; workflow_run and workflow_dispatch
remain the recovery paths.
The shared stable-release concurrency group cannot deadlock: the rerun
waits as pending behind this job, and the sequencer only confirms the
new attempt is queued before it exits and frees the group.