Commit Graph

45854 Commits

Author SHA1 Message Date
teknium1
2bb227c596 docs(telemetry): model rules name AWS ARNs and loopback provider aliases 2026-09-28 12:43:03 -07:00
teknium1
3ba7c850ba fix(telemetry): Desktop provider-connection forms count Gemini/xAI/DeepInfra keys again (M13)
The M13 fix stopped counting any key a tool panel also asks for, since a
bare PUT /api/env cannot tell connecting a provider from configuring a
tool. Desktop onboarding and model settings now send provider_setup=true
and the server counts those saves; the Keys page and tool panels keep
the exclusion.
2026-09-28 12:43:03 -07:00
teknium1
58530cc7f5 fix(telemetry): an opt-out from the CLI also drops the Desktop onboarding latches (m16)
Only the Desktop consent RPC purged desktop_onboarding/; opting out with
hermes config set or the setup wizard left the latches until the Desktop
switch was next flipped. An opted-out process start now removes them
with the other pending markers.
2026-09-28 12:43:03 -07:00
teknium1
0b6ac0b622 fix(telemetry): live relay smoke no longer expects adoption from an agent-created skill
The signals fix (M15) stopped counting Hermes' own skill creation as user
adoption, so the smoke's agent-created skill records no feature_adoption
row; the exact name sets now assert its absence.
2026-09-28 12:43:03 -07:00
teknium1
b67d9ed6fe fix(telemetry): platform.delivery lands in the profile that owns the chat (M6)
In a multiplexed gateway the final send runs after the routed profile
scope is reset, so delivery rows went to the launch profile while the
same turn's reply_latency went to the owner. start_reply_clock now also
remembers the owning home per (platform, chat) in a bounded map that
outlives the clock, and records_delivery submits through it.
2026-09-28 12:43:03 -07:00
teknium1
a2d08fba36 fix(telemetry): loopback provider aliases and AWS ARNs collapse model ids to custom
model_route('lm-studio', id) kept the user's model id because only the
canonical lmstudio was in user_named_model_providers while the alias
passed provider_metric_name as a shipped name. Bedrock ARNs carry the
AWS account id. Both now read custom.
2026-09-28 12:43:03 -07:00
teknium1
d6e4e33d7d fix(telemetry): live relay smoke expects the v5 efficiency and adoption rows
The smoke's exact metric-set checks (SQLite store and export packages) were
never updated for v5: every attended CLI turn now also emits task_cost,
tool_output_truncation, tool_overhead and tool_enabled_unused, and the skill
lifecycle step latches feature_adoption(skills_created). A correct run failed
with "Unexpected SQLite counters", so the documented live gate could no
longer catch a real regression. Both sets now include the five metrics and a
shared _validate_v5_rows asserts their dims in the store and the packages
(task_cost provider/model `custom` for the canary model, api_calls 2,
tool_calls 1; read_file truncation; file toolset used; tool_overhead on cli).

Live smoke (loopback fake OpenAI server, no model spend):
  before: rc=1 AssertionError: Unexpected SQLite counters
  after:  rc=0 "Hermes -> NeMo Relay shared-metrics smoke test passed"
Sabotage: expecting the canary model id in task_cost is caught.
2026-09-28 12:43:03 -07:00
teknium1
d19b204b33 fix(telemetry): relay-shared-metrics doc — valid gateway table, owner rulings
- m10: the v5 relay-platform note sat inside the gateway metrics table, so
  every row after reply_latency (cron.run, startup.latency, update.*,
  process.exit) rendered as one paragraph. The note now follows the table.
  probe_e_doc_table.py before: cron.run/startup.latency/process.exit "not
  found as cell"; after: all inside <td>. Whole doc: 49 metric rows, rendered
  table rows 44 -> 49.
- platform.health for the relay connector stays `relay` (one socket fronts
  several platforms): documented next to the relay note (owner ruling).
- Delegated subagents count in the harness metrics and in cache_break /
  tool_output_truncation (model and tool behaviour); background review and
  curator never do (owner ruling; matches probe_internal.py).
- Model-route rules: lmstudio keeps its provider name and only its model reads
  `custom` (behaviour kept per owner ruling); the local-alias list now matches
  custom_provider_aliases() (adds llama-cpp).
2026-09-28 12:43:03 -07:00
teknium1
d8d9639d8e fix(telemetry): execution_backend skips background review / curator fork calls
hermes.execution_backend.count still counted terminal/browser/code calls made
under the background_review write origin (the review and curator forks),
while hermes.terminal.outcome and file_edit already skip them. Owner ruling:
match terminal.outcome. record_execution_backend now returns early when
is_background_review() is set (one contextvar read, before any work).
Doc: the execution_backend row lists the forks as not counted.

Probe probe_internal.py (terminal_tool inside a background-review origin):
  before: [execution_backend {backend: local, kind: terminal, outcome: success} x1]
  after:  []
test_background_review_fork_terminal_calls_are_not_backend_usage is RED on base
(2 rows != 1), green here.
2026-09-28 12:43:03 -07:00
teknium1
6cc2f4ab92 fix(telemetry): terminal.outcome=timeout comes only from Hermes' deadline
_run_foreground recorded outcome=timeout whenever a backend exception's
text contained "timeout" (an SSH connect timeout, a sandbox API timeout).
That command never reached an exit status, and the doc says timeout comes
only from Hermes' own deadline flag (hermes_timed_out), which
terminal_outcome() already reads. Owner ruling: never from exception text.
The exception path now records nothing; the model-facing result (exit 124)
is unchanged.

Probe (code-read finding; the new test drives the real terminal_tool with a
backend raising "... Connection timeout"):
  before: terminal.outcome {backend: local, command_kind: shell_builtin, outcome: timeout} x1
  after:  no terminal.outcome row; tool result still exit_code 124
test_backend_exception_mentioning_timeout_is_not_a_terminal_timeout is RED on
base, green here.
2026-09-28 12:43:03 -07:00
teknium1
5816a768e4 fix(telemetry): CLI /undo N counts the user turns undo_last really removed
_cmd_undo counted role=="user" rows of the old history, so a compaction
handoff (role user, not a turn) added a phantom wasted_tokens row with
tokens_bucket=unknown. The TUI and gateway already pass the real
turns_undone. undo_last now returns turns_undone (None when nothing was
rewound) and _cmd_undo records exactly that.

Probe probe_cli_undo.py (history = handoff + 2 real turns, /undo 5):
  before: wasted_tokens 10k_to_50k x2 + unknown x1 (3 rows)
  after:  wasted_tokens 10k_to_50k x2 (2 rows)
Test test_cli_undo_n_counts_the_user_turns_actually_undone_not_compaction_handoffs
is RED on base (3 == 2), green here.
2026-09-28 12:43:03 -07:00
teknium1
ab056dba47 fix(telemetry): feature_disabled skips migrations and ${VAR} templates, diffs outside caller locks (m22, m23, m24)
m22: migrations ran save_config -> record_config_saved inside surfaced
processes (`hermes config migrate`, console migrate, dashboard/TUI profile
create), so e.g. v45 appending `connections` to platform_toolsets read as a
user re-enabling a toolset. _persist_migration (the single migration write
path) now wraps save_config in hermes_applied_write(), a ContextVar the hook
honours.

m23: the old side is the raw file (`memory_enabled: ${MEM_ON}`) and the new
side the env-expanded config, so any unrelated save reported `disabled` once a
day forever. A setting whose value on either side is an unexpanded `${...}`
string is now skipped (its value is unknown to the diff).

m24: dashboard handlers call save_config inside config_write_scope
(_CONFIG_MUTATION_LOCK), and the diff (tools_config._get_platform_tools per
platform) + record ran there. record_config_saved now runs only the cheap
gate inline (surface, enabled, surface/home resolution) and hands deep copies
to a non-daemon thread bound to the owning profile (non-daemon so a CLI
command exiting right after its write still records).

Test contract change: test_record_config_saved_needs_a_surface_and_collection
and test_real_config_writes_report_the_move_away_from_default now join the
recording thread before asserting (the record is asynchronous by design).

Probes (review/signals/mine):
  p2_migration.py config|dashboard
    before: re_enabled toolset connections row; after: <no db>
  p4_envref_configset.py
    before: memory.memory_enabled disabled after an unrelated save;
    after: no row (config set curator.enabled false still records)
  p3_lock_v2.py (times config_transitions, try-acquires the web lock)
    before (p3_lock.py): record_config_saved 0.047s/0.004s, lock held=True
    after: diff 0.003-0.008s, lock held=False
Tests test_migrations_and_env_templates_are_not_user_disables and
test_diff_and_record_run_off_the_callers_lock_in_the_owning_profile are RED on
base.
2026-09-28 12:43:03 -07:00
teknium1
85c317b0a9 fix(telemetry): feature_adoption ignores Hermes-internal curator/skill work; age read never raises (M15, m21)
- `curator` latched on the default-on scheduled pass (outcome=success, all
  buckets 0), so it meant "idled through one curator interval". It now needs
  trigger=manual (the user ran it).
- `skills_created` and the v4 milestone `first_skill_created` latched when
  the background-review fork created a skill (provenance=agent_created).
  Both now require provenance != agent_created (foreground creates are
  stamped "learn" -> not agent_created, see tools/skill_usage.record_created).
- days_since_install_bucket caught only sqlite3.Error; a non-numeric
  sessions.started_at raised ValueError into the subscriber on every event,
  re-opened state.db each time and never latched. It now also catches
  OSError/TypeError/ValueError -> `unknown`.

Probes:
  probe_internal_adoption.py
    before: feature_adoption curator + skills_created, milestone first_skill_created
    after:  no feature_adoption rows, no milestone
  probe_statedb_errors.py textstamp
    before: ValueError warnings, 0 rows, settled 0 of 1
    after:  1 state.db read, feature_adoption {unknown, projects}, settled 1 of 1
Tests test_hermes_internal_work_never_latches_adoption and
test_unreadable_first_session_reads_unknown are RED on base.
2026-09-28 12:43:03 -07:00
teknium1
3a0eaff6b2 fix(telemetry): tool_unavailable skips tool_search-deferred built-ins and cron (M14, m20)
tool_unavailable_fields compared against agent.valid_tool_names, the
model-visible array after tool-search assembly. With the default
tools.tool_search.enabled=auto, 14 enabled built-ins (session_search,
todo_list, process_manage, ...) live behind the tool_call bridge, so a direct
call to one was reported as "disabled in this session". Cron agents had
clarify stripped on purpose by the scheduler and still counted.

- Names in the session's tool-search scoped set
  (agent.tool_executor._tool_search_scoped_names: the deferrable names its
  enabled/disabled toolsets allow, cached on the agent) are enabled, not
  unavailable. A session that turned the toolset off still reports it.
- platform == "cron" returns None, like delegated children and background
  review.

Probes:
  probe_deferred_tools.py (real AIAgent, default config)
    before: rows for session_search and todo_list; after: no rows
  probe_cron_narrowed.py
    before: tool_name=clarify row; after: no rows
Test test_deferred_builtins_and_cron_narrowing_are_not_unavailable is RED on
base.
2026-09-28 12:43:03 -07:00
teknium1
23b5f37d82 fix(telemetry): an opted-out start purges pending setup and exit markers (m19)
begin_process's collection-off branch purged only parked update receipts.
A provider_setup marker (pid, start_time, surface, provider) or a process
exit marker written while collection was on stayed on disk through the
opt-out and was reported (`abandoned` / `killed`) once collection came back
on, in a period the user had not consented to.

The same branch now removes process_markers/ and provider_setup_markers/.

Probe p2b_off_then_on.sh:
  before: marker survives two OFF starts; the ON start emits
          `abandoned openrouter`
  after:  markers {} after the first OFF start; the ON start emits only its
          own started/completed rows
Test test_opted_out_start_purges_markers_left_while_opted_in is RED on base.
2026-09-28 12:43:03 -07:00
teknium1
41ce3104a7 fix(telemetry): provider_setup counts only a new provider credential or endpoint (M13, m18)
Every PUT /api/env write of a provider-shaped variable recorded
started+completed: an empty value (a clear), a same-value re-save,
GITHUB_TOKEN/GH_TOKEN (-> copilot), HF_TOKEN, and keys the TTS/STT/image
panels also save (GEMINI_API_KEY, XAI_API_KEY, DEEPINFRA_API_KEY). TUI
model.save_key counted re-saves; editing a custom endpoint counted as a new
setup.

- record_api_key_saved takes the value and the stored value read before the
  save: only a non-empty, changed value counts.
- provider_for_api_key_env returns None for ecosystem tokens (GITHUB_TOKEN,
  GH_TOKEN, HF_TOKEN) and any key a TOOL_CATEGORIES panel asks for: a bare
  key save cannot tell a provider connection from a tool setting (the PUT
  body has no provider flag; adding one is Desktop work, see NOT_COVERED).
  Provider pickers (CLI, TUI model.save_key with its explicit slug) still
  count those providers.
- model.save_key skips a same-key re-save.
- upsert_custom_endpoint records only when the endpoint did not exist.

Probe p4_web_double.py (real route functions):
  before: openrouter started x3/completed x3 (save, re-save, clear),
          copilot x1 from GITHUB_TOKEN, custom x3 (create + 2 edits)
  after:  openrouter x1, no copilot row, custom x1
Tests test_web_forms_count_only_a_new_provider_key_or_endpoint and the
extended test_api_key_env_maps_to_provider_only are RED on base.
2026-09-28 12:43:03 -07:00
teknium1
946ae62a89 fix(telemetry): provider_setup Esc is cancelled and Back continues one flow (M12)
cli_provider_setup treated the setup menus' BaseException control flow as
errors: Esc (_SetupCancelled) recorded failed/other, and Back (_SetupGoBack)
recorded failed/other and then a second `started` when the wizard replayed
the provider picker.

- setup_failure_class maps _SetupCancelled to `cancelled`.
- Back leaves the flow open on the thread-local; the replayed picker
  re-entering the same surface+provider continues it (no new `started`).
  Picking another provider, or leaving the entry point
  (provider_setup_surface exit other than a Back), ends it failed/cancelled.

Probe p5_cli_nav.py (real cmd_model):
  before: back_then_pick started x2, failed/other x1, completed x1;
          escape started x1, failed/other x1
  after:  back_then_pick started x1, completed x1;
          escape started x1, failed/cancelled x1;
          back_then_esc_menu started x1, failed/cancelled x1;
          back_then_other openrouter started+cancelled, anthropic started+completed
Test test_cli_setup_navigation_esc_cancels_and_back_resumes_one_flow is RED
on base (failed/other, extra started).
2026-09-28 12:43:03 -07:00
teknium1
bfa3b84c9f fix(telemetry): Desktop closes action/setting ids locally and records first-run steps after a yes
m13: the renderer stored and sent raw plugin palette/keybinding ids, rage-click
targets and config paths below user-named containers (providers.<name>.api_key);
the backend collapsed them, but only after truncating the action list to 200.
- recordAction collapses anything outside the built-in ids (KEYBIND_ACTIONS +
  KEYBIND_READONLY + the named buttons, minus numbered slots — exactly the
  backend's DESKTOP_ACTION_IDS) to `other` before it is counted or used as a
  rage-click target.
- recordSettingsSaved sends only keys the config schema publishes (the Settings
  page passes its schema: DEFAULT_CONFIG leaves), so nothing below a user-keyed
  container leaves or is kept.
- Backend DAILY_ACTION_ROWS_MAX = |DESKTOP_ACTION_IDS| x |vias| (296): every
  distinct collapsed (action, via) fits, so none is cut before collapsing.

m14: first-run steps happen before the consent question and collection defaults
off, so every step but consent/first_message was dropped. Chose the in-memory
buffer (smaller than removing the steps from contract, schema and TS): while the
answer is unknown/undecided, recordOnboarding/closeOnboardingStep calls are held
in memory (max 100, never persisted or sent) and replayed on the switch turning
on; a decided "off", a profile switch or quitting discards them.
setDesktopMetricsGate takes `decided` (hook passes consent.decided).

Tests (desktop-metrics.test.ts): recordSettingsSaved callers pass a schema
(contract change). New invariants, RED on base: plugin action id / rage target
stored and sent as `other`; a providers.<name> key is not sent; pre-consent
onboarding held, sent on yes (with closes applied), dropped after a decided no.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, raw ids):
  before: localStorage actions {"acme-plugin:/home/alice/clients/bigco|palette":1, ...|click:3};
          wire rage_click target "acme-plugin:/home/alice/clients/bigco", setting
          "providers.acme-corp-internal.api_key"
  after:  localStorage actions {"other|palette":1,"other|click":3}; wire
          [rage_click target "other"] only
2026-09-28 12:43:03 -07:00
teknium1
a0508147c1 fix(telemetry): Desktop metric state is per profile and shared by its windows
M9 + the renderer half of M10, and m15.

M9: the renderer kept one install-wide localStorage record and bound the new
gateway before reading its consent, so a day accumulated under profile A was
flushed to B before B's switch was read (or dropped by B's off), and A's area
latch suppressed B's first use.
- The record is keyed per (connection, profile): `hermes.desktop.metrics.v1:
  <FNV hash>` (no profile name in the key). bindDesktopMetrics takes the
  profile scope; a different scope forgets in-memory state and nulls the gate,
  so nothing is kept or sent until that profile's switch is read.
- The hook pins its requester to the focused (connection, profile) with
  requestGatewayForAgent instead of the session-tile router, so the RPCs land
  in the profile whose switch gates them. main receives the same scope hash.

M10 (renderer): each window cached the record in module memory and the last
debounced persist() won, so peer windows overwrote each other's counts. Every
change now reads the record fresh and writes it back (no debounce, no cache;
persistDesktopMetricsNow is gone). withState calls are no longer nested
(noteMessageSent / abandoned-step reporting read first, record after), since a
nested write would be overwritten by the outer one. A daily flush is pinned to
the scope and requester it started with. Both windows may still send the same
finished day; the backend's durable latch (earlier commit) records it once.

m15: consent was re-read only on attach or profile switch. The hook re-reads
it on window focus; Settings › Privacy applies its save to the gate when it is
scoped to the focused profile by name, not only when unscoped.

Tests (desktop-metrics.test.ts): `stored()` now finds the per-profile key and
setEnabled is expected with the scope argument (contract change). New
invariants, RED on base: profile switch A→B→A keeps/sends nothing of A under B
and A still reports its day; two module instances (windows) of one profile sum
into one day.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, adapted to scopes):
  before: B received before its consent read: [["shared_metrics.desktop_daily",{...view.toggleSidebar 4...}]]
          areas latch A: [terminal_pane] B: []
          two windows stored: {"composer.send|shortcut":4} (W2's session.new lost), W2 daily sends
          composer.send 3 + session.new 2
  after:  B received before its consent read: []; areas latch A: [terminal_pane] B: [terminal_pane]
          stored: {"composer.send|shortcut":4,"session.new|shortcut":2}; W1 and W2 send the same
          aggregate (the backend latch keeps one)
2026-09-28 12:43:03 -07:00
teknium1
aa83040ca4 fix(telemetry): renderer-crash consent and pending crashes are per window and profile
M11. RendererCrashRecorder held one install-wide `enabled` flag that every window
overwrote with its own profile's switch, so an opted-out window's crash was
written to disk and reported into another profile, and an opted-out window
wiped the opted-in profile's pending crash.

- Consent is kept per webContents id (the IPC sender), with a hash of the
  window's focused profile key; set-enabled now carries that key.
- A crash is recorded only when the crashed window's own profile collects;
  each pending entry is tagged with the profile hash (file v2: {entries:[{profile,
  reason}]}; an untagged v1 file is ignored — whose it was is unknown).
- take() returns only the calling window's profile's reasons; one claim per
  profile (the same window may re-claim after its renderer reloaded); ack drops
  only the claimed count. Turning a profile off drops only that profile's entries.
- main.ts passes the window's webContents id to recordRendererGone.

Tests: renderer-crash-metrics.test.ts moves to the per-window API (contract
change: every call names its window, set-enabled names the profile); one new
invariant test (opted-out window records nothing, another profile neither
drains nor purges) — the base API cannot express it, so RED is shown by the
reviewer's probe instead.
Probe (review/desktop/jsrun/.../renderer-crash-metrics.probe.test.ts, adapted to
window ids):
  before: after opted-out window crash, file exists: true {"v":1,"reasons":["crash"]}
          take by window 2: {"reasons":["crash"]}; A crash record after B-window off: false
  after:  file exists: false; take by window 2: null; A crash record after B-window off: true
2026-09-28 12:43:03 -07:00
teknium1
4bb58d3d60 fix(telemetry): switching Desktop collection off drops the backend onboarding latches
m16. shared_metrics.set(enabled=false) purged the renderer's onboarding latch
but left telemetry/shared_metrics/desktop_onboarding/<step>.<event> in the
profile, so a content-free per-profile record outlived the opt-out and the two
sides disagreed on re-opt-in. The Desktop's consent RPC now removes that
directory when it turns collection off (shared_metrics_desktop.
purge_onboarding_latches). Opting out from the CLI (config set / setup wizard)
leaves them until the Desktop's switch is next turned off — see NOT_COVERED.

Test: test_turning_collection_off_drops_the_onboarding_latches — RED on base
(latch file survives), green now.
Probe (review/desktop/probe_optout_set.py):
  before: surviving desktop latches: [.../desktop_onboarding, .../intro.reached, ...]
  after:  surviving desktop latches: [] (only metrics.sqlite3 / outbox remain)
2026-09-28 12:43:03 -07:00
teknium1
1a07b7259a fix(telemetry): Desktop daily report and feature use latch durably per profile, in the usage day's period
M10 + m17 (+ the backend half of M9). The backend's once-per-day latches for
hermes.desktop.{action_use,mode_use} (the finished-day report) and
hermes.desktop.feature_use lived in process memory, so a second window's report,
a backend restart or a pooled backend re-counted the same day/area. The daily
rows also went through the relay's current-day marks, so a day flushed a week
late was counted in the flush day's period.

- record_desktop_daily now writes through SharedMetricsStore.update_rollup_state:
  one write transaction checks a `desktop_daily_reported` state row (the usage
  days already reported, bounded to the last 8 days) and records the rows with
  period_start = the usage day. Days outside [today-8, today] settle unrecorded.
- record_desktop_feature_use uses store.record_counter_once_per_day (the same
  durable per-UTC-day-per-dimension latch feature_disabled uses).
- With collection off, desktop_daily answers recorded:false instead of true, so a
  renderer that has not learnt the switch yet never treats that as "settled" and
  drops the day (the client's own gate purges it when it reads off).
- Neither metric has feature/milestone side effects in the subscriber, so
  writing to the store directly is equivalent apart from the latch and period.

Tests: the fixture mocked record_process_marks_saved; the daily/feature-use
assertions now read the profile's store (their contract changed: rows no longer
go through the relay). The disabled-profile test now expects recorded:false.
The store-refusal test raises from update_rollup_state instead of mocking the
relay's saved count. New invariant test (two windows / a restarted backend, all
in-process state reset between reports): RED on base, green now.

Probe (review/desktop/probe_latch.py, run twice on one home = backend restart):
  before: run 2 → action_use composer.send 2, feature_use terminal_pane 2,
          mode_use 2, all in period 2026-09-28 for usage day 2026-09-20
  after:  run 2 → action_use 1, feature_use 1, mode_use 1; action_use/mode_use in
          period 2026-09-20. Opted-out home: daily -> {'recorded': False}.
2026-09-28 12:43:03 -07:00
teknium1
5c805d648f fix(telemetry): one cache_break row per cold read; opted-out enabled() skips the global lock (m1, m5)
What:
- m1: record_known_cache_break emits only when it is the first announcement for the pending cold
  read (the first cause names the break); several causes before one cold read were one row each.
- m5: enabled() while OFF returns before taking the process-wide _RUNTIME_LOCK when no runtime
  exists for the profile (nothing to tear down).

Tests: compression x2 + model_switch before one cold read -> one compression row (RED on the
previous commit: compression 2 + model_switch 1). Existing cache_break test unchanged and green.

Measured (uncontended, fakehome): enabled() OFF 17.72 -> 17.24 us/call. The lock was the
cross-thread serialisation point, not the bulk of the cost: _raw_config() is 13.0 us of it (see
NOT_COVERED.md).
2026-09-28 12:43:03 -07:00
teknium1
876072459e fix(telemetry): engagement reads its day clock under the write lock; busy writes defer instead of failing (M2, m7)
What:
- M2: engagement.record reads the clock inside the rollup write transaction, so writers apply in
  clock order and no process closes a day another already closed. apply() never moves the stored
  day backwards: an interaction that lands behind the stored day folds into it (the only close path
  is a later clock), instead of being dropped or re-opening a day.
- m7: SharedMetricsStore counter and rollup writes that hit the 250 ms busy timeout are kept in a
  bounded in-memory queue (4096) and land with the next write, the next package export, or an exit
  drain (atexit, 5 s busy timeout, registered only after a deferral). The caller still waits at most
  250 ms; a non-busy error drops only the failing write.

Why: the Relay subscriber thread saw "database is locked" and lost the row (engagement under
multi-process contention; model-call counters under host load).

Tests (RED on the previous commit): a write while another connection holds BEGIN IMMEDIATE ->
no raise, lands with the next write (value 2); a late interaction folds into the newer day and the
closed day reports once.

Stress: 20 parallel copies of the existing
test_cross_process_model_call_updates_are_transactional under -n 12 (load avg 6-15, a second
-n 12 suite running): fixed store 10 rounds -> 200/200 passed; previous store 3 rounds -> 55/60
(5 failed: child exit 1, database is locked).
2026-09-28 12:43:03 -07:00
teknium1
8dfdc46067 fix(telemetry): relay-stamped turns are gateway messages; a multi-platform connector's unknown chat stays relay (M5 part, m9)
What:
- M5: `relay` is a gateway surface alias, so an inbound the connector could not stamp is
  entrypoint gateway_message / surface gateway (it counts as engagement) instead of other/none;
  `relay` joins the task/session gateway platform set (it was collapsing to `plugin`).
- m9: RelayAdapter._metrics_platform labels a chat by its inbound platform, and falls back to the
  connector's primary platform only when the socket fronts exactly one identity; _source_platform
  uses it (egress routing keeps _chat_platform unchanged).

Why: the same turn was `other/none` in task/session while delivery/latency said `discord`; an
unknown chat on a multi-platform connector was misattributed to the primary platform.

Tests: the existing unresolvable-chat stub now implements _metrics_platform (renamed hook). New:
task_start_fields({"platform": "relay"}) -> gateway_message/gateway/relay; unstamped chat on a
two-identity connector -> delivery platform `relay`. Both RED on the previous commit
(other/other/none; discord).
Probe (review/engagement/relay/probe_d_coverage.py), before: source.platform='relay' ->
{'entrypoint': 'other', 'execution_surface': 'other', 'platform': 'none'}, engagement surface None.
After: the contract test above.

Not done: resolving the turn-start task platform to the connector's platform (needs the adapter at
agent construction) and M6 (delivery row profile under multiplex); see NOT_COVERED.md.
2026-09-28 12:43:03 -07:00
teknium1
4c7085d284 fix(telemetry): engagement and switch_after count only a person's turns (M1, M3, M7, M8, m6)
What:
- M1: the engagement turn mark (primary model) needs a user turn (pre_llm_call) on an interactive
  surface or a gateway message; API-server, python embedding and curator turns no longer create
  engagement.day rows or outvote the person's model. close_day emits nothing for a day with no
  engaged surface (the root profile's host row with an active-profile count stays).
- M3 (owner ruling): unattended cron runs are dropped from engagement entirely. surface_day always
  carries an active-minutes bucket, so it cannot express cron presence without minutes; hermes.cron.run
  already counts runs. The `cron` enum value stays in the contract for already-stored rows.
- M7: background review forks (no pre_llm_call) no longer advance model_switch_after's turn run.
- M8: the run observes the route each turn was sent on (its first request), so a provider failover
  neither restarts the run nor names the fallback as the model left.
- m6: documented that the root host row (surfaces_used_count 0) is excluded from days-active counts.

Why: Hermes-internal and unattended activity is not user engagement; switch_after measures how long
a person stayed on the model they chose.

Tests (RED on the previous commit): API/cron-only day -> no day row, mixed day -> primary = person's
model; failover + review fork -> switch_after names the selected model with 2_to_3.

Probe (review/engagement/probe_internal.py, switch/probe_switch.py -k "p3c or p7"), before:
  API-only / curator-only day -> engagement.day {"primary_model": "openai/gpt-5", "surfaces_used_count": "0"}
  cron all day -> day {"active_minutes_bucket": "gte_6h", "active_profile_count_bucket": "1"} + surface_day cron
  2 CLI + 5 API turns -> primary openai/gpt-5
  P3c (6 turns on A, 7th failed over to B) -> {"model": "gpt-5", "turns_before_switch_bucket": "1"}
  P7 review fork -> 4_to_10 (routed=False) / gpt-5-mini "1" (routed=True)
after: the equivalent invariant tests pass (no day row, primary = PUBLIC, selected model 2_to_3).
2026-09-28 12:43:03 -07:00
teknium1
b3d4865328 fix(telemetry): link session segments only at the compression hand-off (B2, M4, m2, m8, m11)
What: the metrics lineage no longer keys off the ambient Portal conversation id (it walks every
parent_session_id: gateway reset/idle expiry, /new, /branch). Compression calls
relay_shared_metrics.rotate_segment(old, new) from the committed rotation
(_notify_context_engine_compression_complete) and from a stale agent adopting the live tip
(_adopt_live_compression_child). The hand-off registers the new id in the lineage, moves the
route run and the spent-turn history into it, and closes the old segment now (or when its in-flight
compression turn ends). A rotated-to id closed before serving a turn, or still open at shutdown,
closes too; a 512-lineage backstop flushes conversations a surface never closed.
- m2: /undo or /retry on the new id reaches the turns earlier segments spent (tokens_bucket known).
- m11: /model right after a rotation finds the conversation's turn run.
- m8: background review forks (same session id, no pre_llm_call) add no turns/calls/messages.

Why: one conversation = one session / context_peak / tool_overhead / tool_enabled_unused row, on
time; N gateway conversations collapsed into one row emitted only at process exit, and rotated
segments stayed in memory until exit.

Tests: two existing tests simulated rotation by sharing a conversation context; their contract
changes to the explicit hand-off (test_a_compressed_conversation..., test_context_peak_is_one_row...).
New: reset/branch chain after a compression + tip-first close order, the compression seam, switch/undo
right after a rotation, review fork volume. RED on base (rotate_segment missing / switch_after [] /
turn bucket '2' for one user turn).

Probe (review/engagement/lineage probe_lineage.py -k c2, probe_branch.py; fix-lineage/probe_handoff.py):
before: c2 3 conversations -> session.count 0 until shutdown, then 1; branch -> 1; 50 compressed
conversations -> 0 rows, 50 lineages/sessions live.
after:  c2 -> 3 rows before shutdown (turns 2,1,1), lineages=0; branch -> 2; 50 compressed -> 50 rows,
0 lineages, no sessions; tip-first close with the old turn running -> 0 then 1; shutdown with an
open tip -> 1 per conversation.
2026-09-28 12:43:03 -07:00
teknium1
ac372200b1 fix(telemetry): network-address model ids collapse; subscriber re-runs the catalog (m12)
What:
- shared_metrics_catalog.model_metric_name also collapses a model id that is
  a network address: a loopback/IPv4 host or localhost at the start, or a
  leading host:port with a 2-5 digit port (so Bedrock's "...-v1:0" and
  OpenRouter's ":free" / Ollama's ":120b" tags stay readable). The existing
  URL/path/weight-file rules move into the same compiled pattern.
- shared_metrics_contract._bounded_dimensions (every mark -> counter
  projection) and model_token_counters re-run provider_metric_name /
  model_metric_name on the mark's provider/model identifier fields and drop
  the row when either would rewrite it ("none" = unset provider passes).
  Record-time only: counter_dimensions_are_valid (also used at packaging) is
  unchanged, so catalog drift can never raise out of _package_metric and
  block a package of rows already stored.

Why: `127.0.0.1:8080/x` on a public provider passed verbatim through every
route metric, and the subscriber validated identifier shape only, so a mark
that skipped the producer's catalog pass stored custom:acme, ollama/<model>,
unknown vendors and even http:// URLs.

Probe (fix-catalog/probe_b1_m12.py):
  before: route(openrouter, 127.0.0.1:8080/x) -> model '127.0.0.1:8080/x'
          (also localhost/qwen, 10.0.0.5/x, gpu-box.lan:8000/qwen verbatim)
          subscriber injected switch_after {127.0.0.1:8080/x, openrouter} -> stored
          subscriber injected switch_after {acme-secret, acmecorp-internal} -> stored
  after:  all of the above -> model 'custom'; both injected marks -> None;
          bedrock anthropic.claude-3-5-sonnet-20241022-v2:0, openai/gpt-4o:free,
          gpt-oss:120b and the legit anthropic mark unchanged.

Test: test_network_address_model_ids_and_raw_injected_marks_never_reach_counters
(RED with the catalog reverted: model_route kept 127.0.0.1:8080/x; RED with the
contract reverted: the injected host:port mark was stored; GREEN with both).
2026-09-28 12:43:03 -07:00
teknium1
802e6797a0 fix(telemetry): provider allowlist comes from shipped providers only (B1)
What: shared_metrics_catalog.provider_names() no longer reads the live
PROVIDER_REGISTRY / models._KNOWN_PROVIDER_NAMES. It is now the built-in
auth rows (BUILTIN_PROVIDER_IDS), HERMES_OVERLAYS, the three static alias
tables, openrouter/custom, and the names+aliases of profiles registered
from the in-tree plugins/model-providers dir (providers._SOURCES ==
"bundled", process-wide layer, independent of the bound profile home).

Why: auth_plugin_providers mirrors every $HERMES_HOME (and pip) provider
plugin, plus its aliases, into PROVIDER_REGISTRY and the picker labels, so a
user-chosen provider name was treated as shipped and its model id was not
collapsed either. Every consumer goes through provider_metric_name, so the
one fix covers model_route and every per-model v4/v5 metric (tokens,
context_peak, friction, tool_quality, loop_guard, tool_recovery,
model_reply_issue, task_cost, wasted_tokens, cache_break, switch_after,
tool_unavailable, engagement.day), provider_setup (all surfaces, incl. the
PUT /api/env key path), the on-disk pending marker, setup.completed,
model_switch/fallback and install.snapshot main_provider. Pip-installed
provider plugins now read custom too (third-party, not shipped); doc says so.

Probe (fix-catalog/probe_b1_m12.py, user plugin acmecorp-internal alias
acme-llm):
  before: provider_metric_name('acmecorp-internal') -> 'acmecorp-internal'
          route: {'model': 'acme-secret-model-v2', 'provider': 'acmecorp-internal'}
          marker on disk: {..., "provider": "acmecorp-internal"}
  after:  provider_metric_name('acmecorp-internal') -> 'custom' ('acme-llm' too)
          route: {'model': 'custom', 'provider': 'custom'}; marker "provider": "custom"
          install main_provider custom; switch/fallback/setup custom
          shipped plugin deepinfra: {'model': 'meta-llama/...', 'provider': 'deepinfra'}

Test: test_user_provider_plugin_name_and_model_never_leave (RED on base:
marker carried "acme-llm"; GREEN after).
2026-09-28 12:43:03 -07:00
teknium1
64480b02c8 docs(telemetry): consent copy names the v5 usage categories
CLI setup and the Desktop consent dialog add one 'how Hermes gets used'
line covering the v5 harness, efficiency, engagement, Desktop feature/
dislike and provider-setup counters, and say setting values never leave.
2026-09-28 12:43:03 -07:00
teknium1
26894219dd style: drop trailing blank line in shared_metrics_harness 2026-09-28 12:43:03 -07:00
teknium1
c13a4abc8d feat(telemetry): v5 signals — tool_unavailable, provider_setup, feature_adoption, feature_disabled
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in
  the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality.
- hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface
  (cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow;
  a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern).
- hermes.feature_adoption.count: once per feature per install at first real use, derived from existing
  counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the
  owning profile's install age.
- hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the
  re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day.
  Surface comes from the user entry point; setup/migrations record nothing.
2026-09-28 12:43:03 -07:00
teknium1
98fa444d16 feat(desktop): wire love/hate telemetry into existing interactions
No new UI: counters hang off interactions the app already has.

- feature_use: panes (terminal/file/browser/review), overlays (palette,
  model/session pickers, switcher, find), voice, Bot Mode (the `bots`
  workspace), skins, projects, full-page routes and each settings view.
- action_use: keybinding dispatch (shortcut), palette items (palette),
  titlebar/sidebar tools that carry a keybinding id, composer send/stop,
  voice, message copy/retry (click, via a typed const table), and the
  native Open Folder menu (menu). Aggregated per day in the app and sent
  once per finished day.
- mode_use: Sessions vs Bot Mode from $workspaceMode / the tile's
  workspace scope; active time = interaction gaps <= 5 min.
- friction: user toast dismissals by closed notice id, error toasts by
  their code-defined summary category, backend drops (settled 3s so a
  Restart/switch cancels), long frames while visible, renderer crashes
  drained from main.
- dislike: quick close (<5s), cancelled flows, settings saves (keys
  only), rage clicks, undo (restored draft, reopen closed tab), shipped
  features switched off.
- onboarding: transitions of the existing first-run stores.

All of it follows the focused profile's collection switch: off (or not
yet known) means nothing is kept in localStorage and nothing is sent,
and switching off purges the local record and tells main to drop its
crash file.
2026-09-28 12:43:03 -07:00
teknium1
0fce280ee9 feat(desktop): consent-gated renderer-crash recorder in the main process
A renderer crash takes the renderer (and its backend socket) with it, so
the renderer cannot report it. Main now observes render-process-gone on
live windows (never user close or intentional teardown), persists only
the bucketed reason (crash/oom/killed/other) under userData while the
renderer has reported the user's opt-in, and hands the list to the
reloaded renderer via take/ack IPC, mirroring the pending update-run
record. Turning collection off deletes the pending file immediately.
2026-09-28 12:43:03 -07:00
teknium1
6c143c238f feat(metrics): hermes.desktop.* contract, schema and gateway RPCs
Desktop love/hate telemetry needs a backend half: six counters
(feature_use, action_use, mode_use, friction, dislike, onboarding) with
closed dimension sets, the v3 JSON schema entries, the public doc rows,
and profile-scoped gateway RPCs the app reports through.

Every value the app sends is collapsed server-side onto the closed set
(unknown -> other); feature use latches per area per UTC day, friction
and dislike are capped per day, the daily report (action counts + mode
time + bot count) latches per day and answers recorded=false on a store
failure so the app keeps the day for a retry. For setting changes the
app sends only the config key: the backend reads the saved value and
compares it to DEFAULT_CONFIG itself, so values never cross the wire;
keys outside DEFAULT_CONFIG read other. All helpers check enabled()
before doing any work.
2026-09-28 12:43:03 -07:00
teknium1
928daa63af feat(telemetry): v5 engagement rollup, turns-before-switch, per-conversation sessions, relay source platform
Engagement (local daily rollup, emitted once the UTC day closes, exactly once per profile):
- hermes.engagement.surface_day.count {surface, active_minutes_bucket} per surface used
  (cli/tui/desktop/gateway/acp/cron); active time = inter-interaction gaps capped at 5 min.
- hermes.engagement.day.count {active_minutes_bucket, surfaces_used_count, primary_provider,
  primary_model, active_profile_count_bucket}; primary model = most attended turns that day.
  Active profiles are folded into one host accumulator in the root profile's store (opaque local
  hashes); only the root profile's row reports the count, other profiles report 0.
- Days-active-per-week / return-by-model are derived server-side; no weekly state, no new id.

Model satisfaction:
- hermes.model_switch_after.count {provider, model, turns_before_switch_bucket}: turns the model
  being left served in the conversation, from every /model site (CLI, TUI, ACP, gateway).

Decisions:
- hermes.session.count counts once per conversation across compression rotation (merged over the
  v4 conversation lineage), and gains message/model-call/tool-call count buckets with a longer
  tail (251_to_1000, gte_1001); rows recorded before the upgrade still package.
- Relay-delivered hermes.gateway.reply_latency and hermes.platform.delivery report the originating
  platform instead of `relay`.
2026-09-28 12:43:03 -07:00
teknium1
607dff390a feat(telemetry): v5 efficiency metrics — turn cost, wasted tokens, tool output truncation, tool overhead, prompt-cache breaks
Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):

- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
  outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
  children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
  calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
  emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
  later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
  result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
  shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
  execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
  used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
  system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
  changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
  warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
  A Hermes-known break suppresses the miss it causes. No prompt hashing.

Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.
2026-09-28 12:43:03 -07:00
teknium1
c5555efca8 test(smoke): expect the per-response model_reply_issue counter
Every primary model response now records hermes.model_reply_issue.count
(issue=none for usable replies), so the smoke turn's two scripted replies
add one row with value 2; the exact-metric-set assertions must include it.
2026-09-28 12:43:03 -07:00
teknium1
8d8836ddb1 feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):

- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
  matcher reports the strategy that landed (or no_match / ambiguous) into a
  context-local probe opened only around the patch/write_file handlers, so we
  learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
  the guardrail already warns/blocks/halts and at turn end for the iteration
  budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
  next_outcome. One row per failed tool call, resolved against the model's
  next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
  a table lookup on the first program word, never the text. Hermes' own
  deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
  backend result, so a command's own `exit 124` reads nonzero, not timeout.
  Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
  truncation only from the structured finish reason; empty/reasoning_only
  from the normalized message; issue=none per response as the denominator
  (model_route counts attempts and files empty replies as failures).

Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
2026-09-28 12:43:03 -07:00
teknium1
c4fe596783 style: split the begin_process one-liners (E702)
The v4 process-exit hooks put an import and a call on one line with a
semicolon in hermes_cli/gateway.py and hermes_cli/main.py.
2026-09-28 12:43:03 -07:00
teknium1
62914682ef refactor(desktop): move shared-metrics wiring out of main.ts
The v4 startup-latency IPC and packaged-update recorder added ~26 lines to
main.ts (18k lines). desktop-shared-metrics.ts now owns the claim IPC, the
recorder construction and its renderer IPC, and the packaged-only apply
wrapper; main.ts keeps one import, one constructor call and two call sites.
The recorder's behaviour is unchanged (update-metrics.ts untouched).
2026-09-28 12:43:03 -07:00
teknium1
88eb76f02e fix(metrics): startup latency counts once per client launch, never after an in-place exec
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat
  call added a row. The TUI and Desktop now send an opaque per-launch id
  (TUI: one per process; Desktop: minted when the once-per-launch main-process
  claim succeeds) and the backend claims (surface, launch_id) once per
  process: a reconnect re-sending the same launch counts once, a new Desktop
  launch against a long-lived backend still counts. Legacy clients without an
  id count once per backend process and surface. The id is only a local latch
  key (bounded set, length-capped), never recorded. Only a usable measurement
  spends the claim. Contract + apps/shared regenerated.
- m10: relaunch() execs in place, keeping the PID, so `sessions browse` ->
  resume counted the picker time as CLI startup. relaunch() stamps its PID in
  HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it
  matches, while children (other PIDs) inheriting the env still count.
- m8 (psutil half): record_process_ready checks the opt-in gate before reading
  the process start time (psutil) or starting a thread. The claim is taken
  first so a later opt-in never records a stale "startup" mid-session.

Test changes to existing assertions: the TUI/Desktop param assertions now
expect launch_id (a new wire param); the RPC surface test sends distinct
launch ids so its three calls stay three launches under the new latch.
2026-09-28 12:43:03 -07:00
teknium1
f6f0d47d2b fix(metrics): local_model_provider_used sees bare aux endpoints and bare-named providers
Two false negatives in the install snapshot's local-server flag (review m11):
- an auxiliary slot under provider auto / no provider with base_url + api_key
  was skipped as "inherits main", but agent.auxiliary_client routes exactly
  that shape to the endpoint as custom; the filter now mirrors the runtime;
- a providers:/custom_providers: entry referenced by bare name (provider:
  home-llama) was only resolved for the custom:<name> spelling. The lookup
  now goes through _get_named_custom_provider for any non-inheriting id, so
  built-ins still shadow entries exactly as at runtime.
2026-09-28 12:43:03 -07:00
teknium1
f5e0a521f4 fix(tui): dashboard Chat-tab TUIs never report startup latency
The dashboard's in-browser Chat tab spawns a TUI per terminal it opens
(HERMES_TUI_DASHBOARD), so every tab counted as a TUI launch and skewed the
tui startup distribution toward "tab opened", not "user launched Hermes".
Skip the report when DASHBOARD_TUI_MODE is set (review finding M9).
2026-09-28 12:43:03 -07:00
teknium1
f7b5ea6f7a fix(telemetry): loop metrics count only the user's real backend work and curator passes
Review findings M7, M8, m6, m7 and the browser half of m8 on the v4 loop metrics.

- M7: TUI/Desktop @-path completion on a non-local backend lists the directory through
  terminal_tool, so every keystroke added a hermes.execution_backend.count row as if the user
  ran a command. Hermes-owned calls now run inside shared_metrics_loop.unmetered_backend_calls()
  (a contextvar checked in record_execution_backend) instead of a flag threaded through
  terminal_tool. It was the only non-_host_local internal terminal_tool caller (bot DM runners
  already use _host_local; the prompt-builder probe calls env.execute directly).
- M8: a scheduled curator tick that finds the run claim held recorded scheduled/skipped every
  tick (5 ticks -> 5 rows for one pass). The holder's own pass is the one counted; the held
  branch now records nothing. Real dry runs still report skipped.
- m6: a foreground timeout returns exit 124 with partial output and no error field, so it was
  recorded as success. exit 124 now records outcome=failed, error_class=timeout.
- m7: vercel_sandbox (and managed_modal) are built-in backends in
  tools/terminal_tool_config._BUILTIN_BACKENDS but collapsed to other. Added to
  TERMINAL_BACKENDS and both schema enums; a test pins every built-in backend to its own bucket.
- m8 (browser): record_browser_call resolved the browser backend on every call even with
  collection off. The resolver is now passed lazily and only runs inside the gated builder.

The existing test asserting a held claim records scheduled/skipped encoded M8 itself and is
replaced by one asserting held ticks record nothing.
2026-09-28 12:43:03 -07:00
teknium1
85ab6b510b fix(telemetry): relay reply latency, secondary-profile disconnects, side-effect-free classification
- Reply latency for relay-delivered turns: the clock starts under the
  inbound's source platform (discord) but the reply leaves through the
  Platform.RELAY adapter, so the stop never found it and the row was never
  recorded. A relay stop now claims the chat's single pending clock.
- Multiplexed secondary-profile fatal errors now record the disconnect
  (reconnects already were), bound explicitly to the owning profile's home:
  the adapter's fatal notification can arrive on a task scoped to the launch
  profile. An unresolvable profile records nothing rather than a guess.
- relay_disabled (the user's relay opt-out revoking its credential) is not a
  lost connection; it no longer counts as a disconnect on either handler.
- Connect/delivery error classification reads stored attributes only
  (inspect.getattr_static), never a property: websockets' deprecated
  ConnectionClosed.code warned and, under -W error, the DeprecationWarning
  replaced the adapter's own exception at the re-raise. Classification also
  moved onto the metrics worker, so nothing it does can reach the caller.

Invariant tests (each red at d2f70bfaa08b): relay-delivered turn records
one latency row per turn (stream and final paths); fatal handlers count
lost connections in the owning profile only and skip relay_disabled; a
connect failure whose exception has a warning/raising status property
reaches the caller unchanged under warnings-as-errors.
2026-09-28 12:43:03 -07:00
teknium1
49aeaa1526 fix(telemetry): name every gateway adapter platform from the Platform enum
hermes.platform.health / delivery / reply_latency built their core platform
vocabulary from hermes_cli.platforms.PLATFORMS (the setup wizard list), which
has no relay, msgraph_webhook or local entry, so those first-party adapters
were reported as the anonymous "plugin" bucket. The adapter vocabulary now
comes from gateway.config.Platform's declared members (loaded on first use:
gateway.config is too heavy for the contract's import); plugin pseudo-members
are not declared members, so unknown plugin platforms still collapse to
"plugin" and catalog/bundled plugin names keep working. The v3 task platform
vocabulary (GATEWAY_PLATFORMS) is unchanged.

Invariant test: every declared Platform member is named as itself, is
contract-valid for the three adapter metrics and validates against the v3
JSON schema counter definitions (red at d2f70bfaa08b: relay/local -> plugin).
2026-09-28 12:43:03 -07:00
teknium1
287564fedd fix(telemetry): update runs counted once, parked only while opted in, recovered files kept until saved
Three defects in how hermes.update.run/stage and hermes.process.exit rows
are recovered from disk:

- Opt-in (M1): the pre-pull park path only checked that the telemetry dir
  existed, so it wrote the WHOLE receipt (argv, step detail) while the user
  was opted out, and a later opt-in counted that run. It now reads consent
  through the already-loaded hermes_cli.config (this interpreter predates
  the checkout swap, so nothing may be imported), parks only when
  collection is on, parks only the fields update_receipt_fields reads (no
  argv, no step text), and purges pending_updates when collection is found
  off. begin_process purges them too on a start with collection off.
- Double count (M2): the completion child records the receipt and the
  parent's boundary finalize (run_completion returned receipt=None) parks
  the same update_id again, so one run became (failed, deps) + (success,
  none). record_update_receipt now claims the update_id with an O_EXCL
  latch under shared_metrics/recorded_updates (last 64 kept); the second
  finalize records nothing.
- Loss (m3): dead-process markers and parked receipts were unlinked before
  their row reached the store; "database is locked" under concurrent starts
  lost them (17/20 rows in the reviewer's race). The claim-by-rename stays,
  but the file is deleted only after record_process_marks_saved confirms
  the rows settled: rows carry a random commit ticket in event metadata
  (allowlisted in the contract), the subscriber tallies tickets whose rows
  persisted, and the reporter flushes the Relay subscribers and compares.
  Otherwise the file is renamed back for the next start. A parked receipt
  whose rows partly landed is not retried (that would double count).

Also: _epoch returns None for a timestamp float() cannot represent
(10**400 raised OverflowError).
2026-09-28 12:43:03 -07:00
teknium1
e50614ed39 fix(telemetry): graceful os._exit stops stamp the process exit marker clean
Every graceful gateway stop was reported by the next start as
hermes.process.exit exit_kind=killed: the gateway always leaves through
_exit_after_graceful_shutdown -> os._exit, which skips the atexit
stamp_exit("clean") that begin_process registered, so the marker still
said "running" and a dead "running" owner reads as killed.

The same shape existed on the two other deliberate os._exit exits: the TUI
gateway's grace timer (_hard_exit, only ever armed by a termination
signal, i.e. a requested stop) and the serve parent-death watchdog (the
designed self-reap once the owning Desktop is gone). Each now stamps
"clean" right before os._exit. Crash and watchdog paths are untouched:
the excepthook's "crash" and shutdown_watchdog's "watchdog" stamps win
because stamp_exit only restamps a marker still in "running".

Live: a real `hermes gateway run` in a scratch home stopped with SIGTERM
now reports (gateway, clean); SIGKILL still reports (gateway, killed).
2026-09-28 12:43:03 -07:00
teknium1
2c9b79a5a3 fix(telemetry): one context_peak row per conversation across compression
Rotating compaction continues a conversation under a new session id, and
each metrics session emitted its own context_peak at close, so one
conversation produced a row per segment (the fuller one plus a low-fill
one), skewing the per-model distribution. The agent turn already
publishes the lineage root (the Portal conversation id); segments sharing
it join one lineage, contribute their peak when they close, and the
merged peak (fullest fill, limit hit anywhere) is emitted once when the
last open segment closes, whatever the close order. Delegated children
share their parent's root but stay their own conversation.

Review finding m2.
2026-09-28 12:43:03 -07:00