95 Commits

Author SHA1 Message Date
ethernet
c13ea774e6 refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.

Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).

Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
2026-09-23 11:41:01 -04:00
teknium1
148d1f6da5 fix(relay): shared-metrics task close skips a scope buried under a concurrent turn
Two concurrent Hermes turns in one session push two task scopes onto the same
physical stack. _finish_task popped its handle unconditionally, so the first turn
to finish hit the native "scope handle is not at the top of the stack" and logged a
full traceback on every overlap (#115471). pop_relay_scope_if_top compares the
handle to the current top inside the task context and skips the pop when a
sibling scope sits above it; the orphan is reclaimed by the existing session-close
drain in _close_scope_handle. Draining the sibling from the task close (as the
candidate PR did) would pop a LIVE turn scope, and guarding pop_relay_scope
itself would break _pop_with_drain, which relies on that raise.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-19 23:42:20 -07:00
teknium1
2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1
398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
Ben Barclay
5a1246f830 fix(observability): attribute ACP and batch execution surfaces
Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:

1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
   but "acp" was absent from EXECUTION_SURFACES, so the contract's
   closed-schema fallback folded every editor session into "other" --
   the bucket meant for genuinely unclassifiable traffic.

2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
   "platform" entirely, so every batch task run reported "unknown"
   despite "batch" already being a first-class surface.

Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".

Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
  entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
  self.platform = "batch" on the runner, and defaulted at the worker call
  site so callers that build a config without it stay attributable

Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.

Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
  M1 revert acp from EXECUTION_SURFACES  -> 3 failed
  M2 revert acp entrypoint mapping only  -> 1 failed
  M3 revert batch passthrough            -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
2026-09-09 11:27:36 +10:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
ecf760db96 simplify(compat): plugins/memory — drop 15 re-exports + 1 alias, repoint 13 test callers
hindsight: drop _PORT_HEALTH_GRACE_ENV/_sanitize_bank_segment facade re-exports (tests -> .embedded/.settings).
honcho: drop 5 tool-schema re-exports, _credential_fingerprint (-> client_cache), _is_auth_error (-> session_auth),
and the _redact_tokens alias (callers renamed to redact_tokens).
openviking: drop 5 _setup re-exports; tests call openviking_module._setup.* directly.
2026-09-03 13:04:16 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
2e36381d85 refactor(hermes_cli): relay_shared_metrics — fold finish_task closing guard 2026-09-02 23:53:49 -07:00
Teknium
a4167f3380 refactor(hermes_cli): shared_metrics store — inline pending-period count 2026-09-02 23:51:33 -07:00
Teknium
aa0eb8b04f refactor(hermes_cli): relay_shared_metrics — shared _run_task_hook for the two boundary entrypoints 2026-09-02 23:49:08 -07:00
Teknium
66f7d5d5d9 refactor(hermes_cli): relay_shared_metrics — merge _task_for into _task_pair, single update_model_call handler 2026-09-02 23:46:27 -07:00
Teknium
38ecfc5fd1 refactor(hermes_cli): shared_metrics_contract — _scoped_dimensions for scope-end projections 2026-09-02 23:42:03 -07:00
Teknium
e7ac74aa12 refactor(hermes_cli): relay_shared_metrics — single _raw_config() reader 2026-09-02 23:38:11 -07:00
Teknium
e1946db88d refactor(hermes_cli): shared_metrics store — reflow remaining SQL/dict literals (trace-identical) 2026-09-02 23:34:26 -07:00
Teknium
3920fa1f06 refactor(hermes_cli): relay_shared_metrics — _elapsed_ms helper, tighter approval/model-call lookups 2026-09-02 23:31:52 -07:00
Teknium
4a78fe794d refactor(hermes_cli): shared_metrics store — compact package insert and gating 2026-09-02 23:28:56 -07:00
Teknium
738dff9dff refactor(hermes_cli): shared_metrics store — hoist counter upsert SQL, flatten active-install guard 2026-09-02 23:23:49 -07:00
Teknium
0361a1f45c refactor(hermes_cli): shared_metrics_sender — drop never-varied only_if_pending flag, flatten identity check 2026-09-02 23:21:38 -07:00
Teknium
1acd3f165f refactor(hermes_cli): relay_shared_metrics — collapse guard ladders in session/turn lookup 2026-09-02 23:18:42 -07:00
Teknium
28f1c872ab refactor(hermes_cli): relay_shared_metrics — hoist pure tool-identity matchers, _scope_handle 2026-09-02 23:07:05 -07:00
Teknium
59b6c305dc refactor(hermes_cli): relay_shared_metrics — module-qualified contract helpers, shared _release() teardown 2026-09-02 23:02:47 -07:00
Teknium
1b4ca18f52 refactor(hermes_cli): relay_shared_metrics — _terminal_flags/_task_parent_handle helpers, compact runtime cache 2026-09-02 22:56:23 -07:00
Teknium
b4cea2dbbd refactor(hermes_cli): shared_metrics_contract — compact dimension tables and signatures 2026-09-02 22:48:37 -07:00
Teknium
6d02842d36 refactor(hermes_cli): shared_metrics store — schema DDL as ordered module tables; fold 5 one-call schema helpers 2026-09-02 22:42:44 -07:00
Teknium
ad079a5f43 refactor(hermes_cli): shared_metrics_sender — dataclass _Response, compact rationale comments (invariants kept) 2026-09-02 22:27:25 -07:00
Teknium
17650193b7 refactor(hermes_cli): relay_shared_metrics — drop write-only _ModelCall.retry_ordinal and unread _ToolCall.task_id 2026-09-02 22:23:49 -07:00
Teknium
9cda9c1f97 refactor(hermes_cli): shared_metrics_contract — kwargs-driven shape matcher, reflowed allowlist constants (values unchanged) 2026-09-02 22:19:53 -07:00
Teknium
adcb4704d4 refactor(hermes_cli): relay_shared_metrics — _forget() for owner-checked index eviction 2026-09-02 21:58:25 -07:00
Teknium
69c419e525 refactor(hermes_cli): shared_metrics store — reflow SQL literals (whitespace-only; normalised SQL trace identical) 2026-09-02 21:49:44 -07:00
Teknium
9316e27edc refactor(hermes_cli): relay_shared_metrics — _abort_session unifies close/deactivate teardown 2026-09-02 21:42:15 -07:00
Teknium
40434ff3cf refactor(hermes_cli): relay_shared_metrics — _admits() gate, shared scope-stack runner, docstring compaction 2026-09-02 21:33:20 -07:00
Teknium
50ed27f263 refactor(hermes_cli): relay_shared_metrics — fold model-call error/end into one updater, _task_pair lookup 2026-09-02 21:15:27 -07:00
Teknium
c9c13fa8f3 refactor(hermes_cli): relay_shared_metrics — _mark()/_guarded() collapse repeated emit and try/except-log blocks 2026-09-02 21:00:46 -07:00
Teknium
1f54585e5b refactor(hermes_cli): observability subscriber/send_config/__init__ — projection table, inline safe-observe, tighter docs 2026-09-02 20:53:57 -07:00
Teknium
6837d2e41b refactor(hermes_cli): shared_metrics_sender — reuse store time helpers, unify defer paths, txn ctx manager 2026-09-02 20:48:41 -07:00
Teknium
5bd425061c refactor(hermes_cli): shared_metrics_contract — table-driven shape match, lifted lookup tables, next()-based bucket 2026-09-02 20:41:00 -07:00
Teknium
da5cc8d1fe refactor(hermes_cli): shared_metrics store — _write() txn ctx, resource-column table, private-path helper 2026-09-02 20:33:40 -07:00
Teknium
4b121f6124 refactor(hermes_cli): compact relay_shared_metrics — event-text helper, _sole() dedupe, derive HANDLED_HOOKS from dispatch table 2026-09-02 20:27:47 -07:00
Jeffrey Quesnelle
48c0c3a873 Merge pull request #96633 from afourniernv/chore/relay-0.8.0
fix(relay): upgrade to 0.8.3
2026-09-02 22:03:27 -04:00
Teknium
cfa327e5dc refactor(hclib): remaining hermes_cli library modules — dead code, unified helpers, flattened branches 2026-09-02 15:03:42 -07:00
Alex Fournier
31e7237adb fix(relay): upgrade to 0.8.1
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-28 08:28:21 -04:00
Ben Barclay
a69a9c351d feat(telemetry): transmit the stable install_id as-is
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.

Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
  derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
  role (unreadable/non-object/id-less payloads still reject rather than
  block the queue) and now records the raw install_id in
  sent_install_id; _body rewrites the payload's install_id from that
  frozen column, keeping byte-identical resends anchored to one
  recorded value.

Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.

Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.

Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.

258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
2026-08-28 15:22:16 +10:00
Ben Barclay
4bdabb21ed fix(telemetry): renew the claim atomically before every POST
Seventh review found the claim-token fix incomplete, and its
reproduction is exact: the pre-POST check was READ-ONLY. A claimant
whose lease expired while suspended still passes it when it wakes
BEFORE anyone reclaims - its token is still in the row - and then a
second process legitimately reclaims while the first one's POST is in
flight. Both send. Reproduced at 60addb16e2: posts ['B', 'A'], both
reporting 'sent'. This is the check-to-POST expiry race, not the
documented mid-POST residual: A's lease was already dead before its
authority check passed.

The check is now an atomic RENEWAL (single CAS UPDATE): it requires the
token to match, the row to be pending, AND the current lease to be
unexpired, and only then extends next_attempt_at a fresh lease into the
future. rowcount == 1 is the only grant. A claimant that wakes past its
own lease fails the unexpired condition and yields even though its
token was never replaced - expiry alone means another process may
claim at any moment, so waking stale is disqualifying regardless of
whether anyone has taken the row yet. The renewed lease (300s) covers
the POST (30s timeout) with margin, and renewal runs before every
retry, not just the first attempt.

Regressions: the reviewer's exact ordering (expired wake before any
reclaim -> zero POSTs, row stays claimable), plus a healthy-claimant
renewal test. Mutation-checked: dropping the lease-unexpired condition
or the token condition each fails the suite.

The at-least-once scope note on _send_one stands: a suspension landing
mid-POST remains client-unfixable; the fixable window is now closed on
both sides (before the check, and between check and POST).

277 tests pass; ruff + footguns clean; staging E2E 202.
2026-08-27 15:21:30 +10:00
Ben Barclay
60addb16e2 fix(telemetry): fence send authority on a per-claim token
Responds to the independent PR review (andrexibiza). Both P1s were
checked against current HEAD rather than taken on authority - the
review was written against 613849c190, before the interval-model
consent replacement landed.

P1-1 (same-UTC-day revoke/re-enable releases refused data): already
fixed by the interval model. The reviewer's exact reproduction - opt in
06:00, revoke 12:00, package collected 18:00, re-enable 20:00 same day
- was re-run at HEAD: the off-window package stays local, and a full-day
aggregate straddling the revocation boundary also stays local (period
containment, timestamp precision). The consent-windows harness already
pins both. The reviewer's related ask that consent-ledger persistence
failures fail closed also holds structurally now: reconciliation derives
state rather than recording transitions, so a lost write means a shorter
confirmed horizon - less is released, never more.

P1-2 (lease has no owner) was VALID at head. Reproduced exactly as
described: A claims, is suspended past the 300s lease, B reclaims and
POSTs, A resumes and POSTs again - and the ingest key is minute-
prefixed, so the duplicate lands as a DISTINCT stored object, making
this worse than a benign idempotent overwrite.

Fix: every claim now mints a claim_token (additive nullable column,
schema version unchanged). Ownership is revalidated immediately before
every external POST, and every settlement, rejection, and backoff write
is compare-and-set on (package_id, claim_token, pending). A lapsed
claimant that resumes yields without transmitting, and its stale
backoff cannot move next_attempt_at under the live claim's lease.

Two deterministic regressions ship with it: expiry -> reclaim -> resume
(the reviewer's schedule), and the subtler stale-backoff-clobber case.

Honest scope, documented on _send_one: delivery remains at-least-once.
The token closes the claim->POST gap; a suspension landing mid-POST
(bytes already on the wire) is not client-revocable. The residual
duplicate is byte-identical content; collapsing it fully needs
package_id-keyed dedupe at the ingest service.

275 tests pass; ruff + footguns clean; staging E2E 202.
2026-08-27 15:04:02 +10:00
Ben Barclay
67d152bc7e fix(telemetry): bound forward-clock damage to the consent horizon
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.

The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.

Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
  call. Honest heartbeats never bind it; a machine off for months
  catches up in a few hook fires (fail-closed latency only); one insane
  sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
  Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
  raw stamp lets an honest clock at revoke time pull a poisoned horizon
  back to the true revoke moment. A rolled-back clock at close time
  only closes earlier - fail-closed.

Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
  (the harness re-implemented the insert; deleting the production line
  survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
  skip was dead code - the store constructor creates the directory
  before the exists() check ran. The probe now checks the default path
  without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
  upgrade (fail-closed; deliberate).

New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.

273 tests pass; ruff and windows-footguns clean; staging E2E 202.
2026-08-27 13:30:25 +10:00
Ben Barclay
5e380d95ba refactor(telemetry): replace consent day-stamp with explicit intervals
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).

Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).

Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
  confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
  clamps window OPENS so a rolled-back clock cannot slide a window
  under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
  only for observed time, so a hand-edited config with no process
  running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
  intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
  undelivered backlog from the earlier consented window (round-5 D4).

The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.

Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).

Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).

Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).

270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
2026-08-27 10:42:31 +10:00
Ben Barclay
613849c190 fix(telemetry): close the consent window on the config transition
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.

BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.

The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.

BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.

Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
  with `if True` survived the entire suite, because every non-http case
  targeted a REMOTE host where the loopback branch rejects anyway. Only a
  non-http scheme on loopback distinguishes the two. Shipped behaviour was
  already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
  Measured against 11 real packages: resource is a stable low-entropy
  tuple and periods are contiguous across a rotation, so for a RARE
  configuration those can bridge windows. The honest claim is that
  rotation raises the cost, not that it makes correlation impossible.

Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.

An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.

259 tests pass. Staging E2E re-run: both packages 202.
2026-08-27 09:47:47 +10:00