The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.
Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.
Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
agent/secret_scope.py has pointed at docs/design/multiplexing-gateway.md
("Workstream A") since the fail-closed secret scope landed, but the file was
never added. This writes the missing doc from the code as it stands today:
the mode flag, scope composition (_profile_runtime_scope seams), the
context-local secret scope and HERMES_HOME override, routing/serving/
persistence/session-lane isolation, the control-plane RPCs, failure modes,
and an honest table of what is still process-global (per-profile MCP
registries are tracked in #67605).
Doc-only change; no code touched. Style follows docs/profile-routing.md.
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.
Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.
Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".
Refs #90837, #90950, #97940, #89332, #45383
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
provider label never carries the vendor token to the alias host (#28660);
route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.
Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
role (unreadable/non-object/id-less payloads still reject rather than
block the queue) and now records the raw install_id in
sent_install_id; _body rewrites the payload's install_id from that
frozen column, keeping byte-identical resends anchored to one
recorded value.
Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.
Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.
Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.
258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
387 commits from main; no conflicts (verified with merge-tree before
merging). Overlap limited to hermes_cli/config_defaults.py and
hermes_cli/setup.py, both auto-merged; all shared-metrics surfaces
untouched by main.
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.
The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.
Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
call. Honest heartbeats never bind it; a machine off for months
catches up in a few hook fires (fail-closed latency only); one insane
sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
raw stamp lets an honest clock at revoke time pull a poisoned horizon
back to the true revoke moment. A rolled-back clock at close time
only closes earlier - fail-closed.
Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
(the harness re-implemented the insert; deleting the production line
survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
skip was dead code - the store constructor creates the directory
before the exists() check ran. The probe now checks the default path
without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
upgrade (fail-closed; deliberate).
New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.
273 tests pass; ruff and windows-footguns clean; staging E2E 202.
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).
Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).
Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
clamps window OPENS so a rolled-back clock cannot slide a window
under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
only for observed time, so a hand-edited config with no process
running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
undelivered backlog from the earlier consented window (round-5 D4).
The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.
Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).
Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).
Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).
270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.
BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.
The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.
BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.
Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
with `if True` survived the entire suite, because every non-http case
targeted a REMOTE host where the loopback branch rejects anyway. Only a
non-http scheme on loopback distinguishes the two. Shipped behaviour was
already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
Measured against 11 real packages: resource is a stable low-entropy
tuple and periods are contiguous across a rotation, so for a RARE
configuration those can bridge windows. The honest claim is that
rotation raises the cost, not that it makes correlation impossible.
Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.
An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.
259 tests pass. Staging E2E re-run: both packages 202.
Third independent review. Both blockers reproduced against a real store
before and after the fix.
BLOCKER 1 — head-of-line starvation. The claim query is LIMIT 1, and a
package already handled this pass was rejected AFTER the fetch, so
_claim_next returned None and send_pending read that as 'queue empty'.
Any row that sorts first and becomes eligible again mid-pass therefore
terminated the pass. This is reachable normally: a 429 with a short
Retry-After, or a pass outliving the 15-minute failure backoff (a legal
pass runs ~1900s). Measured: 10 of 19 healthy packages silently dropped.
The seen-set is now excluded IN SQL, so None genuinely means no eligible work.
Same scenario now delivers 19 of 19.
BLOCKER 2 — revoking consent leaked once it was re-granted. opt_in_period
was write-once, so packages collected while the user had send: false
still had period_start >= the ORIGINAL opt-in day; re-enabling released
the whole refused window. Reproduced: 5 packages from a 5-day opted-out
window transmitted on re-enable. Turning sending off now closes the
consent window, and the next enabled pass opens a new one from that day.
Recorded both in the setup wizard and in the sender itself, because
config.yaml can be hand-edited where the wizard never sees it.
Also: a send_attempts ceiling (a poisoned head row burned ~160 requests
over 30 days, unbounded), _defer clamps to >= 1s so it cannot write a
past deadline, and the dead skipped_not_due field is removed.
Test-quality fixes, since vacuous tests have been the recurring problem:
- the lease test asserted only 'in the future', passing for a 1s lease;
it now requires the lease to outlast one package's worst legal case
- test_shutdown_joins_the_send_thread grepped getsource for a method
name — a change-detector AGENTS.md rejects — and is now behavioural
- gzip determinism was unguarded: both retries in one pass compress in
the same second, so removing mtime=0 was caught by nothing. Now
compares output across a real second boundary.
All five new regressions are mutation-verified: reintroducing each bug
fails its test. The first attempt-ceiling test SURVIVED its mutation
(the seeded row was excluded by another predicate) and was rewritten to
drive the real loop.
251 tests pass. Staging E2E re-run: both packages 202.
Deliver changes_requested review outcomes through kanban subscriptions and
wake the origin for notify+wake / wake modes. Review-specific only: no task
mutation. Reasons are redacted, path-scrubbed and truncated before delivery.
Salvaged from #88694; conflicts with the #87733 wake-kinds expansion resolved
keep-both.
Second independent review found the lease fix incomplete. Reproduced
each finding before fixing.
BLOCKER — the batch lease expired mid-pass. _claim took up to 20 rows
under ONE shared lease, but a single package can legally consume ~96s
(three 30s timeouts plus 1s+5s backoff), so a full batch runs ~1900s
against a 180s lease. Later rows' leases expired while this pass still
held them, and another process re-sent them. Reproduced: 192s elapsed,
pkg-2 POSTed twice.
Packages are now claimed ONE AT A TIME, immediately before being sent,
so a lease only has to cover the package actually in flight. Verified:
same scenario now sends each package exactly once.
HIGH — revoking consent did not stop a running pass. The runtime read
send consent once before starting the thread, so a pass could keep
transmitting for minutes after a user set send: false, contradicting
the documented promise that it 'stops transmission immediately'.
Consent is now re-read before every package and fails CLOSED if it
cannot be established.
MEDIUM — all non-429 4xx were treated as permanent, discarding data.
403 is the ingest service's own origin guard: a Transform Rule or edge
misconfiguration would have permanently dropped every package sent
during the incident. Only 400 (malformed envelope) and 413 (over the
1 MiB cap) are terminal now; everything else retries.
MEDIUM — valid JSON that is not an object blocked the whole queue.
json.loads('["a"]') succeeds, then .get() raised AttributeError inside
the claim transaction, rolling it back and starving every healthy
package behind it. Payload shape and install_id are now validated, and
an unusable row is rejected individually.
LOW — the clock-rollback comment and test name claimed the opposite of
the code. The behaviour is right (a future issued_at means the recorded
age is untrustworthy, so reissue); the wording is now honest about it.
LOW — removed the stale HERMES_TELEMETRY_ENDPOINT reference left in
config_defaults after the override was deleted.
247 tests pass (was 234). Staging E2E re-run: both packages 202.
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.
The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.
The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.
Also from review:
- shutdown() never joined the send thread; the join was only wired into
deactivate(). A short-lived CLI therefore killed an in-flight send at
exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
secrets, and a behavioural override here was a consent hazard: an
inherited variable could silently redirect telemetry a user agreed to
send to Nous. The staging E2E writes the endpoint into its throwaway
profile instead, which also exercises the real config path.
- Added the shared-metrics toggle that AGENTS.md requires
as the third opt-in surface, delegating to the setup prompt so the
consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
so a wrong path or oversized body retried every 15 minutes for 30
days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
send pass, which silently dropped the opt-in day whenever the next
export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
package differ on the wire, so the 'byte-identical retry' E2E was
comparing parsed bodies and could not have caught it. It now compares
raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
said no remote-delivery path exists.
233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
The shared-metrics doc states that a future remote exporter 'must not
reuse the persistent local identifier by default' and 'requires a
separate product and privacy decision covering consent, identity scope,
rotation or keyed pseudonymization, reset behavior, retention, and
deletion'.
That exporter is now being built. Appendix A answers each of those six
items before any code lands, so the reasoning is reviewable on its own
and survives the implementation:
- consent is a separate opt-in from collection, gated on the PERIOD a
package covers rather than when it was created (a period is split
across packages made on different days, so a created_at gate would
send a period's tail while dropping its head and silently
undercount the first day)
- the transmitted identifier is HMAC-SHA256(local-only salt,
install_id), never install_id itself
- the salt rotates every 30 days
- reset gives a new remote identity but cannot unsend
- local retention is unchanged; send state does not extend it
- there is no self-service remote deletion, and the user -> derived-id
lookup that would enable one is deliberately not built
A.7 additionally records what the outbox directory IS (the user's local
history, not a send queue) because misreading it would have led to
deleting user data on acknowledgement.
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
test_contract_doc_conformance enforces that every CapabilityDescriptor
field appears in docs/relay-connector-contract.md §2 so connector authors
mirror the full surface — caught in CI (slice 2/12); the descriptor gained
supports_inchannel_continuable and supports_block_formatting without the
doc rows.
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.
- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
and returns (block_message, modified_args); modify directives
shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
{"action": "modify", "args": {...}} and Claude Code-compatible
{"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).
Salvaged from PR #28953. Best fix for #18988.
Reference prototype for the upcoming 'replace window.confirm/prompt/alert in
the kanban plugin' work. Self-contained HTML — open in any browser, no build
step. Served at docs/design/kanban-dialogs/index.html.
Four variants side-by-side, each rendered in the same kanban context:
- A (Conservative): direct host ConfirmDialog mapping, minimal chrome
- B (Strong-fit, Pro's pick): textarea + SVG icon + inline validation
- B-refined (synthesis, recommended): auto-focus, dual-validation,
cancellable spinner during PATCH, per-task summaries in bulk-many
- C (Divergent): undo toast for non-destructive moves
Decision matrix at the bottom of the page. Copy is verbatim from
web/src/i18n/en.ts (confirmDone, confirmArchive, confirmBlocked,
trash.confirm, completionSummary, etc.).
Design pass: Gemini 3.1 Pro (initial brief) + GPT-OSS 120B (cross-vendor
review). Not shipped to users — review reference only.
Code-level analysis of Pi (earendil-works/pi @ eb79351) and OpenCode
(anomalyco/opencode @ c69abee) plugin architectures across the six
dimensions #64180 specifies, with a 13-row adopt/adapt/avoid table
mapped to #64164/#64161/#64162/#64165/#64229/#64230. Key findings:
neither system has hook timeouts (both shipped hang-class bugs),
OpenCode's permission.ask is typed-but-dead (hook wire-up drift),
Pi treats prompt-cache stability as API contract, and both systems
lack ADRs. Fixes#64180.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013b1XyXitAxV7phGmKWigJX
* fix(gateway): pass live adapters to cron fire webhook's fire_due
The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).
Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.
Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).
* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)
The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.
Restore the invariant that the GATEWAY owns cron execution:
- Dashboard route: after verifying the NAS JWT and resolving the job's
profile, FORWARD the fire to the gateway api_server's own
/api/cron/fire on loopback, NAS bearer preserved (the gateway
re-verifies the JWT — defense in depth, no new trust link), and pass
the gateway's response through. Gateway unreachable → 503 so NAS
retries per the Chronos contract (non-2xx = retryable; the store CAS
de-dupes the eventual double fire). Deliberately NO local-execution
fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
per target profile (config.yaml extra.port → API_SERVER_PORT from
process env or the profile's .env → 8642), with /p/<profile>/ prefix
routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
first boot when absent (never overwrites an operator value), so the
loopback api_server passes its startup guard on hosted images. The
fire route itself is NAS-JWT-authed; the key gates the rest of the
api_server surface. The listener binds 127.0.0.1 by default and the
Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
topology and the 503-retry semantics.
Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.
* fix(cron): read the profile api_server port via the canonical config loader
CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).
Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.
* fix(gateway): only messaging platforms count for the scale-to-zero arm gate
The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).
The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
* fix(relay): stop sibling gateways answering another instance's button press
A Discord button press arrives on the passthrough plane, and the connector
fans a passthrough forward out to EVERY live gateway session of the tenant
(relayServer.routeBusMessage delivers `passthrough` via sessionsByTenant),
unlike a message, which it narrows to the admitted instance set. The prompt
went out from exactly one instance and _pending_prompts is process-local, so
every sibling gateway saw an answer for a prompt it never minted, could not
tell that from its own prompt expiring, and fell through to chat dispatch --
where the option-shaped text ("/c1") is not a real command and run.py replied
"Unknown command `/c1`". One copy per sibling, under the single real ack.
Prompt ids are now minted as `<per-process nonce>.<8 hex>`, so an answer can
be attributed to the process that minted it. A prompt answer is always
consumed, never re-dispatched as chat: a sibling's prompt and a repeat answer
are both dropped silently, and an expired prompt of our own gets a short
"no longer waiting" notice from the owning gateway only.
Ids stay inside the connector codec's contract ([A-Za-z0-9_.-], <=32 chars,
64-byte callback budget -- verified against promptCodec.ts: 52 bytes worst
case with a full-length option id). An id with no nonce segment (a prompt in
flight across an in-place upgrade) is still treated as ours.
Tests: 4 added, each verified to fail without the fix. Full relay suite green
(160 tests).
* style(tests): ruff-format the added relay prompt tests