Commit Graph

231 Commits

Author SHA1 Message Date
pierrenode
3245264668 fix(relay): carry routed profile through the passthrough-plane forward
The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.

Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.

Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
2026-09-02 06:47:45 -07:00
Teknium
b6f0106602 docs(design): note loader-boundary dotenv guard, scoped ${VAR} expansion and scoped .env publish in multiplexing design doc
Accuracy pass on the #89950 doc for the fixes landing alongside it (#77562, #84079, #88441).
2026-09-02 06:19:03 -07:00
Eva
011a60b7cc docs(design): add the multiplexing-gateway design doc referenced by secret_scope
agent/secret_scope.py has pointed at docs/design/multiplexing-gateway.md
("Workstream A") since the fail-closed secret scope landed, but the file was
never added. This writes the missing doc from the code as it stands today:
the mode flag, scope composition (_profile_runtime_scope seams), the
context-local secret scope and HERMES_HOME override, routing/serving/
persistence/session-lane isolation, the control-plane RPCs, failure modes,
and an honest table of what is still process-global (per-profile MCP
registries are tracked in #67605).

Doc-only change; no code touched. Style follows docs/profile-routing.md.
2026-09-02 06:19:03 -07:00
kshitijk4poor
e9160625dc fix(state): also disable SQLite's internal close-time checkpoint on quarantine (py3.12+)
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.

Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
2026-09-02 16:57:21 +05:30
leomcamilo
bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
Ben Barclay
180291162f feat(telemetry): opt-in shared-metrics exporter (#95278)
feat(telemetry): opt-in shared-metrics exporter
2026-09-02 08:35:36 +10:00
Teknium
33797073bb fix: harden startup route salvage — aggregator-slug guard, alias credential ownership, oneshot dedup
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
  routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
  URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
  provider label never carries the vendor token to the alias host (#28660);
  route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
  already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
2026-09-01 12:07:35 -07:00
joaomarcos
4c870951e2 fix(cli): resolve startup model routes before provider defaults
Resolve configured aliases and provider/model inputs before HermesCLI attaches the configured default provider. Keep aggregator namespaces intact, cover oneshot startup, and document the supported CLI forms.
2026-09-01 12:07:35 -07:00
the3asic
18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
joe102084
b028fe632e feat: add cron doctor health check 2026-08-31 10:00:29 -07:00
Ben Barclay
a69a9c351d feat(telemetry): transmit the stable install_id as-is
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.

Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
  derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
  role (unreadable/non-object/id-less payloads still reject rather than
  block the queue) and now records the raw install_id in
  sent_install_id; _body rewrites the payload's install_id from that
  frozen column, keeping byte-identical resends anchored to one
  recorded value.

Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.

Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.

Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.

258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
2026-08-28 15:22:16 +10:00
Ben Barclay
151632c333 Merge branch 'main' into feat/telemetry-exporter
387 commits from main; no conflicts (verified with merge-tree before
merging). Overlap limited to hermes_cli/config_defaults.py and
hermes_cli/setup.py, both auto-merged; all shared-metrics surfaces
untouched by main.
2026-08-27 14:42:41 +10:00
Ben Barclay
67d152bc7e fix(telemetry): bound forward-clock damage to the consent horizon
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.

The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.

Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
  call. Honest heartbeats never bind it; a machine off for months
  catches up in a few hook fires (fail-closed latency only); one insane
  sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
  Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
  raw stamp lets an honest clock at revoke time pull a poisoned horizon
  back to the true revoke moment. A rolled-back clock at close time
  only closes earlier - fail-closed.

Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
  (the harness re-implemented the insert; deleting the production line
  survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
  skip was dead code - the store constructor creates the directory
  before the exists() check ran. The probe now checks the default path
  without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
  upgrade (fail-closed; deliberate).

New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.

273 tests pass; ruff and windows-footguns clean; staging E2E 202.
2026-08-27 13:30:25 +10:00
Ben Barclay
5e380d95ba refactor(telemetry): replace consent day-stamp with explicit intervals
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).

Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).

Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
  confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
  clamps window OPENS so a rolled-back clock cannot slide a window
  under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
  only for observed time, so a hand-edited config with no process
  running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
  intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
  undelivered backlog from the earlier consented window (round-5 D4).

The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.

Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).

Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).

Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).

270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
2026-08-27 10:42:31 +10:00
Ben Barclay
613849c190 fix(telemetry): close the consent window on the config transition
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.

BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.

The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.

BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.

Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
  with `if True` survived the entire suite, because every non-http case
  targeted a REMOTE host where the loopback branch rejects anyway. Only a
  non-http scheme on loopback distinguishes the two. Shipped behaviour was
  already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
  Measured against 11 real packages: resource is a stable low-entropy
  tuple and periods are contiguous across a rotation, so for a RARE
  configuration those can bridge windows. The honest claim is that
  rotation raises the cost, not that it makes correlation impossible.

Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.

An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.

259 tests pass. Staging E2E re-run: both packages 202.
2026-08-27 09:47:47 +10:00
Ben Barclay
8ddff33e4c fix(telemetry): head-of-line starvation and consent-revocation leak
Third independent review. Both blockers reproduced against a real store
before and after the fix.

BLOCKER 1 — head-of-line starvation. The claim query is LIMIT 1, and a
package already handled this pass was rejected AFTER the fetch, so
_claim_next returned None and send_pending read that as 'queue empty'.
Any row that sorts first and becomes eligible again mid-pass therefore
terminated the pass. This is reachable normally: a 429 with a short
Retry-After, or a pass outliving the 15-minute failure backoff (a legal
pass runs ~1900s). Measured: 10 of 19 healthy packages silently dropped.
The seen-set is now excluded IN SQL, so None genuinely means no eligible work.
Same scenario now delivers 19 of 19.

BLOCKER 2 — revoking consent leaked once it was re-granted. opt_in_period
was write-once, so packages collected while the user had send: false
still had period_start >= the ORIGINAL opt-in day; re-enabling released
the whole refused window. Reproduced: 5 packages from a 5-day opted-out
window transmitted on re-enable. Turning sending off now closes the
consent window, and the next enabled pass opens a new one from that day.
Recorded both in the setup wizard and in the sender itself, because
config.yaml can be hand-edited where the wizard never sees it.

Also: a send_attempts ceiling (a poisoned head row burned ~160 requests
over 30 days, unbounded), _defer clamps to >= 1s so it cannot write a
past deadline, and the dead skipped_not_due field is removed.

Test-quality fixes, since vacuous tests have been the recurring problem:
- the lease test asserted only 'in the future', passing for a 1s lease;
  it now requires the lease to outlast one package's worst legal case
- test_shutdown_joins_the_send_thread grepped getsource for a method
  name — a change-detector AGENTS.md rejects — and is now behavioural
- gzip determinism was unguarded: both retries in one pass compress in
  the same second, so removing mtime=0 was caught by nothing. Now
  compares output across a real second boundary.

All five new regressions are mutation-verified: reintroducing each bug
fails its test. The first attempt-ceiling test SURVIVED its mutation
(the seeded row was excluded by another predicate) and was rewritten to
drive the real loop.

251 tests pass. Staging E2E re-run: both packages 202.
2026-08-27 09:09:47 +10:00
DmytroVolodymyrson
a699234f81 fix(kanban): wake controllers on review changes
Deliver changes_requested review outcomes through kanban subscriptions and
wake the origin for notify+wake / wake modes. Review-specific only: no task
mutation. Reasons are redacted, path-scrubbed and truncated before delivery.

Salvaged from #88694; conflicts with the #87733 wake-kinds expansion resolved
keep-both.
2026-08-26 15:31:36 -07:00
Ben Barclay
d0a7144ba1 fix(telemetry): per-row claiming, mid-pass consent re-check, narrower 4xx
Second independent review found the lease fix incomplete. Reproduced
each finding before fixing.

BLOCKER — the batch lease expired mid-pass. _claim took up to 20 rows
under ONE shared lease, but a single package can legally consume ~96s
(three 30s timeouts plus 1s+5s backoff), so a full batch runs ~1900s
against a 180s lease. Later rows' leases expired while this pass still
held them, and another process re-sent them. Reproduced: 192s elapsed,
pkg-2 POSTed twice.

Packages are now claimed ONE AT A TIME, immediately before being sent,
so a lease only has to cover the package actually in flight. Verified:
same scenario now sends each package exactly once.

HIGH — revoking consent did not stop a running pass. The runtime read
send consent once before starting the thread, so a pass could keep
transmitting for minutes after a user set send: false, contradicting
the documented promise that it 'stops transmission immediately'.
Consent is now re-read before every package and fails CLOSED if it
cannot be established.

MEDIUM — all non-429 4xx were treated as permanent, discarding data.
403 is the ingest service's own origin guard: a Transform Rule or edge
misconfiguration would have permanently dropped every package sent
during the incident. Only 400 (malformed envelope) and 413 (over the
1 MiB cap) are terminal now; everything else retries.

MEDIUM — valid JSON that is not an object blocked the whole queue.
json.loads('["a"]') succeeds, then .get() raised AttributeError inside
the claim transaction, rolling it back and starving every healthy
package behind it. Payload shape and install_id are now validated, and
an unusable row is rejected individually.

LOW — the clock-rollback comment and test name claimed the opposite of
the code. The behaviour is right (a future issued_at means the recorded
age is untrustworthy, so reissue); the wording is now honest about it.

LOW — removed the stale HERMES_TELEMETRY_ENDPOINT reference left in
config_defaults after the override was deleted.

247 tests pass (was 234). Staging E2E re-run: both packages 202.
2026-08-26 17:10:43 +10:00
Ben Barclay
49757d5e39 fix(telemetry): address review findings on the shared-metrics sender
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.

The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.

The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.

Also from review:

- shutdown() never joined the send thread; the join was only wired into
  deactivate(). A short-lived CLI therefore killed an in-flight send at
  exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
  secrets, and a behavioural override here was a consent hazard: an
  inherited variable could silently redirect telemetry a user agreed to
  send to Nous. The staging E2E writes the endpoint into its throwaway
  profile instead, which also exercises the real config path.
- Added the  shared-metrics toggle that AGENTS.md requires
  as the third opt-in surface, delegating to the setup prompt so the
  consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
  so a wrong path or oversized body retried every 15 minutes for 30
  days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
  send pass, which silently dropped the opt-in day whenever the next
  export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
  package differ on the wire, so the 'byte-identical retry' E2E was
  comparing parsed bodies and could not have caught it. It now compares
  raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
  said no remote-delivery path exists.

233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
2026-08-26 16:31:52 +10:00
Ben Barclay
e0170c2536 docs(observability): record remote-exporter privacy decisions
The shared-metrics doc states that a future remote exporter 'must not
reuse the persistent local identifier by default' and 'requires a
separate product and privacy decision covering consent, identity scope,
rotation or keyed pseudonymization, reset behavior, retention, and
deletion'.

That exporter is now being built. Appendix A answers each of those six
items before any code lands, so the reasoning is reviewable on its own
and survives the implementation:

- consent is a separate opt-in from collection, gated on the PERIOD a
  package covers rather than when it was created (a period is split
  across packages made on different days, so a created_at gate would
  send a period's tail while dropping its head and silently
  undercount the first day)
- the transmitted identifier is HMAC-SHA256(local-only salt,
  install_id), never install_id itself
- the salt rotates every 30 days
- reset gives a new remote identity but cannot unsend
- local retention is unchanged; send state does not extend it
- there is no self-service remote deletion, and the user -> derived-id
  lookup that would enable one is deliberately not built

A.7 additionally records what the outbox directory IS (the user's local
history, not a send queue) because misreading it would have led to
deleting user data on acknowledgement.
2026-08-26 15:32:26 +10:00
Ben Barclay
a1a1bee5c2 Merge branch 'main' into feat/relay-slack-parity
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
2026-08-20 20:56:03 +10:00
Bryan Bednarski
6f7596e9a2 docs(relay): link supported observability exporters
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:04 -07:00
Alex Fournier
31402f630b fix(relay): complete native plugin cutover
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski
6ec2c0ba8b fix(relay): define process-wide profile policy
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski
3fad83df31 fix(relay): guard native plugin ownership cutover
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski
ad9fb060c5 fix(relay): clarify opt-in plugin layering
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski
8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski
0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski
e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Victor Kyriazakos
afcf4f5214 docs(relay): document the two new descriptor capability bits in the contract §2 table
test_contract_doc_conformance enforces that every CapabilityDescriptor
field appears in docs/relay-connector-contract.md §2 so connector authors
mirror the full surface — caught in CI (slice 2/12); the descriptor gained
supports_inchannel_continuable and supports_block_formatting without the
doc rows.
2026-08-19 14:25:16 +00:00
Nikola Hristov
d083b85591 feat(hooks): pre_tool_call content transformation via modify directive
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.
2026-08-15 23:03:52 -07:00
David Metcalfe
da68ecf4b3 design(kanban): add side-by-side dialog prototype (4 variants)
Reference prototype for the upcoming 'replace window.confirm/prompt/alert in
the kanban plugin' work. Self-contained HTML — open in any browser, no build
step. Served at docs/design/kanban-dialogs/index.html.

Four variants side-by-side, each rendered in the same kanban context:
- A (Conservative): direct host ConfirmDialog mapping, minimal chrome
- B (Strong-fit, Pro's pick): textarea + SVG icon + inline validation
- B-refined (synthesis, recommended): auto-focus, dual-validation,
  cancellable spinner during PATCH, per-task summaries in bulk-many
- C (Divergent): undo toast for non-destructive moves

Decision matrix at the bottom of the page. Copy is verbatim from
web/src/i18n/en.ts (confirmDone, confirmArchive, confirmBlocked,
trash.confirm, completionSummary, etc.).

Design pass: Gemini 3.1 Pro (initial brief) + GPT-OSS 120B (cross-vendor
review). Not shipped to users — review reference only.
2026-08-15 00:33:32 -07:00
terry197913
4e1b2e436c fix: scope plugin manager by resolved hermes home (keyed cache)
fix: remove .codegraph artifacts from commit
2026-08-12 19:13:32 -07:00
Teknium
6bf93c0e38 feat(plugins): add namespaced config and durable state bridge 2026-08-12 16:26:43 -07:00
Topher Ross
223e74345e docs(rfc): align config bridge with #64227 — namespace jail, real APIs, acceptance criteria
- Replace cfg_set references with cfg_get (cfg_set deferred to #64227)
- Add Namespace jail section: key-prefix enforcement, cross-plugin rejection,
  path-traversal rejection, read allow-list
- Rewrite cron sketch: replace CronManager/command with create_job() using
  prompt/script params per cron/jobs.py:1039-1234
- Add #64227 to Related
2026-08-12 16:26:43 -07:00
Topher Ross
516ca4170a docs: plugin config & state bridge design proposal
RFC covering four PluginContext additions driven by kanban-advanced needs:
- ctx.get_config() / ctx.set_config() — typed config access
- ctx.register_config_schema() — schema validation surfaced to hermes doctor
- ctx.cron facade — cron CRUD without subprocess fragility
- provides_config_defaults in plugin.yaml — safe config defaults on install
2026-08-12 16:26:43 -07:00
hans
299996f7e0 docs(rfcs): plugin-architecture research spike — Pi and OpenCode lessons
Code-level analysis of Pi (earendil-works/pi @ eb79351) and OpenCode
(anomalyco/opencode @ c69abee) plugin architectures across the six
dimensions #64180 specifies, with a 13-row adopt/adapt/avoid table
mapped to #64164/#64161/#64162/#64165/#64229/#64230. Key findings:
neither system has hook timeouts (both shipped hang-class bugs),
OpenCode's permission.ask is typed-but-dead (hook wire-up drift),
Pi treats prompt-cache stability as API contract, and both systems
lack ADRs. Fixes #64180.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013b1XyXitAxV7phGmKWigJX
2026-08-12 16:25:01 -07:00
Ben Barclay
bb597e1c02 fix(cron): managed-cron fires execute in the gateway process (live adapters + dashboard forwarder) (#84339)
* fix(gateway): pass live adapters to cron fire webhook's fire_due

The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).

Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.

Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).

* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)

The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.

Restore the invariant that the GATEWAY owns cron execution:

- Dashboard route: after verifying the NAS JWT and resolving the job's
  profile, FORWARD the fire to the gateway api_server's own
  /api/cron/fire on loopback, NAS bearer preserved (the gateway
  re-verifies the JWT — defense in depth, no new trust link), and pass
  the gateway's response through. Gateway unreachable → 503 so NAS
  retries per the Chronos contract (non-2xx = retryable; the store CAS
  de-dupes the eventual double fire). Deliberately NO local-execution
  fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
  per target profile (config.yaml extra.port → API_SERVER_PORT from
  process env or the profile's .env → 8642), with /p/<profile>/ prefix
  routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
  first boot when absent (never overwrites an operator value), so the
  loopback api_server passes its startup guard on hosted images. The
  fire route itself is NAS-JWT-authed; the key gates the rest of the
  api_server surface. The listener binds 127.0.0.1 by default and the
  Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
  compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
  topology and the 503-retry semantics.

Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.

* fix(cron): read the profile api_server port via the canonical config loader

CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).

Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.

* fix(gateway): only messaging platforms count for the scale-to-zero arm gate

The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).

The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
2026-08-12 17:04:44 +10:00
Ben Barclay
a189286aea fix(relay): stop sibling gateways answering another instance's button press (#83677)
* fix(relay): stop sibling gateways answering another instance's button press

A Discord button press arrives on the passthrough plane, and the connector
fans a passthrough forward out to EVERY live gateway session of the tenant
(relayServer.routeBusMessage delivers `passthrough` via sessionsByTenant),
unlike a message, which it narrows to the admitted instance set. The prompt
went out from exactly one instance and _pending_prompts is process-local, so
every sibling gateway saw an answer for a prompt it never minted, could not
tell that from its own prompt expiring, and fell through to chat dispatch --
where the option-shaped text ("/c1") is not a real command and run.py replied
"Unknown command `/c1`". One copy per sibling, under the single real ack.

Prompt ids are now minted as `<per-process nonce>.<8 hex>`, so an answer can
be attributed to the process that minted it. A prompt answer is always
consumed, never re-dispatched as chat: a sibling's prompt and a repeat answer
are both dropped silently, and an expired prompt of our own gets a short
"no longer waiting" notice from the owning gateway only.

Ids stay inside the connector codec's contract ([A-Za-z0-9_.-], <=32 chars,
64-byte callback budget -- verified against promptCodec.ts: 52 bytes worst
case with a full-length option id). An id with no nonce segment (a prompt in
flight across an in-place upgrade) is still treated as ours.

Tests: 4 added, each verified to fail without the fix. Full relay suite green
(160 tests).

* style(tests): ruff-format the added relay prompt tests
2026-08-12 10:04:26 +10:00
ethernet
0e2b5f3835 chore: remove old plan files 2026-08-09 15:17:01 -04:00
GodsBoy
8cb066404e fix(plugins): address portable MCP review feedback 2026-08-07 09:44:21 -07:00
GodsBoy
e288d93fc1 fix(review): harden portable plugin boundaries 2026-08-07 09:44:21 -07:00
GodsBoy
c5117655b6 feat(plugins): validate portable agent packages 2026-08-07 09:44:21 -07:00
HexLab98
2d0c2682c7 docs: document the gateway agent cache memory budget
Record why the cache needs a third bound and what the pressure pass will and
will not shed, so an operator tuning agent.agent_cache knows which knob to
reach for. Adds the config keys to the session-lifecycle appendix and a user
guide section covering the "auto" cgroup-derived budget.
2026-08-07 21:02:33 +05:30
Alex Fournier
535a59c5c0 docs(observability): clarify active profile identity
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-04 15:13:34 -07:00
Alex Fournier
942d731553 Merge updated client resource metrics into active-install metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-02 20:13:05 -07:00
Alex Fournier
d0322fad69 Merge updated skill metrics into client resource metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-02 20:12:50 -07:00
Alex Fournier
884c2daa1c Merge updated tool metrics into skill metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>

# Conflicts:
#	tests/tools/test_skills_hub.py
2026-08-02 20:12:37 -07:00
Alex Fournier
14c8bd646c Merge updated model metrics into tool metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-02 20:11:20 -07:00
Alex Fournier
a97abcd55a Merge upstream main into model metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>

# Conflicts:
#	tests/agent/test_auxiliary_relay.py
2026-08-02 20:10:47 -07:00