Commit Graph

616 Commits

Author SHA1 Message Date
teknium1
8d8836ddb1 feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):

- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
  matcher reports the strategy that landed (or no_match / ambiguous) into a
  context-local probe opened only around the patch/write_file handlers, so we
  learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
  the guardrail already warns/blocks/halts and at turn end for the iteration
  budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
  next_outcome. One row per failed tool call, resolved against the model's
  next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
  a table lookup on the first program word, never the text. Hermes' own
  deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
  backend result, so a command's own `exit 124` reads nonzero, not timeout.
  Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
  truncation only from the structured finish reason; empty/reasoning_only
  from the normalized message; issue=none per response as the denominator
  (model_route counts attempts and files empty replies as failures).

Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
2026-09-28 12:43:03 -07:00
teknium1
88eb76f02e fix(metrics): startup latency counts once per client launch, never after an in-place exec
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat
  call added a row. The TUI and Desktop now send an opaque per-launch id
  (TUI: one per process; Desktop: minted when the once-per-launch main-process
  claim succeeds) and the backend claims (surface, launch_id) once per
  process: a reconnect re-sending the same launch counts once, a new Desktop
  launch against a long-lived backend still counts. Legacy clients without an
  id count once per backend process and surface. The id is only a local latch
  key (bounded set, length-capped), never recorded. Only a usable measurement
  spends the claim. Contract + apps/shared regenerated.
- m10: relaunch() execs in place, keeping the PID, so `sessions browse` ->
  resume counted the picker time as CLI startup. relaunch() stamps its PID in
  HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it
  matches, while children (other PIDs) inheriting the env still count.
- m8 (psutil half): record_process_ready checks the opt-in gate before reading
  the process start time (psutil) or starting a thread. The claim is taken
  first so a later opt-in never records a stale "startup" mid-session.

Test changes to existing assertions: the TUI/Desktop param assertions now
expect launch_id (a new wire param); the RPC surface test sends distinct
launch ids so its three calls stay three launches under the new latch.
2026-09-28 12:43:03 -07:00
teknium1
f7b5ea6f7a fix(telemetry): loop metrics count only the user's real backend work and curator passes
Review findings M7, M8, m6, m7 and the browser half of m8 on the v4 loop metrics.

- M7: TUI/Desktop @-path completion on a non-local backend lists the directory through
  terminal_tool, so every keystroke added a hermes.execution_backend.count row as if the user
  ran a command. Hermes-owned calls now run inside shared_metrics_loop.unmetered_backend_calls()
  (a contextvar checked in record_execution_backend) instead of a flag threaded through
  terminal_tool. It was the only non-_host_local internal terminal_tool caller (bot DM runners
  already use _host_local; the prompt-builder probe calls env.execute directly).
- M8: a scheduled curator tick that finds the run claim held recorded scheduled/skipped every
  tick (5 ticks -> 5 rows for one pass). The holder's own pass is the one counted; the held
  branch now records nothing. Real dry runs still report skipped.
- m6: a foreground timeout returns exit 124 with partial output and no error field, so it was
  recorded as success. exit 124 now records outcome=failed, error_class=timeout.
- m7: vercel_sandbox (and managed_modal) are built-in backends in
  tools/terminal_tool_config._BUILTIN_BACKENDS but collapsed to other. Added to
  TERMINAL_BACKENDS and both schema enums; a test pins every built-in backend to its own bucket.
- m8 (browser): record_browser_call resolved the browser backend on every call even with
  collection off. The resolver is now passed lazily and only runs inside the gated builder.

The existing test asserting a held claim records scheduled/skipped encoded M8 itself and is
replaced by one asserting held ticks record nothing.
2026-09-28 12:43:03 -07:00
teknium1
287564fedd fix(telemetry): update runs counted once, parked only while opted in, recovered files kept until saved
Three defects in how hermes.update.run/stage and hermes.process.exit rows
are recovered from disk:

- Opt-in (M1): the pre-pull park path only checked that the telemetry dir
  existed, so it wrote the WHOLE receipt (argv, step detail) while the user
  was opted out, and a later opt-in counted that run. It now reads consent
  through the already-loaded hermes_cli.config (this interpreter predates
  the checkout swap, so nothing may be imported), parks only when
  collection is on, parks only the fields update_receipt_fields reads (no
  argv, no step text), and purges pending_updates when collection is found
  off. begin_process purges them too on a start with collection off.
- Double count (M2): the completion child records the receipt and the
  parent's boundary finalize (run_completion returned receipt=None) parks
  the same update_id again, so one run became (failed, deps) + (success,
  none). record_update_receipt now claims the update_id with an O_EXCL
  latch under shared_metrics/recorded_updates (last 64 kept); the second
  finalize records nothing.
- Loss (m3): dead-process markers and parked receipts were unlinked before
  their row reached the store; "database is locked" under concurrent starts
  lost them (17/20 rows in the reviewer's race). The claim-by-rename stays,
  but the file is deleted only after record_process_marks_saved confirms
  the rows settled: rows carry a random commit ticket in event metadata
  (allowlisted in the contract), the subscriber tallies tickets whose rows
  persisted, and the reporter flushes the Relay subscribers and compares.
  Otherwise the file is renamed back for the next start. A parked receipt
  whose rows partly landed is not retried (that would double count).

Also: _epoch returns None for a timestamp float() cannot represent
(10**400 raised OverflowError).
2026-09-28 12:43:03 -07:00
teknium1
2c9b79a5a3 fix(telemetry): one context_peak row per conversation across compression
Rotating compaction continues a conversation under a new session id, and
each metrics session emitted its own context_peak at close, so one
conversation produced a row per segment (the fuller one plus a low-fill
one), skewing the per-model distribution. The agent turn already
publishes the lineage root (the Portal conversation id); segments sharing
it join one lineage, contribute their peak when they close, and the
merged peak (fullest fill, limit hit anywhere) is emitted once when the
last open segment closes, whatever the close order. Delegated children
share their parent's root but stay their own conversation.

Review finding m2.
2026-09-28 12:43:03 -07:00
teknium1
093c10a2ee fix(telemetry): local-server aliases and unknown providers never ship model ids
provider_names() accepts ollama/local/vllm/llamacpp/llama.cpp as published
names (they are aliases of the generic `custom` provider), so every
provider/model metric kept the raw config.yaml model id under them, and a
missing provider (`unknown`) kept it too. Map the aliases Hermes itself
routes to `custom` (derived from the providers/auth/models alias tables)
to `custom`, and report the model as `custom` whenever the provider is
unknown: nothing proves such an id is public. Shipped providers with
public ids are unchanged. Covers model_route, model_tokens (primary and
auxiliary), fallback, model_switch, setup.completed, install snapshot
main_provider, friction, tool_quality and context_peak, which all go
through this pair.

Gateway /model also labelled the configured model with switch_model's
`openrouter` default when config.yaml names no provider; it now reports
the provider actually configured/overridden (None -> `unknown`). And the
switch row + switch_away friction bind the routed profile's home on a
multiplexed runner (slash dispatch installs no profile scope), as the
/retry and /undo friction already do.

Review findings B2, M3.
2026-09-28 12:43:03 -07:00
teknium1
fd14f50b5d feat(metrics): hermes.update.run/stage derived from the final update receipt
The update pipeline already writes one machine-readable receipt per run, so the
run/stage counters are derived from it in the process that finalizes it instead
of instrumenting stages for metrics. The receipt gains minimal stage END marks
(plan, snapshot, apply[mode], deps, build, restart; verify is inferred from the
fleet matrix), an initiator fact (desktop when a live orchestrator claim
names another pid) and the pre-update commit date, which is all the derivation
needs for outcome, failed stage, duration and from-version-age buckets.

A pre-pull interpreter must never import pulled code: when it is the finalizer
(same pid, checkout sha moved) it parks the receipt (stdlib only) and the next
Hermes start records it. Collection off => nothing imported beyond a config read.
The completion-process test stub gains the new record_stage API the real
module now exposes (no assertion changed).
2026-09-28 12:43:03 -07:00
teknium1
2b3e1c3995 feat(metrics): startup latency and install version-lag/hardware fields
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.

hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.

The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
2026-09-28 12:43:03 -07:00
teknium1
2efa4f3caa feat(telemetry): per-model tool-call quality, friction and context-peak counters
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.

hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.

hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.

hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
2026-09-28 12:43:03 -07:00
teknium1
f4075a20d6 feat(telemetry): gateway platform health, delivery, first-reply latency and cron run metrics
The gateway and cron ticker were blind spots in shared metrics: we could not
tell which messaging platforms fail to connect or drop, how often replies fail
to reach users, how long users wait for a first reply, or whether scheduled
jobs run, fail, get skipped or get missed while Hermes was down.

New opt-in counters (all through the existing enabled() / record_process_mark
gate, recorded fire-and-forget on one background worker that runs in a copy of
the caller's context so the owning profile gets the row):

- hermes.platform.health: platform, event (connect_ok/connect_failed/
  reconnect/disconnect), error_class. One seam in the runner
  (_connect_adapter_with_timeout, used by cold start, multiplex secondaries and
  the reconnect watcher) plus the fatal-error handler; classified from
  exception types, HTTP statuses and Hermes's own fatal codes, never text.
- hermes.platform.delivery: platform, outcome, failure_class. One row per
  logical BasePlatformAdapter._send_with_retry call (retries and plain-text
  fallback included), from SendResult's typed fields or the exception type.
- hermes.gateway.reply_latency: platform, first_response_bucket. Clock starts
  when _handle_message accepts a non-internal turn; stops on the stream
  consumer's first delivered text or send_final_ledgered. Busy acks and
  command replies never stop it.
- hermes.cron.run: outcome (success/failed/missed/skipped), delivery_kind,
  duration_bucket. Recorded at the write-once terminal execution row
  (finish_execution) and where the due scan drops an occurrence (catch-up
  disabled, expired one-shot). Job names, prompts and targets never leave.

Platform names: platforms Hermes ships under plugins/platforms/ now report by
name, and a plugin-catalog platform reports its catalog entry name only when
the installer-owned .install-metadata.json record proves a catalog install of
the plugin dir that defines the registered adapter factory (a URL install
cannot claim a catalog name via its tree or manifest). Everything else stays
"plugin". GATEWAY_PLATFORMS becomes catalog-backed (_CatalogValues), and the
v3 schema's platform fields accept published catalog identifiers.
2026-09-28 12:43:03 -07:00
teknium1
2e56f40c16 feat(telemetry): learning-loop, delegation and execution-backend metric contract
Adds four opt-in shared-metrics counters so we can see whether the learning loop
actually runs and pays off, how wide delegate_task fan-outs go, and which
sandboxes carry real work:

- hermes.memory.op.count: op, provider (builtin / bundled plugin / plugin),
  origin (foreground / background_review), outcome (success / failed / rejected)
- hermes.curator.run.count: trigger, outcome, archived/created/merged/patched buckets
- hermes.delegation.run.count: subagent_count_bucket, depth, mode, outcome
- hermes.execution_backend.count: kind (terminal/browser/code), backend, outcome,
  error_class

Every dimension is a closed enum or COUNT_BUCKETS value; unknown providers and
backends collapse to plugin / other. Builders live in the new
shared_metrics_loop.py sibling (stdlib-only imports so tool hot paths can use it)
and go through the existing enabled() gate. Off-turn producers (curator thread,
async delegation units) record in the profile captured when the work started.
Schema and the "what is collected" doc gain matching entries.
2026-09-28 12:43:03 -07:00
teknium1
9edbf627f0 fix(desktop): shared-metrics question is a composer offer strip, never a launch modal
The first-run consent dialog blocked the composer on every undecided
profile, so a healthy install no longer opened straight to chat (caught by
the Desktop first-run E2E). The question now lives in the composer status
stack beside the free-tier strip: Send to Nous / Local only / No thanks /
Details, no focus steal, one owner across split composers, one offer at a
time. Details opens the full explainer on request; closing it decides
nothing. The backend's `decided` stays the only latch.
2026-09-28 12:43:03 -07:00
teknium1
3d62ae2233 feat(telemetry): decision-data docs, consent copy, smoke coverage and install-age for new installs
- Docs: a Decision-data metrics section mapping each metric to the product question it answers;
  skill-load and snapshot paragraphs describe the new public-name and install-age fields.
- setup consent copy names the new data classes (session length, token totals, command and
  catalog names, bucketed setup counts).
- The real-CLI smoke now asserts the session summary, token sums, TTFT bucket and milestones in
  both the SQLite store and the exported, schema-validated package.
- A profile with no sessions yet is a brand-new install (lt_1h), not unknown: the first task's
  milestone and the first snapshot fire before state.db has a session row.
- Two invariant tests: session summaries + once-per-install milestones; token sums per model and
  auxiliary task with private task names collapsed.
2026-09-28 12:43:03 -07:00
teknium1
b8732b768e feat(telemetry): shared metrics v3 with failure classes, tool usage, platform and daily snapshot
The fleet rollups could say that tasks and tools failed, but not why, on which
model, on which messaging platform, or which features installs actually use.

- model_route rows gain call_role, outcome and error_class (the classifier's
  FailoverReason for the last failed attempt; a success row keeps the error it
  recovered from).
- task_run rows gain platform (built-in gateway platform, plugin, or none);
  task_run.finished gains failure_class from a closed set.
- new hermes.tool.usage.count: tool_name (only names from the static
  toolsets.TOOLSETS snapshot; mcp/plugin otherwise), outcome, error_class.
- new hermes.install.snapshot, latched to once per rolling 24h like
  client.active: memory provider plus bucketed counts of MCP servers,
  plugins, skills, cron jobs and profiles; no names.
- package schema v3; v2-shaped pending counters still validate and drain.
- consent copy and developer docs list the new categories; the relay smoke
  asserts v3 shapes and disables title generation (its third model request
  broke the smoke's request-count check on main).
2026-09-28 12:43:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
01d85137bd fix(review): MCP owner rebuild keeps the launch profile's env-only credentials (#119092 review)
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.

Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
2026-09-28 05:24:26 -07:00
Siddharth Balyan
98278833e2 Compaction follows compression.threshold again: no default 256K token cap (#117915, #125235) (#126064)
* fix(compression): compaction follows the ratio again; no default token cap

compression.threshold_tokens defaulted to 256000 since #115986, so every
window above ~341K compacted at 256K regardless of compression.threshold
or model_thresholds: a 1M-window model compacted at 25% of its window,
and threshold: 0.8 still compacted at 256K. No single token count suits
windows from 64K to 1M+, so the cap goes back to opt-in (null) and the
ratio decides, as it did before #115986.

The template seeder and `hermes doctor --fix` copied the 256000 default
into config.yaml, where it reads as a user choice and would keep capping
those installs. Config v47 drops threshold_tokens only when it equals
256000; any other explicit cap and an explicit null are preserved.

The rest of #115986 stays: the model-switch warning quoting the real
trigger and the shared _derive_trigger are correct with any default.
Tests that encoded 256K now derive expectations from DEFAULT_CONFIG.

* docs(compression): threshold_tokens is an optional cap, default null

User guide, developer guide and delegation page described the 256K
default cap; they now describe the ratio trigger as the default and
threshold_tokens as an opt-in cost ceiling.
2026-09-28 07:29:40 +00:00
Brooklyn Nicholson
929de15c73 docs(desktop): correct WorkspacePageHeaderControl older-host note
Some checks failed
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
A named import of a missing SDK export fails to link through the runtime
shim, so the plugin never loads; it is not undefined. Point authors at a
namespace-import feature-detect or the raw Contribute form.

Refs #123597

Originally authored by Justin Haynes (@jhaynes).
2026-09-27 13:18:46 -05:00
Brooklyn Nicholson
aa25f9e85f fix(desktop): Kanban board switcher outside the full-page layout
A Kanban board opened in a split route tile rendered no board switcher:
the board contributed it to WORKSPACE_PAGE_HEADER_AREA unconditionally,
and only the workspace pane paints that area. The tile's contribution
also leaked into another page's header and shared its id with the full
page's, so closing the tile removed the page's switcher.

Add WorkspacePageHeaderControl (exported via the plugin SDK). The
workspace pane's render provides a private host context; inside it the
control projects into the page header, anywhere else it renders inline.
The board mounts BoardSwitcher once, through it, in its own header row.

Fixes #123597

Originally authored by Justin Haynes (@jhaynes).
2026-09-27 13:18:46 -05:00
Adolanium
d1167fdff7 fix(desktop): match custom:<key> providers in the settings and bot model pickers
model.info and saved profiles report a user-defined provider as custom:<key>, but the catalog row uses the bare key as its slug. Settings > Model and the Bot Mode picker compared the two with ===, so a saved custom provider never found its row. Settings showed a duplicate custom:<key> entry and a Set up provider button, and the bot editor fell back to the manual form.

Both now match rows with catalogProviderMatches, like the composer picker already does. Settings uses a small findCatalogProvider helper for every row lookup, including the aux and MoA slots and the endpoint passed on Set to main. catalogProviderMatches is now exported through the plugin SDK so the bot picker can use it.
2026-09-27 12:49:56 -05:00
Adolanium
e46d4c0ade fix(compression): let the next summary build on a fallback handoff
A deterministic fallback summary replaced the older handoff in the transcript but never updated _previous_summary. The next compaction kept the stale in-memory summary and dropped the fallback row from its window, so the fallback's user asks, files and last dropped turns never reached the summarizer. Store the fallback body in _previous_summary the same way a normal summary is stored.

(cherry picked from commit 35417d2e1ffbb775c3eaff17b26623896afa56c1)
2026-09-27 18:20:40 +05:30
MongLong0214
bb17b1f74c fix(compression): keep the pending round's images when it fits
Compaction spared a pending tool round's text results but not its
images. The image-retirement pass kept only the newest three image
results across the transcript, so four parallel vision calls in one
unread round lost the oldest image before the model saw any of them.
On a single-prompt run the final media pass lost three of the four:
compaction re-appends the task as a user row after the round, and the
pending round was looked up after that row was added, so nothing was
spared.

Both image passes now skip a pending round that fits the hard share,
and the pending round is found before the task row is re-appended.
Completed rounds and a pending round over the hard share keep the
existing policy.

(cherry picked from commit eb7661f4365f009d5f9ef85e26f2f4b6ac83632e)
2026-09-26 23:45:36 +05:30
MongLong0214
65f185c594 fix(compression): keep the pending tool round when two steers follow it
The pending-round check skipped only one trailing /steer row. Two can
land in the same iteration: one is appended when the tool batch ends,
and a steer sent after that is drained before the next request and
inserted right after the newest tool result. The transcript then ends
tool -> steer -> steer, so the check stopped on the second steer row,
found no pending round, and preflight compaction replaced the unread
tool output with the one-line pressure stub.

The check now skips every contiguous trailing steer row, identified the
same way as before, and still stops on any other user row. The existing
regression gains a case with two steer rows after the round.

(cherry picked from commit 455a1a3868b3843d15401781363ce796bae96b5b)
2026-09-26 23:45:36 +05:30
MongLong0214
18eb09f5db fix(compression): keep the unanswered tool round verbatim through mid-turn compaction
Preflight compaction can fire right after a tool round, before the model
has read its results. The protected-tail passes then treated that round
like any old output: pass 2 and the pressure pass (#61932) replaced its
results with one-line stubs, and when the newest row alone exceeded the
tail ceiling the cut landed at the end of the transcript and summarised
the round with its whole turn. The model then re-ran the call, side
effects included, or answered without the output.

_pending_tool_round finds the tool results the transcript ends with,
skipping one trailing /steer row (a steer is delivered after the newest
result, before the next API call). Both passes spare that round, and the
tail cut aligns from the row before the end so the group stays whole.
The one exception is a round that alone exceeds the tail's hard share
(20%) of the input budget, the window minus the output reservation: it
still gives way, as #61932 requires. _effective_input_window is
extracted from _compute_threshold_tokens so both use the same budget;
thresholds are unchanged.

(cherry picked from commit 9d74e22379cd7dc39636c522175772d0165db4ac)
2026-09-26 23:45:36 +05:30
Yuan Li
a5e8ad9f02 fix(agent): bound sustained summary-overload aborts so they cannot guarantee a session wipe
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.

After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.

Fixes #123167

(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
2026-09-26 22:42:38 +05:30
kshitijk4poor
11c50f05d0 refactor(compression): Astra native-compaction gate reuses is_astra_model
The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.

Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
2026-09-26 19:27:41 +05:30
ethernet
b5d583c4ac feat(release): start the gates and the signed candidates together
Every gate and every candidate now needs only admit, so the signed macOS
and Windows builds and the docker image stop waiting behind the full CI
run. acceptance stays the one join: it still needs ci, docker and all six
candidates, so what can be promoted is unchanged.

The four gates that skipped under --skip-tests only because ci skipped
(nix, termux-checks, windows-live, install-e2e) and pm-bundle needed their
own condition. SKIPPED_BY requires the gate to observe 'skipped', so
dropping the edge alone would have left them running and blocked the
release instead of failing it.
2026-09-25 12:50:53 -04:00
ethernet
29f6f63f22 docs: describe the daily canary schedule and drop hindsight from built-in providers
canary-release.yml runs on a daily schedule and on manual dispatch, not on each push to main. Hindsight lives in plugin-catalog/, not plugins/memory/.
2026-09-24 15:25:18 -04:00
ethernet
dc11e3b3bc feat(release): add --skip-bundles and --skip-tests to stable releases
`release.py release` gains two flags. They can be used together.

--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.

--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.

The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.

The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.

A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:

- The next version was derived from it, so the next cut would reuse the
  version. It now takes the newer of the R2 head and the newest
  published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
  tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
  their publication pass, so a bundle-less release would re-advance
  every 15 minutes. The head is now the newer of the R2 head and the
  published release whose final tag binds the Docker stable alias
  digest.

`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.

Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:

- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]

Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
2026-09-24 13:31:33 -04:00
ethernet
845b61f2b8 fix(release): rerun failed stable runs on the failure event, drop the cron
Stable Release Publication ran every 15 minutes (96 runs a day, each
checking out full history, setting up node and buildx, logging into
Docker Hub, and taking the release-signing environment) only because the
sequencer held a failed run for a 15-minute backoff that the failure
event could never satisfy, so the cron was what actually retried.

Drop the backoff: the reconcile pass started by a failed Stable Release
reruns its failed jobs right away. MAX_ATTEMPTS burning, oldest-first
retry ordering, the attempt-entry check, and the needs_retarget repair
stay. The schedule trigger goes; workflow_run and workflow_dispatch
remain the recovery paths.

The shared stable-release concurrency group cannot deadlock: the rerun
waits as pending behind this job, and the sequencer only confirms the
new attempt is queued before it exits and frees the group.
2026-09-24 11:50:30 -04:00
ethernet
5c062d0cc3 docs: clarify Python updater range and test sandbox marker 2026-09-24 09:20:50 -04:00
ethernet
c70129205f Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 05:10:12 -04:00
ethernet
ca06a8dec9 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/bootstrap-installer/src-tauri/src/update.rs
2026-09-24 05:03:23 -04:00
teknium1
be59f0411e fix(plugins): broadcast_plugin_event from the turn-isolation child reaches the Desktop
Review finding (MAJOR-2): with dashboard.turn_isolation the plugin's agent-side
code (tools, hooks) runs in the compute-host child, where _live_transports is
empty and _broadcast_global_event dropped the frame at logger.debug. The child
now registers its _HostTransport as a live transport for the life of run_host,
so a session-less broadcast rides the existing host pipe; the parent bridge
(_relay_compute_host_rpc) recognises an event frame with no session_id and
fans it out via _broadcast_global_event instead of write_json, which would
have dropped it on hermes serve's stdio.

Minor: the late `from tui_gateway.server import …` could raise ImportError in
a plugin-only process despite the "safe from any handler" docstring — caught
and logged as a warning.

Docs: the SDK page now tables exactly which process each call site runs in
and what it reaches (hermes serve / compute-host child / stdio TUI / gateway
run + chat + cron = nobody).
2026-09-24 02:03:03 -07:00
teknium1
3593b63563 docs(plugins): broadcast_plugin_event contract, delivery scope and rss-reader migration
The listener receives the GatewayEvent (destructure payload), names are
dotted, delivery is per process (hermes serve = every Desktop window), and
the exact rss-reader switch away from the jsonl queue + 3 s poll. One line
in the catalog trust model on the lint's regex/RegExp allowance.
2026-09-24 02:03:03 -07:00
tobenwarrior
8034d493dc feat(plugins): public event bridge for plugin backends
A plugin backend pushing an update to its own desktop half imports
tui_gateway.server._broadcast_global_event — a core-private whose signature
is not a plugin contract. Give plugin authors a public, documented door.

hermes_cli.plugin_events.broadcast_plugin_event(plugin_id, event, payload)
emits plugin.<plugin_id>.<event> on the app's global event stream (the one
host.onEvent subscribes to), with the plugin id forced into the namespace and
the bare event name validated — a mangled name would strand the desktop half
waiting on a name nobody emits. Fire-and-forget, safe from any request
handler.

Wishlist item 8 of #116305; unblocks rss-reader (#115972).
2026-09-24 02:03:03 -07:00
teknium1
c2645ef14c fix(desktop): SandboxedFrame takes an explicit prop allowlist, not iframe attrs
Review finding (MAJOR) on #120927: the props type extended
`ComponentProps<'iframe'>` and the component spread them onto the element,
so a plugin could pass `allow="camera; microphone"` — the electron
`setPermissionRequestHandler`/`setPermissionCheckHandler` grant media
capture without looking at the requesting frame's origin, so that hands a
third-party site the app's mic/camera — plus `srcDoc` (replaces the `src`
contract with caller markup), `name`, `allowFullScreen`, `csp`,
`credentialless`. Props are now an explicit interface (className, style,
onLoad, onError, ref, sandbox, src, title) and nothing is spread, so none
of those reach the DOM even through a cast.

minor: `src` is scheme-checked — only http(s)/data render, anything else
(file:, blob:, javascript:, relative) renders nothing with a console.warn,
matching what the docs already promised.

minor: the SDK surface exports only `SandboxedFrame` + `SandboxedFrameProps`;
`sanitizeFrameSandbox`/`SANDBOXED_FRAME_DEFAULT_SANDBOX` stay module-internal.

Tests: the existing posture test now also asserts allow/srcdoc/name/
allowfullscreen never land on the element; one new test covers the scheme
refusal. Both were red on 17d1b046.
2026-09-24 01:53:34 -07:00
teknium1
a48557baff feat(desktop): trim SandboxedFrame tests to invariants; document the allowlist contract
Nine change-detector tests become four invariants (one per behaviour, two
behaviours per module): every realm-escaping / unknown / mixed-case token is
dropped and an emptied set falls back to the default posture; allowlisted
tokens survive deduped; the rendered frame carries the posture; props cannot
re-open it.

Docs: TS signature, the exact allowlist and the strip list the ruling names
(allow-same-origin, allow-top-navigation*, allow-popups*, allow-modals,
allow-storage-access-by-user-activation), teardown, and the rss-reader
(#115972) migration off its stubbed /preview → openExternal chain.
2026-09-24 01:53:34 -07:00
tobenwarrior
d81db86d61 feat(desktop): sandboxed embed primitive for plugins
A reader-style plugin embeds external content by mounting a raw Electron
<webview> on the app's persist:hermes-preview partition — sharing the app's
cookies and storage. Give plugin authors the app's own guest-content posture
instead.

SandboxedFrame renders a sandboxed iframe: opaque origin, allow-scripts by
default, no-referrer, lazy. Realm-escaping tokens (allow-same-origin,
top-navigation, popups, modals, storage-access) are stripped even when a
caller asks for them — a frame with no sandbox attribute is fully
privileged, so an empty result falls back to the default posture.

Wishlist item 9 of #116305; unblocks rss-reader (#115972).
2026-09-24 01:53:34 -07:00
ethernet
9824294c8b Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 04:45:34 -04:00
brooklyn!
efafed5baf fix(desktop): Kanban attachment download through the one gated resolver
Fold the Kanban attachment download onto the existing gateway file-save
resolver (lib/media.ts downloadGatewayMediaFile / captureGatewayFileDownload)
instead of the source PR's second resolver (api/file-download.ts): one auth
shape, one owner-scope contract, for every Desktop gateway-file save.

- lib/media.ts: downloadGatewayMediaFile now accepts an explicit owner scope
  ({ connectionId, profile }) so a capture snapshot (Kanban) and the ambient
  $connection path (artifacts, chat previews) share one function. Adds
  downloadGatewayFileWithFeedback (the toast wrapper) and
  captureGatewayFileDownload (snapshots capabilityScoped() at read time).
- store/file-actions.ts: downloadRemoteFile is now a thin call into
  downloadGatewayFileWithFeedback -- the same fileMenu.downloadSaved /
  downloadFailed toasts the Files panel already used; cancel stays silent.
- plugins/kanban/drawer.tsx: rebased AttachmentDownload/AttachmentsSection
  onto main's Dialog-based drawer (aside "Attachments" section), using the
  app's boxless text-button treatment (size="inline" variant="text") to
  match the drawer's other inline actions, with a Tip for a long filename.
- sdk/index.ts: keeps captureGatewayFileDownload as the plugin SDK export
  (now sourced from lib/media, not a second module); the plugin docs keep
  it as a void-returning, non-throwing capture (toasts already fire inside).
- types.ts: keeps KanbanAttachment.stored_path.
- Deletes api/file-download.ts and its resolver; drawer/file-actions tests
  adapted for the new call shape plus a local-mode ownership case.

Root cause: the Kanban drawer only ever rendered the attachment filename
with no action, so on macOS (and everywhere else) there was no way to fetch
the stored file at all -- the backend already returns stored_path
(plugins/kanban/dashboard/plugin_api.py) but nothing in the renderer used it.

Tests: apps/desktop `npx vitest run --project ui src/plugins/kanban
src/lib/media src/api src/store/file-actions` (112 passed); the new drawer /
media-file-download / file-actions cases fail on main first (confirmed
against a scratch main worktree) and pass here. `npm run typecheck` clean.

Limits: attachments with no stored_path (older backend rows) keep the
control disabled rather than guessing a workspace path, matching the source
PR's compatibility stance.

Fixes: https://github.com/NousResearch/hermes-agent/issues/85672
Supersedes: https://github.com/NousResearch/hermes-agent/pull/107370
Supersedes: https://github.com/NousResearch/hermes-agent/pull/110161
Supersedes: https://github.com/NousResearch/hermes-agent/pull/87727

Co-authored-by: Johan Roest <229638764+jroest@users.noreply.github.com>
2026-09-24 03:42:42 -05:00
Johan Roest
d9a44a0ed4 fix(desktop): download Kanban attachments through authenticated gateway
(cherry picked from commit 1f9e64b0ed27f68885ad0a211914e7c10736c0ab)
2026-09-24 03:42:42 -05:00
teknium1
55c52f2902 fix(desktop): appearance extras use the pane boundary and render on the top-level page only
Review findings (minor) on #120912:

* `AppearanceExtraSlot` wrapped each contribution in `ContribBoundary
  variant="chip"`, so a page-level plugin card that threw collapsed into a
  bar-item chip meant for toolbar slots. Use the default `pane` variant —
  the canonical ErrorState with Retry, matching every other zone body.

* The slot was mounted unconditionally, so every deep-link subpage
  (`settings/appearance/<section>`, which shows exactly one built-in
  section) also grew the plugin cards. Gate it on `subpage === undefined`
  like the rest of the page's top-level-only chrome.

Tests: the boundary test now asserts the pane fallback (red on c322f692);
one new page-level test renders `AppearanceSettings` with and without a
subpage and asserts the extra mounts only on the top-level page (red on
c322f692). Docs updated to the new contract.
2026-09-24 01:34:59 -07:00
teknium1
84ac69e884 feat(desktop): keep ColorSwatches export as on main; document APPEARANCE_AREAS.extra contract
Drop the salvaged ColorSwatches re-export hunk: main already exports the
component from the SDK, and the plain-JS plugins this slot serves do not need
the props type, so the shared sdk/index.ts hunk stays a single added line.

Docs: TS signature, arbitration (every registration mounts in its own
boundary — no first-wins, plugins cannot suppress each other), teardown
(loader-owned disposer, nothing persisted) and the exact migration for
better-session-appearance (#115961) and hermes-appearance-hub (#116049).
2026-09-24 01:34:59 -07:00
tobenwarrior
8d71403ffa feat(desktop): appearance-settings slot and the app's swatch grid for plugins
A plugin adding appearance controls injects nodes into Settings → Appearance
and drives the app's widgets through React internals. Give it a seam.

APPEARANCE_AREAS.extra renders contributions at the end of the Appearance
page (own error boundary each, chip fallback), and ColorSwatches — the grid
the profile rail and project dialog already use — joins the public SDK
surface with its props type, so a plugin picks colours with the app's own
control and its own onChange (pair it with host.sessions.setColor for
session colours).

Wishlist item 7 of #116305; unblocks better-session-appearance (#115961).
2026-09-24 01:34:59 -07:00
teknium1
4cec3f85e7 fix(desktop): host.pluginDecisions hands out frozen copies, not the live store object
Review finding (MAJOR) on the typed bridge: `pluginDecisions.get()`,
`.value` and the `subscribe`/`listen` callback argument returned the
store's own object. `pluginActive()` reads that same reference and
`saveDecisions({...$pluginDecisions.get(), [id]: enabled})` spreads it,
so `host.pluginDecisions.get()['other'] = false` silently disabled
another plugin and the next toggle persisted it — the declined `set()`
by another door.

Every value the read-only view hands out is now `Object.freeze({...v})`;
assignment throws under strict mode and the store is untouched. The
existing read-only test gains the mutation assertion (red on c1b740ac).
Docs: frozen-copy contract spelled out; cheat-sheet signature for
`host.profiles.list(scope?: ProfileScope)` made explicit.
2026-09-24 01:25:11 -07:00
teknium1
7daf8180f7 docs(desktop): capabilities bridge section; pluginDecisions documented read-only
Signatures, the profile-scoping rule, the WHY of the declined set() writer,
teardown note, and the exact better-capabilities migration table.
2026-09-24 01:25:11 -07:00
tobenwarrior
bd55633248 feat(desktop): typed skills/toolsets/profiles bridge for plugins
Plugins configuring capabilities called window.hermesDesktop.api raw and
read/wrote the persisted pluginDecisions map directly. Give them the same
doors the Capabilities page uses.

host.skills.list/setEnabled, host.toolsets.list/setEnabled, and
host.profiles.list wrap the app's own api modules — same endpoints, same
profile scoping (a ProfileScope configures any profile without swapping the
app-wide active one) — and host.pluginDecisions exposes the decisions map
with set() going through the live toggle (deactivate/activate included),
never a raw storage write.

Wishlist item 6 of #116305; unblocks better-capabilities (#115960).
2026-09-24 01:25:11 -07:00
teknium1
6b2c23ae42 fix(desktop): model-pill provider must return a string; non-strings fall through
Review finding (MAJOR) on #120919: useComposerModelPillLabel accepted any
truthy return with `if (label)`. ModelPill drops the value straight into JSX
with no error boundary, so a plugin returning an object/array/number threw
"Objects are not valid as a React child" and blanked the whole composer —
exactly the failure mode the hook's throw-swallowing was meant to prevent.
Only a non-empty, non-whitespace string is now a label; anything else
declines like null.

Review finding (minor): the two existing tests are tightened instead of
adding new ones. The compact-mode "providers skipped" assertion was
`queryByText(...)` on a chevron-only render and passed regardless; it is now
a spy that must not be called. The throw test now also covers the
non-string return (red on b559ac9c: React child error) and two non-null
providers: the first registered string wins and the later provider's spy is
never invoked.

Docs: arbitration section states the string-only contract and that
`reasoningEffort` is always a string ('' when none).
2026-09-24 01:12:58 -07:00
teknium1
1b77007a36 docs(desktop-sdk): model-pill label providers — signatures, arbitration, migration
Spell out the contract plugins will lean on: the context fields, the
"registry order, first non-null wins, throw = decline" rule, that compact
mode never consults providers, and that teardown is the contribution's
own disposer (nothing for ctx.onDispose). Include the exact migration for
compact-reasoning-label, noting that the core label no longer carries the
effort word, so the plugin's regex strip is already a no-op and the hook
is the place to compute a label rather than rewrite rendered text.
2026-09-24 01:12:58 -07:00