The M13 fix stopped counting any key a tool panel also asks for, since a
bare PUT /api/env cannot tell connecting a provider from configuring a
tool. Desktop onboarding and model settings now send provider_setup=true
and the server counts those saves; the Keys page and tool panels keep
the exclusion.
m13: the renderer stored and sent raw plugin palette/keybinding ids, rage-click
targets and config paths below user-named containers (providers.<name>.api_key);
the backend collapsed them, but only after truncating the action list to 200.
- recordAction collapses anything outside the built-in ids (KEYBIND_ACTIONS +
KEYBIND_READONLY + the named buttons, minus numbered slots — exactly the
backend's DESKTOP_ACTION_IDS) to `other` before it is counted or used as a
rage-click target.
- recordSettingsSaved sends only keys the config schema publishes (the Settings
page passes its schema: DEFAULT_CONFIG leaves), so nothing below a user-keyed
container leaves or is kept.
- Backend DAILY_ACTION_ROWS_MAX = |DESKTOP_ACTION_IDS| x |vias| (296): every
distinct collapsed (action, via) fits, so none is cut before collapsing.
m14: first-run steps happen before the consent question and collection defaults
off, so every step but consent/first_message was dropped. Chose the in-memory
buffer (smaller than removing the steps from contract, schema and TS): while the
answer is unknown/undecided, recordOnboarding/closeOnboardingStep calls are held
in memory (max 100, never persisted or sent) and replayed on the switch turning
on; a decided "off", a profile switch or quitting discards them.
setDesktopMetricsGate takes `decided` (hook passes consent.decided).
Tests (desktop-metrics.test.ts): recordSettingsSaved callers pass a schema
(contract change). New invariants, RED on base: plugin action id / rage target
stored and sent as `other`; a providers.<name> key is not sent; pre-consent
onboarding held, sent on yes (with closes applied), dropped after a decided no.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, raw ids):
before: localStorage actions {"acme-plugin:/home/alice/clients/bigco|palette":1, ...|click:3};
wire rage_click target "acme-plugin:/home/alice/clients/bigco", setting
"providers.acme-corp-internal.api_key"
after: localStorage actions {"other|palette":1,"other|click":3}; wire
[rage_click target "other"] only
M9 + the renderer half of M10, and m15.
M9: the renderer kept one install-wide localStorage record and bound the new
gateway before reading its consent, so a day accumulated under profile A was
flushed to B before B's switch was read (or dropped by B's off), and A's area
latch suppressed B's first use.
- The record is keyed per (connection, profile): `hermes.desktop.metrics.v1:
<FNV hash>` (no profile name in the key). bindDesktopMetrics takes the
profile scope; a different scope forgets in-memory state and nulls the gate,
so nothing is kept or sent until that profile's switch is read.
- The hook pins its requester to the focused (connection, profile) with
requestGatewayForAgent instead of the session-tile router, so the RPCs land
in the profile whose switch gates them. main receives the same scope hash.
M10 (renderer): each window cached the record in module memory and the last
debounced persist() won, so peer windows overwrote each other's counts. Every
change now reads the record fresh and writes it back (no debounce, no cache;
persistDesktopMetricsNow is gone). withState calls are no longer nested
(noteMessageSent / abandoned-step reporting read first, record after), since a
nested write would be overwritten by the outer one. A daily flush is pinned to
the scope and requester it started with. Both windows may still send the same
finished day; the backend's durable latch (earlier commit) records it once.
m15: consent was re-read only on attach or profile switch. The hook re-reads
it on window focus; Settings › Privacy applies its save to the gate when it is
scoped to the focused profile by name, not only when unscoped.
Tests (desktop-metrics.test.ts): `stored()` now finds the per-profile key and
setEnabled is expected with the scope argument (contract change). New
invariants, RED on base: profile switch A→B→A keeps/sends nothing of A under B
and A still reports its day; two module instances (windows) of one profile sum
into one day.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, adapted to scopes):
before: B received before its consent read: [["shared_metrics.desktop_daily",{...view.toggleSidebar 4...}]]
areas latch A: [terminal_pane] B: []
two windows stored: {"composer.send|shortcut":4} (W2's session.new lost), W2 daily sends
composer.send 3 + session.new 2
after: B received before its consent read: []; areas latch A: [terminal_pane] B: [terminal_pane]
stored: {"composer.send|shortcut":4,"session.new|shortcut":2}; W1 and W2 send the same
aggregate (the backend latch keeps one)
M11. RendererCrashRecorder held one install-wide `enabled` flag that every window
overwrote with its own profile's switch, so an opted-out window's crash was
written to disk and reported into another profile, and an opted-out window
wiped the opted-in profile's pending crash.
- Consent is kept per webContents id (the IPC sender), with a hash of the
window's focused profile key; set-enabled now carries that key.
- A crash is recorded only when the crashed window's own profile collects;
each pending entry is tagged with the profile hash (file v2: {entries:[{profile,
reason}]}; an untagged v1 file is ignored — whose it was is unknown).
- take() returns only the calling window's profile's reasons; one claim per
profile (the same window may re-claim after its renderer reloaded); ack drops
only the claimed count. Turning a profile off drops only that profile's entries.
- main.ts passes the window's webContents id to recordRendererGone.
Tests: renderer-crash-metrics.test.ts moves to the per-window API (contract
change: every call names its window, set-enabled names the profile); one new
invariant test (opted-out window records nothing, another profile neither
drains nor purges) — the base API cannot express it, so RED is shown by the
reviewer's probe instead.
Probe (review/desktop/jsrun/.../renderer-crash-metrics.probe.test.ts, adapted to
window ids):
before: after opted-out window crash, file exists: true {"v":1,"reasons":["crash"]}
take by window 2: {"reasons":["crash"]}; A crash record after B-window off: false
after: file exists: false; take by window 2: null; A crash record after B-window off: true
CLI setup and the Desktop consent dialog add one 'how Hermes gets used'
line covering the v5 harness, efficiency, engagement, Desktop feature/
dislike and provider-setup counters, and say setting values never leave.
No new UI: counters hang off interactions the app already has.
- feature_use: panes (terminal/file/browser/review), overlays (palette,
model/session pickers, switcher, find), voice, Bot Mode (the `bots`
workspace), skins, projects, full-page routes and each settings view.
- action_use: keybinding dispatch (shortcut), palette items (palette),
titlebar/sidebar tools that carry a keybinding id, composer send/stop,
voice, message copy/retry (click, via a typed const table), and the
native Open Folder menu (menu). Aggregated per day in the app and sent
once per finished day.
- mode_use: Sessions vs Bot Mode from $workspaceMode / the tile's
workspace scope; active time = interaction gaps <= 5 min.
- friction: user toast dismissals by closed notice id, error toasts by
their code-defined summary category, backend drops (settled 3s so a
Restart/switch cancels), long frames while visible, renderer crashes
drained from main.
- dislike: quick close (<5s), cancelled flows, settings saves (keys
only), rage clicks, undo (restored draft, reopen closed tab), shipped
features switched off.
- onboarding: transitions of the existing first-run stores.
All of it follows the focused profile's collection switch: off (or not
yet known) means nothing is kept in localStorage and nothing is sent,
and switching off purges the local record and tells main to drop its
crash file.
A renderer crash takes the renderer (and its backend socket) with it, so
the renderer cannot report it. Main now observes render-process-gone on
live windows (never user close or intentional teardown), persists only
the bucketed reason (crash/oom/killed/other) under userData while the
renderer has reported the user's opt-in, and hands the list to the
reloaded renderer via take/ack IPC, mirroring the pending update-run
record. Turning collection off deletes the pending file immediately.
Desktop love/hate telemetry needs a backend half: six counters
(feature_use, action_use, mode_use, friction, dislike, onboarding) with
closed dimension sets, the v3 JSON schema entries, the public doc rows,
and profile-scoped gateway RPCs the app reports through.
Every value the app sends is collapsed server-side onto the closed set
(unknown -> other); feature use latches per area per UTC day, friction
and dislike are capped per day, the daily report (action counts + mode
time + bot count) latches per day and answers recorded=false on a store
failure so the app keeps the day for a retry. For setting changes the
app sends only the config key: the backend reads the saved value and
compares it to DEFAULT_CONFIG itself, so values never cross the wire;
keys outside DEFAULT_CONFIG read other. All helpers check enabled()
before doing any work.
The v4 startup-latency IPC and packaged-update recorder added ~26 lines to
main.ts (18k lines). desktop-shared-metrics.ts now owns the claim IPC, the
recorder construction and its renderer IPC, and the packaged-only apply
wrapper; main.ts keeps one import, one constructor call and two call sites.
The recorder's behaviour is unchanged (update-metrics.ts untouched).
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat
call added a row. The TUI and Desktop now send an opaque per-launch id
(TUI: one per process; Desktop: minted when the once-per-launch main-process
claim succeeds) and the backend claims (surface, launch_id) once per
process: a reconnect re-sending the same launch counts once, a new Desktop
launch against a long-lived backend still counts. Legacy clients without an
id count once per backend process and surface. The id is only a local latch
key (bounded set, length-capped), never recorded. Only a usable measurement
spends the claim. Contract + apps/shared regenerated.
- m10: relaunch() execs in place, keeping the PID, so `sessions browse` ->
resume counted the picker time as CLI startup. relaunch() stamps its PID in
HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it
matches, while children (other PIDs) inheriting the env still count.
- m8 (psutil half): record_process_ready checks the opt-in gate before reading
the process start time (psutil) or starting a thread. The claim is taken
first so a later opt-in never records a stale "startup" mid-session.
Test changes to existing assertions: the TUI/Desktop param assertions now
expect launch_id (a new wire param); the RPC surface test sends distinct
launch ids so its three calls stay three launches under the new latch.
CLI setup and the Desktop consent details now list update results and timing,
crashes, startup/reply speed, messaging-platform health and the coarse machine
facts (RAM range, GPU type, version age), so the opt-in describes everything the
v4 counters record. All Desktop locales updated.
Desktop's packaged updaters (electron-updater, App Installer, Store) never run
`hermes update`, so the receipt cannot count them. Main persists a pending
record before apply (the installer may quit the app before apply returns),
decides a hand-off on the next launch (version changed => success), and the
renderer sends it once the gateway is open; the file is removed only after the
RPC resolved. Checkout hand-offs are excluded: their receipt already counts
them as kind=desktop. The backend buckets raw words into bounded dims.
The backend buckets and records hermes.startup.latency, but only the clients
know when they were launched. The Ink TUI reports Node process uptime at its
first gateway.ready; Desktop reports Electron main-process uptime when the
renderer's gateway first connects. The Desktop latch lives in the main process
because renderer reloads, extra windows and backend reconnects all re-run the
renderer boot. Both calls are fire-and-forget and swallow errors so older
backends without the method are unaffected.
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.
hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.
The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.
hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.
hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.
hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
useSlashCommand's dispatcher entry fires shared_metrics.slash_command once
per user-typed command, before resolving its surface, so locally handled
commands (pickers, /skin, /journey, /new, /branch, /resume) count as well
as the ones that reach slash.exec. Fire-and-forget: errors (an older
gateway without the method) are swallowed and nothing waits on it.
Not counted: the alias re-dispatch (runSlash(..., recordInput=false)) and
programmatic calls, which pass typed: false (the transcript's Compress
button, through the tile delegate's new executeSlash option).
Existing dispatcher tests that assert the exact gateway traffic of a
command now ignore the metrics call.
The gateway counted slash commands at its client boundary (successful
slash.exec / command.dispatch), so every command the Desktop or Ink TUI
handles locally (/new, /branch, /skin, /resume, overlays, pickers) never
reached shared metrics. Counting moves to the clients' dispatcher entry:
the new fire-and-forget RPC records {command, execution_surface} for one
user-typed command (surface desktop|tui from the same client detection),
scoped to the session's profile, always answering {ok: true}.
_count_slash_command is removed from rpc_dispatch so each typed command
lands exactly once, from the client. Contracts regenerated.
The first-run consent dialog blocked the composer on every undecided
profile, so a healthy install no longer opened straight to chat (caught by
the Desktop first-run E2E). The question now lives in the composer status
stack beside the free-tier strip: Send to Nous / Local only / No thanks /
Details, no focus steal, one owner across split composers, one offer at a
time. Details opens the full explainer on request; closing it decides
nothing. The backend's `decided` stays the only latch.
Desktop had no way to read or change the shared-metrics opt-ins. The new
profile-scoped RPCs read/write telemetry.shared_metrics.{enabled,send} in the
focused profile's config.yaml through the one config writer, keep the setup
wizard's invariant (send is forced off unless collection is on) and reconcile
the consent windows on every change via the wizard's own helper. first_run
answers also record the desktop setup-completed metric (lazy import until the
events module lands).
Issue #63125 asks for the cached vs uncached read rate at a glance in
the model picker. The backend already ships it: _apply_pricing fills
ModelPricing.cache from input_cache_read, and the CLI picker renders it
as a third column — the catalog menu was the only surface dropping it.
- ModelPrice appends the cached rate (dim, `·$0.16`) next to the
uncached input/output pair; its hover title names it, and the row's
tooltip gains the cached-read line. Uncached read stays the inline
`input` rate, so the cached-vs-uncached comparison reads at a glance.
- The hardcoded "free" label and English-only tooltip move to
shell.modelMenu copy (free / cacheRead / priceTitle) so the row
localizes; fr/de/es are completeness-checked, so all locales that
carry the modelMenu block get the new keys.
- A collapsed `-fast` family now falls back to the fast sibling's price
when the base id is unpriced (the sharing the comment already
claimed — review point on #84831).
Refs #63125
Review follow-ups on #84831:
- Treat null input/output as unknown: render nothing rather than
"null/null" when a provider reports a non-free model without rates
(_apply_pricing ships "" for unknown, but the wire type is
string | null, so guard both).
- Only show the sale tag for a positive discount_percent, so a negative
value can't render as a double-minus.
- Fallback title uses "—" for a missing half instead of "undefined".
- New test: no price span at all when the provider carries no pricing
for the model.
The model.options payload has carried per-model ModelPricing for providers
that support live pricing (Nous Portal included) since the picker requested
it, but the catalog menu never displayed it. Add a compact price label to
each model row: input/output $/Mtok, a green "free" label for free-tier
models, and a sale tag when the portal reports a discounted list price.
Verified against a live payload probe: the nous provider ships 35 pricing
entries (e.g. anthropic/claude-sonnet-5 -> $1.60/$8.00, 20% discount).
fr/de/es overlays are completeness-checked against the en key set, so the
new updates.copyFullLog key must ship there or catalog-completeness fails;
ar/ru carry the neighbouring moreChanges key, so they translate too.
Swap the hand-rolled copied-state button for the shared CopyButton
primitive (haptics, aria-label, error path, unmount cleanup included).
The update overlay previously showed at most 6 commit subjects from the
changelog, with no way to access the full list or copy it for translation.
Changes:
- Increase git log limit from 40 to all commits (remove -n 40)
- Add 'Copy full changelog' button that copies all commits as formatted
text (type(scope): subject - author) to clipboard
- Use parseCommitHeader to reconstruct full convention commit format
including scope and author
- Add i18n keys: en 'Copy full changelog', zh '复制完整更新日志'
Closes#48654
The first-pick bug was the Select's controlled value already being the
Custom sentinel, which the value swap fixes on its own; the Input's
autoFocus handles focus once it mounts, so the rAF refocus effect is not
needed. Extend the regression test to type into the field and to check
that clearing it and blurring returns to the dropdown.
Live settle addresses a folded tool turn by its final row, hydration by its
first, so after a tool-using turn the tips never matched and the first send
bounced with "Chat out of date" in a single window.
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
The banner was text-xs amber-200 — easy to miss in dark mode and near
illegible in light mode. It now uses the app's warn-badge contrast
pair (text-amber-600 dark:text-amber-300) at text-sm with a size-4
icon, stronger border/background, and medium weight, so it reads at
a glance in both themes. The persistent variant gains a dismiss button
that records the fingerprint acknowledgement; the post-switch notice
stays undismissable because it announces a change that just happened.
Co-authored-by: kyssta-exe <25470058+kyssta-exe@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
A dismissal is bound to the exact pin configuration it acknowledged —
main provider plus every slot's task/provider/model/base_url — stored
per profile. Any slot edit, main switch, or endpoint repoint changes
the fingerprint and re-arms the banner, so the silent-credit-burn
protection is never lost. base_url is normalized (trimmed, trailing
slashes stripped) so equivalent endpoints share a fingerprint, and
StaleAuxAssignment grows an optional base_url the persistent banner's
mapping now carries (switch echoes still omit it).
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
The queue panel showed each queued prompt as a single truncated line;
reading a long entry required loading it back into the composer. Previews
now wrap across two lines, and long or multiline entries get an in-place
expand/collapse toggle (aria-expanded chevron) that reveals a scrollable
full-text view. Edit/send/delete actions are unchanged.
Adds a QueuePanel test covering toggle presence (long/multiline only) and
the expand/collapse round-trip. Copy added in every locale that carries
the neighbouring queue keys.
Original work by YonganZhang in
https://github.com/NousResearch/hermes-agent/pull/46007, rebased onto the
current TypeScript sources (Codicon + iconSize tokens, steer/resume
actions preserved).
Co-authored-by: Yongan Zhang <374456248@qq.com>
Fixes https://github.com/NousResearch/hermes-agent/issues/45664
The composer model pill capped its width at max-w-40 (160px), so long
model names with session-state suffixes ("DeepSeek V4 Flash · Med")
truncated with an ellipsis even when the row had room. Drop the cap: the
pill sizes to its label, and the existing shrink + min-w-0 + truncate
squeeze path still yields width between collapse stages, so long names
only clip when the row is genuinely out of space.
Extends the model-pill test to lock the unbounded-width behaviour.
Original idea by @Aio777 in
https://github.com/NousResearch/hermes-agent/pull/49341 (raised the cap;
this removes it as requested by the issue).
Co-authored-by: aryan <alokesh1@sheffield.ac.uk>
Fixes https://github.com/NousResearch/hermes-agent/issues/49340
LogTail gains in-place search: matches are highlighted, the active one is
scrolled into view, and a "3/6" counter with up/down (Enter, Shift+Enter,
Cmd+G) steps through them, reusing the find bar's key handling and labels.
The Command Center system log and the MCP log pane share it, so there is one
log component instead of a Command Center fork.
Clearing a search re-engages follow-the-tail instead of leaving the pane
parked where the last match was.
WARNING/ERROR/CRITICAL lines get a thin gutter mark, read from the level
column of the log format, so an INFO line mentioning ERROR is not flagged.
Command Center sections use the same breadcrumb header as Settings.
Co-authored-by: Sonnenwerk <a.majuskel@gmail.com>