192e574efef1845ee81bbc182af803ff2432f04f
653 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6397a84a91 |
fix(sessions): un-hide an auto-archived chat when it is resumed or compressed
The idle sweep archives a whole compression lineage. When the chat later resumed and compressed, the new tip was inserted with archived=0 under the archived root, and the sidebar admits a lineage by its root's flag, so the live chat stayed hidden. Record sweep provenance in a new sessions.auto_archived column. Publishing a compression child or reopening a session under a sweep-only archive un-hides the lineage; a deliberate archive (sidebar, API, CLI) clears the provenance and keeps the chat hidden. Compression children now inherit the parent's archive state so a manually archived lineage stays uniform. Fixes #117713 |
||
|
|
dd66c132f2 |
fix(sessions): keep tool-call uids paired through rewrites, folds of repeated ids and chunked copies
Found by the pre-merge correctness and hermes-specific review of this stack: - A fold after a row whose one response repeated a provider id (one shared uid) appended the absorbed turn's uid in the wrong slot. The shared uid now fills each of the row's own occurrences before the fold appends (per_occurrence_tool_call_uids). A single response keeps one shared uid for a repeated id: every result carries it, so no call looks unanswered. - A digest-less re-flush of a restored fold (replay heal, adopt after a digest mismatch) let the stored pre-fold map overwrite the longer live list, so the later occurrence's results paired with nothing. A live list that starts with every stored occurrence is kept (stored uid or list). - A dict that lost its uid map (a clone) rewrote tool_call_uids to NULL. The row-addressed rewrite now fills stored uids for the calls the dict still names; a call it dropped keeps none. - append_messages_batch(chunk_rows=...) reset the pairing index per chunk, so a result in the chunk after its call was stored without tool_call_uid (branch/seed copies). The whole batch is paired before it is chunked. - _merge_assistant_into keeps its dict guard on the absorbed turn's map (a plugin-built dict with a non-dict map no longer raises). |
||
|
|
5d8b17f865 |
fix(sessions): an assistant fold that discards the later text records no merge witness
_merge_assistant_into never joins multimodal (list) content, so a non-empty later turn beside a list (or a list beside non-empty text) is dropped, yet the fold still recorded its uid in _absorbed_message_uids, a false claim that its text lives on. The merge now reports whether the later text was kept and the witness follows it; the retired row id is still recorded. Reported by @andrexibiza on #126307. |
||
|
|
d9808867c7 |
fix(sessions): an assistant fold keeps every occurrence uid of a reused provider id
Folding two assistant turns that reuse a provider id (llama.cpp-style constant ids) leaves
both calls in tool_calls, but the {id: uid} union kept only the later uid, so the first
occurrence lost its identity. The id now maps to a list of uids, one per occurrence in
tool_calls order (merge_tool_call_uids); index_tool_call_uids pairs each call with its own
occurrence, so a following result still pairs with the nearest call. Provider ids are never
rewritten. Persisted and restored as-is (the JSON map accepts list values).
Reported by @andrexibiza on #126307.
|
||
|
|
9d680cb95c | docs(sessions): the bounded legacy backfill and branch copies keeping message_uid | ||
|
|
c85beecacd |
fix(sessions): a superseded verification candidate is not a merge-witness constituent
Alternation repair replaces a provisional verification candidate with the final answer instead of folding it. Its row is still retired through _absorbed_row_ids, but its uid must not appear in the replacement's _absorbed_message_uids: the witness names text that lives on in the composite. (cherry picked from commit fd284d1fd35371215a616efab965ef117348f3c0) |
||
|
|
c2e2150d78 |
fix(sessions): close the identity gaps the review found and prove the recovery paths
- Strip persistence-only fields from a select_context() selection: the request copy was stripped before the hook, but an engine handing back the conversation_messages clones put timestamp, _row_id and the new ids on the wire (pre-existing for the old fields). - Every row carries a message_uid whoever wrote it: an AFTER INSERT trigger mints one for a row inserted without it (an older build writing into a v31 store); the one-time backfill is gated by its own state_meta marker so a store whose schema version cannot advance (no FTS5) never rescans. - One merge-witness helper (record_absorbed_message) for every fold of a durable message: alternation repair's user and assistant merges, the compressor's in-flight restatement onto the carrier, the real user anchor folded into a scaffolding turn (the anchor's uid leads), micro-compaction's adjacent-user merge. - tool_call_uids follows tool_calls: it is a payload column, so an assistant merge's unioned map is rewritten with the row instead of being overwritten by the stale stored map on adoption. - A result pairs with the nearest call naming its provider id: each assistant row's ids shadow earlier occurrences in the batch, restore and live-list indexes, a call row without a map leaves its result unpaired, and nothing pairs across a user turn. - A composite rewind returns and installs the replacement row's uid on the live scaffold. - Tests: import_foreign_history, hermes sessions recover and backup import carry the ids; the ACP restore test accounts for the uid on restored rows. (cherry picked from commit b30636a49cb29fdd21ba4100eef9d810897338be) |
||
|
|
76d0ac6caa |
feat(sessions): give tool calls and their results a per-occurrence uid
Provider tool-call ids repeat: Hermes mints deterministic call_<12hex> ids for identical calls and models
reuse ids, so tool_call_id cannot identify an occurrence. An assistant row now carries
messages.tool_call_uids, a {tool_call_id: uuid4} map for its tool_calls, and each tool-result row carries
the matching messages.tool_call_uid. The provider-facing tool_calls JSON is untouched, so row identity,
display identity and the CAS digest are unchanged and nothing nested reaches the wire.
Minted at the assistant row's first insert; paired onto the result in the same batch, from the live list
when the result is flushed later, or on restore from the preceding assistant row; copied by every clone
and re-insert; restored as the live _tool_call_uids / _tool_call_uid; stripped from provider requests.
(cherry picked from commit aa49c0f59788676d4beca288181711c558fca165)
|
||
|
|
f51d0a7233 |
feat(sessions): persist the consecutive-user merge witness on the survivor row
A merge survivor's _absorbed_message_uids lived only in memory, so an engine that restarted after a merge was back to parsing "\n\n" to learn which stored messages a composite user row contains. Store the list in messages.absorbed_message_uids (an owned column: the survivor's row-addressed rewrite and every re-insert carry it), restore it as the live key, and round-trip it through export/import. (cherry picked from commit 0bc05f2656229b1010f330e1998684b04d5f963f) |
||
|
|
751d8526e3 |
feat(sessions): give every persisted message a durable message_uid that reaches context engines
The physical messages.id is re-issued by every copy (in-place compaction generations and their concurrent-tail clones, rotation children, replace_messages), and _row_id is opt-in on restore, so a context engine that keeps its own per-message state had nothing stable to key on across a restart or a compaction boundary and re-identified rows by content and timestamp. Add messages.message_uid (uuid4 hex), minted once at a row's first insert and stamped on the caller's dict, copied by every SQL clone, kept by every re-insert of the same dict, left alone by row-addressed rewrites, restored on every projection, and stripped from provider requests. A consecutive-user merge survivor keeps its own uid and records the absorbed uids in _absorbed_message_uids. Schema v31 backfills a uid onto rows that predate the column. (cherry picked from commit 20d67ffb5e50e0ed956e56ddc7f8e28d36059e2e) |
||
|
|
9bcbe7b5df |
feat(i18n): pluggable, layered language packs across core, Desktop and TUI (#126296)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales
* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter
ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.
* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs
* chore(tui): split en catalog siblings by lane (slash sibling)
* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports
* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts
* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list
* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes
* feat(plugins): report language-pack layers in the mid-run activation summary
* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n
StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.
* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter
- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
layers partial packs (nested or flat dotted) over bundled/en via
mergeTranslations; a string over a function-valued en entry becomes a
positional {0}/{1} formatter; $appLocaleVersion bumps so translators
re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
registry as source 'backend' (method-not-found is silent); re-synced on
socket open, display.language change and profile switch. A saved pack-only
language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).
* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)
Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.
* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()
- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces
* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)
* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()
Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.
locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).
* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()
- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions
* i18n(platforms): route Google Chat and Teams user-facing text through t()
Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.
Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.
* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()
LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.
* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)
Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.
* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)
_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.
* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)
RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.
* i18n(cli): live-work dock, subagent monitor and render copy through t()
cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.
* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()
* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)
get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.
* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output
* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time
Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).
* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)
cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.
* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()
- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
at call time; category labels via slash.category.*; help/alias/usage suffixes
via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
keeps its English constant; callers use history_unreadable() ->
gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.
* i18n(telegram): route adapter chat copy through t()
Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.
Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.
* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()
- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor
* i18n(discord): route adapter chat copy through t()
Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().
* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)
* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()
- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
_DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
Updating/Generating) are one full template per variant; plurals use
<key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).
* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English
* test(i18n): pin Telegram/Discord adapter catalog wiring
Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.
* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge
* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint
* i18n(tr): translate bundled catalog + tui pack
* i18n(ja): translate bundled catalog + tui pack
* i18n(ko): translate bundled catalog + tui pack
* i18n(zh): translate bundled catalog + tui pack
* i18n(fr): translate bundled catalog + tui pack
* i18n(af): translate bundled catalog + tui pack
* i18n(uk): translate bundled catalog + tui pack
* i18n(ar): translate bundled catalog + tui pack
* i18n(pt): translate bundled catalog + tui pack
* i18n(it): translate bundled catalog + tui pack
* i18n(es): translate bundled catalog + tui pack
* i18n(zh-hant): translate bundled catalog + tui pack
* i18n(ru): translate bundled catalog + tui pack
* i18n(hu): translate bundled catalog + tui pack
* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)
* i18n(de): translate bundled catalog + tui pack
* i18n(ga): translate bundled catalog + tui pack
* test(i18n): fixture matches _normalize_lang(lang, home) signature
* i18n(tui): scaffold userMessages/slashCmd en siblings
* i18n(tui): wire secure prompts + content tables
* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays
Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.
* i18n(tui): wire slash ops/wake replies
* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)
* i18n(tui): wire slash core/debug/setup replies
* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)
* i18n(tui): wire slash session/topup/subscription replies
* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)
* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys
* i18n(tui): wire userMessages copy through the userMessages namespace
* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json
* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)
Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).
* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)
* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)
* tui: i18n-export-en script (English templates for pack translators)
* docs(i18n): bundled TUI packs are the bottom layer of the tui surface
* i18n(ru): translate TUI pack
* i18n(ar): translate TUI pack
* i18n(es): translate TUI pack
* i18n(pt): translate TUI pack
* i18n(ko): translate TUI pack
* i18n(de): translate TUI pack
* i18n(ja): translate TUI pack
* i18n(fr): translate TUI pack
* i18n(tr): translate TUI pack
* i18n(it): translate TUI pack
* i18n(zh): translate TUI pack
* i18n(zh-hant): translate TUI pack
* i18n(hu): translate TUI pack
* i18n(uk): translate TUI pack
1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.
Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).
* i18n(ga): translate TUI pack
* i18n(af): translate TUI pack
* plugin_guard: locale catalogs in language packs step down the agent-config family
A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.
* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)
* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header
Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.
* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel
- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)
* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})
* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property
* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal
* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export
The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.
* test: unbreak two main-red timing tests the PR merge-ref inherits
- test_local_runtime racing fake publishes the modern state record (legacy pid-only
records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
a fixed 0.35s, which a loaded CI runner does not always meet
* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)
* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak
SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.
* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings
* chore(i18n): regenerate desktop key catalog after main sync
---------
Co-authored-by: Teknium <teknium@nousresearch.com>
|
||
|
|
9b7db3105b |
fix(telemetry): fewer rows per busy day and a smaller local store
A volume simulation through the real store showed two costs worth cutting. Rows: hermes.task_run.finished carried duration, retry, model-call and tool-call buckets on top of outcome/end reason/surface, so almost every task became its own row; hermes.tool_call.count did the same with latency and retry buckets. The terminal rows now keep only the fields they are filtered by, and the same end events feed two small split counters: hermes.task_run.duration (surface, outcome, duration, retries) and hermes.tool_call.latency (tool category, latency, retries). Per-task call counts already ride on hermes.task_cost.count. v3 still accepts the v2 field sets of both counters so rows recorded before an upgrade drain. Local store: every package lived three times for 30 days (counter rows, the payload column, a pretty-printed outbox file). Outbox files are now compact JSON, and once the ingest has accepted or refused a package the database drops its copy of the body; the file stays as the local history the docs promise, and pending packages keep their body for resends. Measured (simulated day, skewed draws, before -> after): heavy 2,788 -> 2,347 rows, 21.0 -> 17.7 KB on the wire, 102 -> 44 MB local over 30 days extreme 6,561 -> 5,047 rows, 46.0 -> 33.6 KB, 272 -> 102 MB typical 663 -> 650 rows, 7.1 -> 6.8 KB, 24 -> 13 MB |
||
|
|
1763066334 |
fix(telemetry): Azure deployment names leave only when they are a public model id
Azure calls a model by a deployment name its owner chose, which can name a team, customer or environment. On azure* providers the model id now passes only when it is in Hermes' static catalogs or the local models.dev cache (never a network call); anything else reads custom. |
||
|
|
2bb227c596 | docs(telemetry): model rules name AWS ARNs and loopback provider aliases | ||
|
|
3ba7c850ba |
fix(telemetry): Desktop provider-connection forms count Gemini/xAI/DeepInfra keys again (M13)
The M13 fix stopped counting any key a tool panel also asks for, since a bare PUT /api/env cannot tell connecting a provider from configuring a tool. Desktop onboarding and model settings now send provider_setup=true and the server counts those saves; the Keys page and tool panels keep the exclusion. |
||
|
|
d19b204b33 |
fix(telemetry): relay-shared-metrics doc — valid gateway table, owner rulings
- m10: the v5 relay-platform note sat inside the gateway metrics table, so every row after reply_latency (cron.run, startup.latency, update.*, process.exit) rendered as one paragraph. The note now follows the table. probe_e_doc_table.py before: cron.run/startup.latency/process.exit "not found as cell"; after: all inside <td>. Whole doc: 49 metric rows, rendered table rows 44 -> 49. - platform.health for the relay connector stays `relay` (one socket fronts several platforms): documented next to the relay note (owner ruling). - Delegated subagents count in the harness metrics and in cache_break / tool_output_truncation (model and tool behaviour); background review and curator never do (owner ruling; matches probe_internal.py). - Model-route rules: lmstudio keeps its provider name and only its model reads `custom` (behaviour kept per owner ruling); the local-alias list now matches custom_provider_aliases() (adds llama-cpp). |
||
|
|
d8d9639d8e |
fix(telemetry): execution_backend skips background review / curator fork calls
hermes.execution_backend.count still counted terminal/browser/code calls made
under the background_review write origin (the review and curator forks),
while hermes.terminal.outcome and file_edit already skip them. Owner ruling:
match terminal.outcome. record_execution_backend now returns early when
is_background_review() is set (one contextvar read, before any work).
Doc: the execution_backend row lists the forks as not counted.
Probe probe_internal.py (terminal_tool inside a background-review origin):
before: [execution_backend {backend: local, kind: terminal, outcome: success} x1]
after: []
test_background_review_fork_terminal_calls_are_not_backend_usage is RED on base
(2 rows != 1), green here.
|
||
|
|
ab056dba47 |
fix(telemetry): feature_disabled skips migrations and ${VAR} templates, diffs outside caller locks (m22, m23, m24)
m22: migrations ran save_config -> record_config_saved inside surfaced
processes (`hermes config migrate`, console migrate, dashboard/TUI profile
create), so e.g. v45 appending `connections` to platform_toolsets read as a
user re-enabling a toolset. _persist_migration (the single migration write
path) now wraps save_config in hermes_applied_write(), a ContextVar the hook
honours.
m23: the old side is the raw file (`memory_enabled: ${MEM_ON}`) and the new
side the env-expanded config, so any unrelated save reported `disabled` once a
day forever. A setting whose value on either side is an unexpanded `${...}`
string is now skipped (its value is unknown to the diff).
m24: dashboard handlers call save_config inside config_write_scope
(_CONFIG_MUTATION_LOCK), and the diff (tools_config._get_platform_tools per
platform) + record ran there. record_config_saved now runs only the cheap
gate inline (surface, enabled, surface/home resolution) and hands deep copies
to a non-daemon thread bound to the owning profile (non-daemon so a CLI
command exiting right after its write still records).
Test contract change: test_record_config_saved_needs_a_surface_and_collection
and test_real_config_writes_report_the_move_away_from_default now join the
recording thread before asserting (the record is asynchronous by design).
Probes (review/signals/mine):
p2_migration.py config|dashboard
before: re_enabled toolset connections row; after: <no db>
p4_envref_configset.py
before: memory.memory_enabled disabled after an unrelated save;
after: no row (config set curator.enabled false still records)
p3_lock_v2.py (times config_transitions, try-acquires the web lock)
before (p3_lock.py): record_config_saved 0.047s/0.004s, lock held=True
after: diff 0.003-0.008s, lock held=False
Tests test_migrations_and_env_templates_are_not_user_disables and
test_diff_and_record_run_off_the_callers_lock_in_the_owning_profile are RED on
base.
|
||
|
|
85c317b0a9 |
fix(telemetry): feature_adoption ignores Hermes-internal curator/skill work; age read never raises (M15, m21)
- `curator` latched on the default-on scheduled pass (outcome=success, all
buckets 0), so it meant "idled through one curator interval". It now needs
trigger=manual (the user ran it).
- `skills_created` and the v4 milestone `first_skill_created` latched when
the background-review fork created a skill (provenance=agent_created).
Both now require provenance != agent_created (foreground creates are
stamped "learn" -> not agent_created, see tools/skill_usage.record_created).
- days_since_install_bucket caught only sqlite3.Error; a non-numeric
sessions.started_at raised ValueError into the subscriber on every event,
re-opened state.db each time and never latched. It now also catches
OSError/TypeError/ValueError -> `unknown`.
Probes:
probe_internal_adoption.py
before: feature_adoption curator + skills_created, milestone first_skill_created
after: no feature_adoption rows, no milestone
probe_statedb_errors.py textstamp
before: ValueError warnings, 0 rows, settled 0 of 1
after: 1 state.db read, feature_adoption {unknown, projects}, settled 1 of 1
Tests test_hermes_internal_work_never_latches_adoption and
test_unreadable_first_session_reads_unknown are RED on base.
|
||
|
|
3a0eaff6b2 |
fix(telemetry): tool_unavailable skips tool_search-deferred built-ins and cron (M14, m20)
tool_unavailable_fields compared against agent.valid_tool_names, the
model-visible array after tool-search assembly. With the default
tools.tool_search.enabled=auto, 14 enabled built-ins (session_search,
todo_list, process_manage, ...) live behind the tool_call bridge, so a direct
call to one was reported as "disabled in this session". Cron agents had
clarify stripped on purpose by the scheduler and still counted.
- Names in the session's tool-search scoped set
(agent.tool_executor._tool_search_scoped_names: the deferrable names its
enabled/disabled toolsets allow, cached on the agent) are enabled, not
unavailable. A session that turned the toolset off still reports it.
- platform == "cron" returns None, like delegated children and background
review.
Probes:
probe_deferred_tools.py (real AIAgent, default config)
before: rows for session_search and todo_list; after: no rows
probe_cron_narrowed.py
before: tool_name=clarify row; after: no rows
Test test_deferred_builtins_and_cron_narrowing_are_not_unavailable is RED on
base.
|
||
|
|
23b5f37d82 |
fix(telemetry): an opted-out start purges pending setup and exit markers (m19)
begin_process's collection-off branch purged only parked update receipts.
A provider_setup marker (pid, start_time, surface, provider) or a process
exit marker written while collection was on stayed on disk through the
opt-out and was reported (`abandoned` / `killed`) once collection came back
on, in a period the user had not consented to.
The same branch now removes process_markers/ and provider_setup_markers/.
Probe p2b_off_then_on.sh:
before: marker survives two OFF starts; the ON start emits
`abandoned openrouter`
after: markers {} after the first OFF start; the ON start emits only its
own started/completed rows
Test test_opted_out_start_purges_markers_left_while_opted_in is RED on base.
|
||
|
|
41ce3104a7 |
fix(telemetry): provider_setup counts only a new provider credential or endpoint (M13, m18)
Every PUT /api/env write of a provider-shaped variable recorded
started+completed: an empty value (a clear), a same-value re-save,
GITHUB_TOKEN/GH_TOKEN (-> copilot), HF_TOKEN, and keys the TTS/STT/image
panels also save (GEMINI_API_KEY, XAI_API_KEY, DEEPINFRA_API_KEY). TUI
model.save_key counted re-saves; editing a custom endpoint counted as a new
setup.
- record_api_key_saved takes the value and the stored value read before the
save: only a non-empty, changed value counts.
- provider_for_api_key_env returns None for ecosystem tokens (GITHUB_TOKEN,
GH_TOKEN, HF_TOKEN) and any key a TOOL_CATEGORIES panel asks for: a bare
key save cannot tell a provider connection from a tool setting (the PUT
body has no provider flag; adding one is Desktop work, see NOT_COVERED).
Provider pickers (CLI, TUI model.save_key with its explicit slug) still
count those providers.
- model.save_key skips a same-key re-save.
- upsert_custom_endpoint records only when the endpoint did not exist.
Probe p4_web_double.py (real route functions):
before: openrouter started x3/completed x3 (save, re-save, clear),
copilot x1 from GITHUB_TOKEN, custom x3 (create + 2 edits)
after: openrouter x1, no copilot row, custom x1
Tests test_web_forms_count_only_a_new_provider_key_or_endpoint and the
extended test_api_key_env_maps_to_provider_only are RED on base.
|
||
|
|
946ae62a89 |
fix(telemetry): provider_setup Esc is cancelled and Back continues one flow (M12)
cli_provider_setup treated the setup menus' BaseException control flow as
errors: Esc (_SetupCancelled) recorded failed/other, and Back (_SetupGoBack)
recorded failed/other and then a second `started` when the wizard replayed
the provider picker.
- setup_failure_class maps _SetupCancelled to `cancelled`.
- Back leaves the flow open on the thread-local; the replayed picker
re-entering the same surface+provider continues it (no new `started`).
Picking another provider, or leaving the entry point
(provider_setup_surface exit other than a Back), ends it failed/cancelled.
Probe p5_cli_nav.py (real cmd_model):
before: back_then_pick started x2, failed/other x1, completed x1;
escape started x1, failed/other x1
after: back_then_pick started x1, completed x1;
escape started x1, failed/cancelled x1;
back_then_esc_menu started x1, failed/cancelled x1;
back_then_other openrouter started+cancelled, anthropic started+completed
Test test_cli_setup_navigation_esc_cancels_and_back_resumes_one_flow is RED
on base (failed/other, extra started).
|
||
|
|
bfa3b84c9f |
fix(telemetry): Desktop closes action/setting ids locally and records first-run steps after a yes
m13: the renderer stored and sent raw plugin palette/keybinding ids, rage-click
targets and config paths below user-named containers (providers.<name>.api_key);
the backend collapsed them, but only after truncating the action list to 200.
- recordAction collapses anything outside the built-in ids (KEYBIND_ACTIONS +
KEYBIND_READONLY + the named buttons, minus numbered slots — exactly the
backend's DESKTOP_ACTION_IDS) to `other` before it is counted or used as a
rage-click target.
- recordSettingsSaved sends only keys the config schema publishes (the Settings
page passes its schema: DEFAULT_CONFIG leaves), so nothing below a user-keyed
container leaves or is kept.
- Backend DAILY_ACTION_ROWS_MAX = |DESKTOP_ACTION_IDS| x |vias| (296): every
distinct collapsed (action, via) fits, so none is cut before collapsing.
m14: first-run steps happen before the consent question and collection defaults
off, so every step but consent/first_message was dropped. Chose the in-memory
buffer (smaller than removing the steps from contract, schema and TS): while the
answer is unknown/undecided, recordOnboarding/closeOnboardingStep calls are held
in memory (max 100, never persisted or sent) and replayed on the switch turning
on; a decided "off", a profile switch or quitting discards them.
setDesktopMetricsGate takes `decided` (hook passes consent.decided).
Tests (desktop-metrics.test.ts): recordSettingsSaved callers pass a schema
(contract change). New invariants, RED on base: plugin action id / rage target
stored and sent as `other`; a providers.<name> key is not sent; pre-consent
onboarding held, sent on yes (with closes applied), dropped after a decided no.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, raw ids):
before: localStorage actions {"acme-plugin:/home/alice/clients/bigco|palette":1, ...|click:3};
wire rage_click target "acme-plugin:/home/alice/clients/bigco", setting
"providers.acme-corp-internal.api_key"
after: localStorage actions {"other|palette":1,"other|click":3}; wire
[rage_click target "other"] only
|
||
|
|
a0508147c1 |
fix(telemetry): Desktop metric state is per profile and shared by its windows
M9 + the renderer half of M10, and m15.
M9: the renderer kept one install-wide localStorage record and bound the new
gateway before reading its consent, so a day accumulated under profile A was
flushed to B before B's switch was read (or dropped by B's off), and A's area
latch suppressed B's first use.
- The record is keyed per (connection, profile): `hermes.desktop.metrics.v1:
<FNV hash>` (no profile name in the key). bindDesktopMetrics takes the
profile scope; a different scope forgets in-memory state and nulls the gate,
so nothing is kept or sent until that profile's switch is read.
- The hook pins its requester to the focused (connection, profile) with
requestGatewayForAgent instead of the session-tile router, so the RPCs land
in the profile whose switch gates them. main receives the same scope hash.
M10 (renderer): each window cached the record in module memory and the last
debounced persist() won, so peer windows overwrote each other's counts. Every
change now reads the record fresh and writes it back (no debounce, no cache;
persistDesktopMetricsNow is gone). withState calls are no longer nested
(noteMessageSent / abandoned-step reporting read first, record after), since a
nested write would be overwritten by the outer one. A daily flush is pinned to
the scope and requester it started with. Both windows may still send the same
finished day; the backend's durable latch (earlier commit) records it once.
m15: consent was re-read only on attach or profile switch. The hook re-reads
it on window focus; Settings › Privacy applies its save to the gate when it is
scoped to the focused profile by name, not only when unscoped.
Tests (desktop-metrics.test.ts): `stored()` now finds the per-profile key and
setEnabled is expected with the scope argument (contract change). New
invariants, RED on base: profile switch A→B→A keeps/sends nothing of A under B
and A still reports its day; two module instances (windows) of one profile sum
into one day.
Probe (review/desktop/jsrun/.../desktop-metrics.probe.test.ts, adapted to scopes):
before: B received before its consent read: [["shared_metrics.desktop_daily",{...view.toggleSidebar 4...}]]
areas latch A: [terminal_pane] B: []
two windows stored: {"composer.send|shortcut":4} (W2's session.new lost), W2 daily sends
composer.send 3 + session.new 2
after: B received before its consent read: []; areas latch A: [terminal_pane] B: [terminal_pane]
stored: {"composer.send|shortcut":4,"session.new|shortcut":2}; W1 and W2 send the same
aggregate (the backend latch keeps one)
|
||
|
|
aa83040ca4 |
fix(telemetry): renderer-crash consent and pending crashes are per window and profile
M11. RendererCrashRecorder held one install-wide `enabled` flag that every window
overwrote with its own profile's switch, so an opted-out window's crash was
written to disk and reported into another profile, and an opted-out window
wiped the opted-in profile's pending crash.
- Consent is kept per webContents id (the IPC sender), with a hash of the
window's focused profile key; set-enabled now carries that key.
- A crash is recorded only when the crashed window's own profile collects;
each pending entry is tagged with the profile hash (file v2: {entries:[{profile,
reason}]}; an untagged v1 file is ignored — whose it was is unknown).
- take() returns only the calling window's profile's reasons; one claim per
profile (the same window may re-claim after its renderer reloaded); ack drops
only the claimed count. Turning a profile off drops only that profile's entries.
- main.ts passes the window's webContents id to recordRendererGone.
Tests: renderer-crash-metrics.test.ts moves to the per-window API (contract
change: every call names its window, set-enabled names the profile); one new
invariant test (opted-out window records nothing, another profile neither
drains nor purges) — the base API cannot express it, so RED is shown by the
reviewer's probe instead.
Probe (review/desktop/jsrun/.../renderer-crash-metrics.probe.test.ts, adapted to
window ids):
before: after opted-out window crash, file exists: true {"v":1,"reasons":["crash"]}
take by window 2: {"reasons":["crash"]}; A crash record after B-window off: false
after: file exists: false; take by window 2: null; A crash record after B-window off: true
|
||
|
|
4bb58d3d60 |
fix(telemetry): switching Desktop collection off drops the backend onboarding latches
m16. shared_metrics.set(enabled=false) purged the renderer's onboarding latch but left telemetry/shared_metrics/desktop_onboarding/<step>.<event> in the profile, so a content-free per-profile record outlived the opt-out and the two sides disagreed on re-opt-in. The Desktop's consent RPC now removes that directory when it turns collection off (shared_metrics_desktop. purge_onboarding_latches). Opting out from the CLI (config set / setup wizard) leaves them until the Desktop's switch is next turned off — see NOT_COVERED. Test: test_turning_collection_off_drops_the_onboarding_latches — RED on base (latch file survives), green now. Probe (review/desktop/probe_optout_set.py): before: surviving desktop latches: [.../desktop_onboarding, .../intro.reached, ...] after: surviving desktop latches: [] (only metrics.sqlite3 / outbox remain) |
||
|
|
1a07b7259a |
fix(telemetry): Desktop daily report and feature use latch durably per profile, in the usage day's period
M10 + m17 (+ the backend half of M9). The backend's once-per-day latches for
hermes.desktop.{action_use,mode_use} (the finished-day report) and
hermes.desktop.feature_use lived in process memory, so a second window's report,
a backend restart or a pooled backend re-counted the same day/area. The daily
rows also went through the relay's current-day marks, so a day flushed a week
late was counted in the flush day's period.
- record_desktop_daily now writes through SharedMetricsStore.update_rollup_state:
one write transaction checks a `desktop_daily_reported` state row (the usage
days already reported, bounded to the last 8 days) and records the rows with
period_start = the usage day. Days outside [today-8, today] settle unrecorded.
- record_desktop_feature_use uses store.record_counter_once_per_day (the same
durable per-UTC-day-per-dimension latch feature_disabled uses).
- With collection off, desktop_daily answers recorded:false instead of true, so a
renderer that has not learnt the switch yet never treats that as "settled" and
drops the day (the client's own gate purges it when it reads off).
- Neither metric has feature/milestone side effects in the subscriber, so
writing to the store directly is equivalent apart from the latch and period.
Tests: the fixture mocked record_process_marks_saved; the daily/feature-use
assertions now read the profile's store (their contract changed: rows no longer
go through the relay). The disabled-profile test now expects recorded:false.
The store-refusal test raises from update_rollup_state instead of mocking the
relay's saved count. New invariant test (two windows / a restarted backend, all
in-process state reset between reports): RED on base, green now.
Probe (review/desktop/probe_latch.py, run twice on one home = backend restart):
before: run 2 → action_use composer.send 2, feature_use terminal_pane 2,
mode_use 2, all in period 2026-09-28 for usage day 2026-09-20
after: run 2 → action_use 1, feature_use 1, mode_use 1; action_use/mode_use in
period 2026-09-20. Opted-out home: daily -> {'recorded': False}.
|
||
|
|
8dfdc46067 |
fix(telemetry): relay-stamped turns are gateway messages; a multi-platform connector's unknown chat stays relay (M5 part, m9)
What:
- M5: `relay` is a gateway surface alias, so an inbound the connector could not stamp is
entrypoint gateway_message / surface gateway (it counts as engagement) instead of other/none;
`relay` joins the task/session gateway platform set (it was collapsing to `plugin`).
- m9: RelayAdapter._metrics_platform labels a chat by its inbound platform, and falls back to the
connector's primary platform only when the socket fronts exactly one identity; _source_platform
uses it (egress routing keeps _chat_platform unchanged).
Why: the same turn was `other/none` in task/session while delivery/latency said `discord`; an
unknown chat on a multi-platform connector was misattributed to the primary platform.
Tests: the existing unresolvable-chat stub now implements _metrics_platform (renamed hook). New:
task_start_fields({"platform": "relay"}) -> gateway_message/gateway/relay; unstamped chat on a
two-identity connector -> delivery platform `relay`. Both RED on the previous commit
(other/other/none; discord).
Probe (review/engagement/relay/probe_d_coverage.py), before: source.platform='relay' ->
{'entrypoint': 'other', 'execution_surface': 'other', 'platform': 'none'}, engagement surface None.
After: the contract test above.
Not done: resolving the turn-start task platform to the connector's platform (needs the adapter at
agent construction) and M6 (delivery row profile under multiplex); see NOT_COVERED.md.
|
||
|
|
4c7085d284 |
fix(telemetry): engagement and switch_after count only a person's turns (M1, M3, M7, M8, m6)
What:
- M1: the engagement turn mark (primary model) needs a user turn (pre_llm_call) on an interactive
surface or a gateway message; API-server, python embedding and curator turns no longer create
engagement.day rows or outvote the person's model. close_day emits nothing for a day with no
engaged surface (the root profile's host row with an active-profile count stays).
- M3 (owner ruling): unattended cron runs are dropped from engagement entirely. surface_day always
carries an active-minutes bucket, so it cannot express cron presence without minutes; hermes.cron.run
already counts runs. The `cron` enum value stays in the contract for already-stored rows.
- M7: background review forks (no pre_llm_call) no longer advance model_switch_after's turn run.
- M8: the run observes the route each turn was sent on (its first request), so a provider failover
neither restarts the run nor names the fallback as the model left.
- m6: documented that the root host row (surfaces_used_count 0) is excluded from days-active counts.
Why: Hermes-internal and unattended activity is not user engagement; switch_after measures how long
a person stayed on the model they chose.
Tests (RED on the previous commit): API/cron-only day -> no day row, mixed day -> primary = person's
model; failover + review fork -> switch_after names the selected model with 2_to_3.
Probe (review/engagement/probe_internal.py, switch/probe_switch.py -k "p3c or p7"), before:
API-only / curator-only day -> engagement.day {"primary_model": "openai/gpt-5", "surfaces_used_count": "0"}
cron all day -> day {"active_minutes_bucket": "gte_6h", "active_profile_count_bucket": "1"} + surface_day cron
2 CLI + 5 API turns -> primary openai/gpt-5
P3c (6 turns on A, 7th failed over to B) -> {"model": "gpt-5", "turns_before_switch_bucket": "1"}
P7 review fork -> 4_to_10 (routed=False) / gpt-5-mini "1" (routed=True)
after: the equivalent invariant tests pass (no day row, primary = PUBLIC, selected model 2_to_3).
|
||
|
|
b3d4865328 |
fix(telemetry): link session segments only at the compression hand-off (B2, M4, m2, m8, m11)
What: the metrics lineage no longer keys off the ambient Portal conversation id (it walks every parent_session_id: gateway reset/idle expiry, /new, /branch). Compression calls relay_shared_metrics.rotate_segment(old, new) from the committed rotation (_notify_context_engine_compression_complete) and from a stale agent adopting the live tip (_adopt_live_compression_child). The hand-off registers the new id in the lineage, moves the route run and the spent-turn history into it, and closes the old segment now (or when its in-flight compression turn ends). A rotated-to id closed before serving a turn, or still open at shutdown, closes too; a 512-lineage backstop flushes conversations a surface never closed. - m2: /undo or /retry on the new id reaches the turns earlier segments spent (tokens_bucket known). - m11: /model right after a rotation finds the conversation's turn run. - m8: background review forks (same session id, no pre_llm_call) add no turns/calls/messages. Why: one conversation = one session / context_peak / tool_overhead / tool_enabled_unused row, on time; N gateway conversations collapsed into one row emitted only at process exit, and rotated segments stayed in memory until exit. Tests: two existing tests simulated rotation by sharing a conversation context; their contract changes to the explicit hand-off (test_a_compressed_conversation..., test_context_peak_is_one_row...). New: reset/branch chain after a compression + tip-first close order, the compression seam, switch/undo right after a rotation, review fork volume. RED on base (rotate_segment missing / switch_after [] / turn bucket '2' for one user turn). Probe (review/engagement/lineage probe_lineage.py -k c2, probe_branch.py; fix-lineage/probe_handoff.py): before: c2 3 conversations -> session.count 0 until shutdown, then 1; branch -> 1; 50 compressed conversations -> 0 rows, 50 lineages/sessions live. after: c2 -> 3 rows before shutdown (turns 2,1,1), lineages=0; branch -> 2; 50 compressed -> 50 rows, 0 lineages, no sessions; tip-first close with the old turn running -> 0 then 1; shutdown with an open tip -> 1 per conversation. |
||
|
|
ac372200b1 |
fix(telemetry): network-address model ids collapse; subscriber re-runs the catalog (m12)
What:
- shared_metrics_catalog.model_metric_name also collapses a model id that is
a network address: a loopback/IPv4 host or localhost at the start, or a
leading host:port with a 2-5 digit port (so Bedrock's "...-v1:0" and
OpenRouter's ":free" / Ollama's ":120b" tags stay readable). The existing
URL/path/weight-file rules move into the same compiled pattern.
- shared_metrics_contract._bounded_dimensions (every mark -> counter
projection) and model_token_counters re-run provider_metric_name /
model_metric_name on the mark's provider/model identifier fields and drop
the row when either would rewrite it ("none" = unset provider passes).
Record-time only: counter_dimensions_are_valid (also used at packaging) is
unchanged, so catalog drift can never raise out of _package_metric and
block a package of rows already stored.
Why: `127.0.0.1:8080/x` on a public provider passed verbatim through every
route metric, and the subscriber validated identifier shape only, so a mark
that skipped the producer's catalog pass stored custom:acme, ollama/<model>,
unknown vendors and even http:// URLs.
Probe (fix-catalog/probe_b1_m12.py):
before: route(openrouter, 127.0.0.1:8080/x) -> model '127.0.0.1:8080/x'
(also localhost/qwen, 10.0.0.5/x, gpu-box.lan:8000/qwen verbatim)
subscriber injected switch_after {127.0.0.1:8080/x, openrouter} -> stored
subscriber injected switch_after {acme-secret, acmecorp-internal} -> stored
after: all of the above -> model 'custom'; both injected marks -> None;
bedrock anthropic.claude-3-5-sonnet-20241022-v2:0, openai/gpt-4o:free,
gpt-oss:120b and the legit anthropic mark unchanged.
Test: test_network_address_model_ids_and_raw_injected_marks_never_reach_counters
(RED with the catalog reverted: model_route kept 127.0.0.1:8080/x; RED with the
contract reverted: the injected host:port mark was stored; GREEN with both).
|
||
|
|
802e6797a0 |
fix(telemetry): provider allowlist comes from shipped providers only (B1)
What: shared_metrics_catalog.provider_names() no longer reads the live
PROVIDER_REGISTRY / models._KNOWN_PROVIDER_NAMES. It is now the built-in
auth rows (BUILTIN_PROVIDER_IDS), HERMES_OVERLAYS, the three static alias
tables, openrouter/custom, and the names+aliases of profiles registered
from the in-tree plugins/model-providers dir (providers._SOURCES ==
"bundled", process-wide layer, independent of the bound profile home).
Why: auth_plugin_providers mirrors every $HERMES_HOME (and pip) provider
plugin, plus its aliases, into PROVIDER_REGISTRY and the picker labels, so a
user-chosen provider name was treated as shipped and its model id was not
collapsed either. Every consumer goes through provider_metric_name, so the
one fix covers model_route and every per-model v4/v5 metric (tokens,
context_peak, friction, tool_quality, loop_guard, tool_recovery,
model_reply_issue, task_cost, wasted_tokens, cache_break, switch_after,
tool_unavailable, engagement.day), provider_setup (all surfaces, incl. the
PUT /api/env key path), the on-disk pending marker, setup.completed,
model_switch/fallback and install.snapshot main_provider. Pip-installed
provider plugins now read custom too (third-party, not shipped); doc says so.
Probe (fix-catalog/probe_b1_m12.py, user plugin acmecorp-internal alias
acme-llm):
before: provider_metric_name('acmecorp-internal') -> 'acmecorp-internal'
route: {'model': 'acme-secret-model-v2', 'provider': 'acmecorp-internal'}
marker on disk: {..., "provider": "acmecorp-internal"}
after: provider_metric_name('acmecorp-internal') -> 'custom' ('acme-llm' too)
route: {'model': 'custom', 'provider': 'custom'}; marker "provider": "custom"
install main_provider custom; switch/fallback/setup custom
shipped plugin deepinfra: {'model': 'meta-llama/...', 'provider': 'deepinfra'}
Test: test_user_provider_plugin_name_and_model_never_leave (RED on base:
marker carried "acme-llm"; GREEN after).
|
||
|
|
c13a4abc8d |
feat(telemetry): v5 signals — tool_unavailable, provider_setup, feature_adoption, feature_disabled
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality. - hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface (cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow; a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern). - hermes.feature_adoption.count: once per feature per install at first real use, derived from existing counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the owning profile's install age. - hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day. Surface comes from the user entry point; setup/migrations record nothing. |
||
|
|
6c143c238f |
feat(metrics): hermes.desktop.* contract, schema and gateway RPCs
Desktop love/hate telemetry needs a backend half: six counters (feature_use, action_use, mode_use, friction, dislike, onboarding) with closed dimension sets, the v3 JSON schema entries, the public doc rows, and profile-scoped gateway RPCs the app reports through. Every value the app sends is collapsed server-side onto the closed set (unknown -> other); feature use latches per area per UTC day, friction and dislike are capped per day, the daily report (action counts + mode time + bot count) latches per day and answers recorded=false on a store failure so the app keeps the day for a retry. For setting changes the app sends only the config key: the backend reads the saved value and compares it to DEFAULT_CONFIG itself, so values never cross the wire; keys outside DEFAULT_CONFIG read other. All helpers check enabled() before doing any work. |
||
|
|
928daa63af |
feat(telemetry): v5 engagement rollup, turns-before-switch, per-conversation sessions, relay source platform
Engagement (local daily rollup, emitted once the UTC day closes, exactly once per profile):
- hermes.engagement.surface_day.count {surface, active_minutes_bucket} per surface used
(cli/tui/desktop/gateway/acp/cron); active time = inter-interaction gaps capped at 5 min.
- hermes.engagement.day.count {active_minutes_bucket, surfaces_used_count, primary_provider,
primary_model, active_profile_count_bucket}; primary model = most attended turns that day.
Active profiles are folded into one host accumulator in the root profile's store (opaque local
hashes); only the root profile's row reports the count, other profiles report 0.
- Days-active-per-week / return-by-model are derived server-side; no weekly state, no new id.
Model satisfaction:
- hermes.model_switch_after.count {provider, model, turns_before_switch_bucket}: turns the model
being left served in the conversation, from every /model site (CLI, TUI, ACP, gateway).
Decisions:
- hermes.session.count counts once per conversation across compression rotation (merged over the
v4 conversation lineage), and gains message/model-call/tool-call count buckets with a longer
tail (251_to_1000, gte_1001); rows recorded before the upgrade still package.
- Relay-delivered hermes.gateway.reply_latency and hermes.platform.delivery report the originating
platform instead of `relay`.
|
||
|
|
607dff390a |
feat(telemetry): v5 efficiency metrics — turn cost, wasted tokens, tool output truncation, tool overhead, prompt-cache breaks
Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):
- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
A Hermes-known break suppresses the miss it causes. No prompt hashing.
Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.
|
||
|
|
8d8836ddb1 |
feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through record_process_mark (disabled config => zero rows, zero files): - hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy matcher reports the strategy that landed (or no_match / ambiguous) into a context-local probe opened only around the patch/write_file handlers, so we learn which strategies earn their keep without counting internal callers. - hermes.loop_guard.count: provider, model, signal, detector. Hooked where the guardrail already warns/blocks/halts and at turn end for the iteration budget; latched once per turn per (signal, detector). - hermes.tool_recovery.count: provider, model, tool, next_tool, next_outcome. One row per failed tool call, resolved against the model's next round (same tool first) or the turn end (no_tool_call / gave_up). - hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is a table lookup on the first program word, never the text. Hermes' own deadline/interrupt now set hermes_timed_out / hermes_interrupted on the backend result, so a command's own `exit 124` reads nonzero, not timeout. Disjoint dims from hermes.execution_backend.count (served vs not). - hermes.model_reply_issue.count: provider, model, issue. Refusal and truncation only from the structured finish reason; empty/reasoning_only from the normalized message; issue=none per response as the denominator (model_route counts attempts and files empty replies as failures). Background review / curator loops and _host_local / unmetered terminal calls never count. Helpers live in shared_metrics_harness.py; call sites are one or two lines each. |
||
|
|
88eb76f02e |
fix(metrics): startup latency counts once per client launch, never after an in-place exec
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat call added a row. The TUI and Desktop now send an opaque per-launch id (TUI: one per process; Desktop: minted when the once-per-launch main-process claim succeeds) and the backend claims (surface, launch_id) once per process: a reconnect re-sending the same launch counts once, a new Desktop launch against a long-lived backend still counts. Legacy clients without an id count once per backend process and surface. The id is only a local latch key (bounded set, length-capped), never recorded. Only a usable measurement spends the claim. Contract + apps/shared regenerated. - m10: relaunch() execs in place, keeping the PID, so `sessions browse` -> resume counted the picker time as CLI startup. relaunch() stamps its PID in HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it matches, while children (other PIDs) inheriting the env still count. - m8 (psutil half): record_process_ready checks the opt-in gate before reading the process start time (psutil) or starting a thread. The claim is taken first so a later opt-in never records a stale "startup" mid-session. Test changes to existing assertions: the TUI/Desktop param assertions now expect launch_id (a new wire param); the RPC surface test sends distinct launch ids so its three calls stay three launches under the new latch. |
||
|
|
f7b5ea6f7a |
fix(telemetry): loop metrics count only the user's real backend work and curator passes
Review findings M7, M8, m6, m7 and the browser half of m8 on the v4 loop metrics. - M7: TUI/Desktop @-path completion on a non-local backend lists the directory through terminal_tool, so every keystroke added a hermes.execution_backend.count row as if the user ran a command. Hermes-owned calls now run inside shared_metrics_loop.unmetered_backend_calls() (a contextvar checked in record_execution_backend) instead of a flag threaded through terminal_tool. It was the only non-_host_local internal terminal_tool caller (bot DM runners already use _host_local; the prompt-builder probe calls env.execute directly). - M8: a scheduled curator tick that finds the run claim held recorded scheduled/skipped every tick (5 ticks -> 5 rows for one pass). The holder's own pass is the one counted; the held branch now records nothing. Real dry runs still report skipped. - m6: a foreground timeout returns exit 124 with partial output and no error field, so it was recorded as success. exit 124 now records outcome=failed, error_class=timeout. - m7: vercel_sandbox (and managed_modal) are built-in backends in tools/terminal_tool_config._BUILTIN_BACKENDS but collapsed to other. Added to TERMINAL_BACKENDS and both schema enums; a test pins every built-in backend to its own bucket. - m8 (browser): record_browser_call resolved the browser backend on every call even with collection off. The resolver is now passed lazily and only runs inside the gated builder. The existing test asserting a held claim records scheduled/skipped encoded M8 itself and is replaced by one asserting held ticks record nothing. |
||
|
|
287564fedd |
fix(telemetry): update runs counted once, parked only while opted in, recovered files kept until saved
Three defects in how hermes.update.run/stage and hermes.process.exit rows are recovered from disk: - Opt-in (M1): the pre-pull park path only checked that the telemetry dir existed, so it wrote the WHOLE receipt (argv, step detail) while the user was opted out, and a later opt-in counted that run. It now reads consent through the already-loaded hermes_cli.config (this interpreter predates the checkout swap, so nothing may be imported), parks only when collection is on, parks only the fields update_receipt_fields reads (no argv, no step text), and purges pending_updates when collection is found off. begin_process purges them too on a start with collection off. - Double count (M2): the completion child records the receipt and the parent's boundary finalize (run_completion returned receipt=None) parks the same update_id again, so one run became (failed, deps) + (success, none). record_update_receipt now claims the update_id with an O_EXCL latch under shared_metrics/recorded_updates (last 64 kept); the second finalize records nothing. - Loss (m3): dead-process markers and parked receipts were unlinked before their row reached the store; "database is locked" under concurrent starts lost them (17/20 rows in the reviewer's race). The claim-by-rename stays, but the file is deleted only after record_process_marks_saved confirms the rows settled: rows carry a random commit ticket in event metadata (allowlisted in the contract), the subscriber tallies tickets whose rows persisted, and the reporter flushes the Relay subscribers and compares. Otherwise the file is renamed back for the next start. A parked receipt whose rows partly landed is not retried (that would double count). Also: _epoch returns None for a timestamp float() cannot represent (10**400 raised OverflowError). |
||
|
|
2c9b79a5a3 |
fix(telemetry): one context_peak row per conversation across compression
Rotating compaction continues a conversation under a new session id, and each metrics session emitted its own context_peak at close, so one conversation produced a row per segment (the fuller one plus a low-fill one), skewing the per-model distribution. The agent turn already publishes the lineage root (the Portal conversation id); segments sharing it join one lineage, contribute their peak when they close, and the merged peak (fullest fill, limit hit anywhere) is emitted once when the last open segment closes, whatever the close order. Delegated children share their parent's root but stay their own conversation. Review finding m2. |
||
|
|
093c10a2ee |
fix(telemetry): local-server aliases and unknown providers never ship model ids
provider_names() accepts ollama/local/vllm/llamacpp/llama.cpp as published names (they are aliases of the generic `custom` provider), so every provider/model metric kept the raw config.yaml model id under them, and a missing provider (`unknown`) kept it too. Map the aliases Hermes itself routes to `custom` (derived from the providers/auth/models alias tables) to `custom`, and report the model as `custom` whenever the provider is unknown: nothing proves such an id is public. Shipped providers with public ids are unchanged. Covers model_route, model_tokens (primary and auxiliary), fallback, model_switch, setup.completed, install snapshot main_provider, friction, tool_quality and context_peak, which all go through this pair. Gateway /model also labelled the configured model with switch_model's `openrouter` default when config.yaml names no provider; it now reports the provider actually configured/overridden (None -> `unknown`). And the switch row + switch_away friction bind the routed profile's home on a multiplexed runner (slash dispatch installs no profile scope), as the /retry and /undo friction already do. Review findings B2, M3. |
||
|
|
fd14f50b5d |
feat(metrics): hermes.update.run/stage derived from the final update receipt
The update pipeline already writes one machine-readable receipt per run, so the run/stage counters are derived from it in the process that finalizes it instead of instrumenting stages for metrics. The receipt gains minimal stage END marks (plan, snapshot, apply[mode], deps, build, restart; verify is inferred from the fleet matrix), an initiator fact (desktop when a live orchestrator claim names another pid) and the pre-update commit date, which is all the derivation needs for outcome, failed stage, duration and from-version-age buckets. A pre-pull interpreter must never import pulled code: when it is the finalizer (same pid, checkout sha moved) it parks the receipt (stdlib only) and the next Hermes start records it. Collection off => nothing imported beyond a config read. The completion-process test stub gains the new record_stage API the real module now exposes (no assertion changed). |
||
|
|
2b3e1c3995 |
feat(metrics): startup latency and install version-lag/hardware fields
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.
hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.
The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
|
||
|
|
2efa4f3caa |
feat(telemetry): per-model tool-call quality, friction and context-peak counters
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.
hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.
hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.
hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
|
||
|
|
f4075a20d6 |
feat(telemetry): gateway platform health, delivery, first-reply latency and cron run metrics
The gateway and cron ticker were blind spots in shared metrics: we could not tell which messaging platforms fail to connect or drop, how often replies fail to reach users, how long users wait for a first reply, or whether scheduled jobs run, fail, get skipped or get missed while Hermes was down. New opt-in counters (all through the existing enabled() / record_process_mark gate, recorded fire-and-forget on one background worker that runs in a copy of the caller's context so the owning profile gets the row): - hermes.platform.health: platform, event (connect_ok/connect_failed/ reconnect/disconnect), error_class. One seam in the runner (_connect_adapter_with_timeout, used by cold start, multiplex secondaries and the reconnect watcher) plus the fatal-error handler; classified from exception types, HTTP statuses and Hermes's own fatal codes, never text. - hermes.platform.delivery: platform, outcome, failure_class. One row per logical BasePlatformAdapter._send_with_retry call (retries and plain-text fallback included), from SendResult's typed fields or the exception type. - hermes.gateway.reply_latency: platform, first_response_bucket. Clock starts when _handle_message accepts a non-internal turn; stops on the stream consumer's first delivered text or send_final_ledgered. Busy acks and command replies never stop it. - hermes.cron.run: outcome (success/failed/missed/skipped), delivery_kind, duration_bucket. Recorded at the write-once terminal execution row (finish_execution) and where the due scan drops an occurrence (catch-up disabled, expired one-shot). Job names, prompts and targets never leave. Platform names: platforms Hermes ships under plugins/platforms/ now report by name, and a plugin-catalog platform reports its catalog entry name only when the installer-owned .install-metadata.json record proves a catalog install of the plugin dir that defines the registered adapter factory (a URL install cannot claim a catalog name via its tree or manifest). Everything else stays "plugin". GATEWAY_PLATFORMS becomes catalog-backed (_CatalogValues), and the v3 schema's platform fields accept published catalog identifiers. |
||
|
|
2e56f40c16 |
feat(telemetry): learning-loop, delegation and execution-backend metric contract
Adds four opt-in shared-metrics counters so we can see whether the learning loop actually runs and pays off, how wide delegate_task fan-outs go, and which sandboxes carry real work: - hermes.memory.op.count: op, provider (builtin / bundled plugin / plugin), origin (foreground / background_review), outcome (success / failed / rejected) - hermes.curator.run.count: trigger, outcome, archived/created/merged/patched buckets - hermes.delegation.run.count: subagent_count_bucket, depth, mode, outcome - hermes.execution_backend.count: kind (terminal/browser/code), backend, outcome, error_class Every dimension is a closed enum or COUNT_BUCKETS value; unknown providers and backends collapse to plugin / other. Builders live in the new shared_metrics_loop.py sibling (stdlib-only imports so tool hot paths can use it) and go through the existing enabled() gate. Off-turn producers (curator thread, async delegation units) record in the profile captured when the work started. Schema and the "what is collected" doc gain matching entries. |
||
|
|
9edbf627f0 |
fix(desktop): shared-metrics question is a composer offer strip, never a launch modal
The first-run consent dialog blocked the composer on every undecided profile, so a healthy install no longer opened straight to chat (caught by the Desktop first-run E2E). The question now lives in the composer status stack beside the free-tier strip: Send to Nous / Local only / No thanks / Details, no focus steal, one owner across split composers, one offer at a time. Details opens the full explainer on request; closing it decides nothing. The backend's `decided` stays the only latch. |
||
|
|
3d62ae2233 |
feat(telemetry): decision-data docs, consent copy, smoke coverage and install-age for new installs
- Docs: a Decision-data metrics section mapping each metric to the product question it answers; skill-load and snapshot paragraphs describe the new public-name and install-age fields. - setup consent copy names the new data classes (session length, token totals, command and catalog names, bucketed setup counts). - The real-CLI smoke now asserts the session summary, token sums, TTFT bucket and milestones in both the SQLite store and the exported, schema-validated package. - A profile with no sessions yet is a brand-new install (lt_1h), not unknown: the first task's milestone and the first snapshot fire before state.db has a session row. - Two invariant tests: session summaries + once-per-install milestones; token sums per model and auxiliary task with private task names collapsed. |