Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):
- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
A Hermes-known break suppresses the miss it causes. No prompt hashing.
Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.