Files
hermes-agent/hermes_cli
teknium1 8d8836ddb1 feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):

- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
  matcher reports the strategy that landed (or no_match / ambiguous) into a
  context-local probe opened only around the patch/write_file handlers, so we
  learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
  the guardrail already warns/blocks/halts and at turn end for the iteration
  budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
  next_outcome. One row per failed tool call, resolved against the model's
  next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
  a table lookup on the first program word, never the text. Hermes' own
  deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
  backend result, so a command's own `exit 124` reads nonzero, not timeout.
  Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
  truncation only from the structured finish reason; empty/reasoning_only
  from the normalized message; issue=none per response as the denominator
  (model_route counts attempts and files empty replies as failures).

Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
2026-09-28 12:43:03 -07:00
..
…
…
…
…
…
…
…