Files
hermes-agent/tests
teknium1 060cd7f9bb fix(state): trim the futile-holder FTS diagnostic to shape and fix the remedy text
Slim redo of the mechanism from #106410 on top of its pick (no wrappers, no
persisted "kind" enum, no process-local flag that dies with the process):

- Futility = the SAME holder PID set has blocked >= _FTS_HOLDER_FUTILE_ATTEMPTS
  (10) deferrals over >= _FTS_HOLDER_FUTILE_SECONDS (30 min); tracked as
  holders_since/holders_attempts in the persisted fts_rebuild_deferral record
  and reset whenever the holder set changes. The 3-deferral/60 s escalate
  window is the orphan-reap gate and stays as is.
- ONE escalated ERROR line names each holder pid + cmdline and the remedy that
  can actually be followed from inside a gateway session: stop ONLY the other
  holder; this process's own retry admits the rebuild within 60 s. The old
  "with the gateway stopped" advice was unrunnable from a gateway-hosted
  session (the gateway is the session) and is gone from both log and doctor.
- hermes doctor renders the futile record distinctly.
- retry_deferred_fts_recovery: a capped backoff earned by holder set X no
  longer applies once the live holder set differs from X, so stopping the
  other service is followed by a retry on the next tick, not up to an hour
  later (the issue's 16-min wait).
- Tests trimmed from 5 to 2 invariants (futile line + doctor entry after N
  same-holder deferrals; backoff reset when the holder set changes); the
  contributor's control tests for changing PIDs / orphan reap are covered by
  the existing test_repeated_deferrals_reap_inactive_orphan_then_rebuild.

The "canonical writes and LIKE search remain available" WARNING is kept
because it is true on origin/main: a stale open drops every FTS trigger, so
the messages INSERT succeeds (probed live with a real state.db + a second
process holding it). Writes fail only when a peer re-publishes triggers over
the corrupt index — a separate class, not this diagnostic.

Refs #106393
2026-09-09 09:19:57 -07:00
..