Five test files and five evals scripts merged since this branch's base still import the
moved names via gateway.platforms.base. They work (base.py imports the names for its own
use) but the PR's invariant is that in-tree code imports from the defining module.
Diagnosing the 1,393-agent run's cache misses took a state.db join and a live
probe because none of the three were on the one line we log per call:
write=<n> cache_creation tokens; a write costs 50x a read, so this is the
money, and "read stuck, write large" on consecutive lines is a
routing miss visible without a probe
id=<id> the provider's response id (Anthropic msg_..., OpenRouter/Nous
gen-...): what a provider needs to look a request up
upstream=<n> who actually served, when the route reports it (OpenRouter's
`provider`); how we learned GMI was not serving
Fields are appended to the existing line and omitted when absent, so every
existing parser (evals/postmortem, the two cache_prefix probes) keeps
matching; the forensics parser reads them when present.
The streamed chat path fabricated `id="stream-<uuid>"` and dropped the
chunks' id/provider on the floor; it now keeps the first chunk's id and
provider, falling back to the fabricated id only when the stream never sent
one. Nothing depended on the prefix.
Live (Nous, Fable 5.1, both wires): chat -> `write=3759
id=gen-1788727882-... upstream=Anthropic`; native -> `id=gen-1788727891-...`
(the Anthropic Message object's id was already real). Tests: fields present
and omitted, prefix unchanged; forensics parser reads new and old lines.
N concurrent real AIAgent sessions on a growing tool loop; per-call cache_read /
cache_creation / response id / upstream / prefix shas; consecutive pairs classified
ideal / stuck / collapse. argparse (--provider nous|openrouter|anthropic, --wire,
--pin, --settle, --ttl, --model); wire defaults to what Hermes would pick. Summary
JSON carries bad pairs with both response ids. Smoke-run against Nous from the repo
path (picked chat via nous_api_mode; 6/6 ideal).
evals/postmortem/ turns the one-off audit behind tracking issue #103563 into
something anyone with a Hermes state.db copy (and optionally rotated
agent.log*) can run on their own fan-out:
forensics/ common.py discovers the run tree (root = most descendants,
compression-rollover children excluded so cost buckets stay
disjoint), fits pricing from estimated_cost_usd, and five lanes
recompute the OBSERVED figures: tokens (buckets, depth/duration
shares, context reconstruction, excess-cache-write proxy, cap
replay), logcalls (per-call cache behaviour from agent.log with
coverage printed first; strict and loose plateau definitions
reported separately), delegation (timeouts, orphaned children,
polling hours, batch-join withheld child-hours, truncated
summaries), tools (hardline blocks, foreground refusals,
whole-file rewrites), goal_loop (nudges, parked barrier), rework
(public-surface drop at PR open + post-open commit inventory).
Every figure is labeled OBSERVED or MODELED.
live_ab/ the per-PR A/Bs (real code paths, fake providers, temp
HERMES_HOME), paths from argv.
review_probes/ the independent /review's probes, credited and adapted; each
reproduced a round-1 defect and the fixed head must pass it.
run.py runs the offline probes against one or two checkouts and prints
PASS/FAIL side by side (--live adds the ones that spend cents).
tests/ synthetic-DB smoke test for the lanes and runner.
On the run's DB the lanes reproduce the tracking issue's population exactly
(1,394 sessions, 93,284 calls, $19,302.59; cache_write $11,159.76) and on
main vs an integration checkout of the 13 PRs the runner shows every probe
FAIL -> PASS (two guard-only probes pass on both, noted in run.py).
The trajectories are deliberately not shipped: the DB holds 51,956 home
paths, 5,341 e-mails, private IPs, chat ids and real-shaped credentials in
tool output. The lane reports and recomputed JSON are in a secret gist
linked from #103563.