Eval probes wrote fixtures, receipts and evidence to hard-coded /tmp paths and
two live A/B tasks literally instructed the model to work in /tmp. They now
derive locations from tempfile.gettempdir() / os.tmpdir() (env overrides kept),
usage examples use relative output names, and the desktop e2e screenshot dirs,
perf scripts and the short-session repro fixture stop naming /tmp. Also adds
the explicit encoding= the windows-footgun check wants in the touched files.
Five test files and five evals scripts merged since this branch's base still import the
moved names via gateway.platforms.base. They work (base.py imports the names for its own
use) but the PR's invariant is that in-tree code imports from the defining module.
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.
Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
evals/token_accounting/replay_gates.py runs the REAL AIAgent turn loop against a local fake
chat-completions server with scripted usage.prompt_tokens, in three shapes (CLI same-object
history, gateway JSON-reloaded history, SessionDB close/reopen + fresh agent) x two arms
(transcript 3x threshold by bytes/4 while real usage is under; transcript tiny while real usage is
over). Acceptance: no gate fires when real usage is under threshold regardless of estimate
inflation; every gate fires once real usage is over.
origin/main 5f406d88ea: cli/gateway/restore inflated all FAIL (1 local compaction each on the
estimate). This branch: 6/6 PASS, anchor restored in the fresh process.