Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
The redelivery hook keyed on the canonical flood_control:<seconds> result, so
two real refusals slipped past it and armed no timer, leaving the reply for the
next restart. A short wait that outlived the send retries raised instead of
failing closed, and an edit refused again after its inline wait returned the
platform's raw text. Both now fail closed canonically, the second carrying the
new delay rather than the first refusal's.
The ledger also accepts a row still carrying the platform's own wording, so a
row persisted by an unnormalized path is dated from the delay it states instead
of the generic default. Without that a boot sweep claims it at once and spends
its one attempt inside the penalty. Matching requires the flood wording as well
as a delay, so an unrelated retry suggestion is never read as a flood.
Six new assertions fail without this change. 770 passed across the ledger,
Telegram, send-retry and queued suites.
Round-2 review. Making _arm_flood_timers_for_waiting_rows read the ledger in a
worker thread introduced a race: a flood refusal recorded during that thread's
yield had its own _schedule_flood_redelivery request declined (the timer still
held the slot), and the timer then cleared its slot on a snapshot taken before
the new row committed, leaving that reply with no timer until the next reconnect
or restart. The read is a single indexed SELECT; doing it synchronously keeps the
arm decision atomic with respect to concurrent schedule requests. A test drives a
refusal during the timer's own redelivery send and asserts the same timer arms it.
Second review pass on the flood-retry change.
- sweep_recoverable returns a dead owner's not-yet-due flood row flagged
`adopted` (with its `not_before`) instead of dropping it from the result,
so _claim_pending_obligations clears its session's resume_pending flag like
every other claimed row; the answer is in the ledger and the turn must not
be re-run at boot. _redeliver_claimed_obligations skips adopted rows and
still arms the timer for them.
- A legacy row without adapter_profile is normalised to 'default' when the
boot sweep claims or adopts it: the runtime sweep matches profiles exactly
and the timer asks for 'default', so an adopted NULL row could only wake the
timer without ever being sent.
- A sleeping timer is replaced by a shorter refusal's timer; the shorter one
re-arms for the longer row after its sweep. A timer already sweeping is
never cancelled from outside.
- pending_flood_retries is read in a worker thread like the other ledger calls.
- Adoption tests seed a distinct dead-owner stamp and assert ownership moves.
- Scratch tags and long comment lines removed.
Review on the first cut reproduced two gaps.
The raw UTF-16 length of the stored reply said nothing about how many
requests the adapter made: MarkdownV2 escaping turns 3000 dots into 6000
units, two Telegram messages, and the send result does not say which chunk
the platform refused. Drop the length-based certainty (and its helper and
constant) and mark every flood redelivery with the rate-limit marker, at
runtime and at boot. The cost is a marker on a reply nothing of which was
delivered; the alternative was a silent duplicate.
Releasing an unsent runtime claim always wrote send_path_degraded, so a flood
row whose resume flag could not be cleared (or whose adapter vanished before
dispatch) left the flood timer's list and stayed stranded until a reconnect or
restart. The claimed row now carries its pre-claim error and the release
writes it back, so the row waits the platform's figure once more and is sent
on the next timer.
The Telegram adapter fails a flood-controlled final send closed as
flood_control:<seconds> so the send coroutine never sleeps a long penalty
(#91969), on the understanding that the delivery ledger owns the wait. The
ledger did not: sweep_failed_for_runtime only replayed send_path_degraded
rows, so a flood-refused row sat in 'failed' until the next restart, whose
sweep_recoverable then redelivered it hours late prefixed with "Recovered
reply, the gateway restarted during delivery, so this may be a duplicate".
Observed: a reply refused at 08:58 UTC (both the MarkdownV2 send and the
plain fallback got flood_control:185) arrived at 12:15 UTC after a restart,
labelled as a possible duplicate although the platform had never accepted it.
Ledger (gateway/delivery_ledger.py):
- flood_control rows are runtime-retryable, but only once their own deadline
has passed: the refusal's updated_at plus the platform's wait
(flood_not_before). Neither an early timer nor a reconnect sweep spends a
redelivery attempt inside the penalty window.
- A flood refusal of a reply that fits in one Telegram message (4096 UTF-16
units) proves non-delivery, so that redelivery carries no duplicate marker.
A chunked reply may have had its first chunk accepted before the refusal
and keeps the marker.
- Claiming a flood row clears the stale refusal (last_error NULL, state
'attempting'), so a resend interrupted before mark_delivered is seen as
uncertain by the next boot and gets the marker.
- At boot, a dead owner's flood row that is not yet due is adopted (owner
re-stamped, no attempt spent) instead of being resent early.
- pending_flood_retries() lists this process's waiting flood rows per adapter
identity with the earliest deadline.
Runner (gateway/run_startup.py):
- _schedule_flood_redelivery arms one timer per adapter identity that runs
the existing runtime sweep after the wait (capped at 15 minutes per sleep;
the row's deadline, not the timer, decides eligibility, so a capped timer
wakes early, sends nothing, and re-arms for the remainder). The slot stays
occupied until the timer ends and only the running timer may arm its
successor into it, so a refusal during the sweep can never strand the row.
- _arm_flood_timers_for_waiting_rows runs after every redelivery pass (boot
and runtime), covering adopted rows, rows skipped as not yet due, and rows
refused again.
Adapter (gateway/platforms/base.py): _finalize_delivery_obligation arms the
timer on a flood_control failure, best-effort inside the existing try.
Tests: 35 in tests/gateway/test_delivery_ledger_flood_retry.py, with a
controllable clock shared by ledger and runner. Each guard was mutation
tested. Existing ledger, reconnect and redelivery suites (131 tests) pass.
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.
Fixes#104724
The ACP wire contract has session/new as a separate round trip from
session/prompt precisely so a client can open a session before it knows a
prompt is coming, and at least one shipping client opens sessions it will
never prompt: bb discovers a model catalog by spawning a throwaway
`hermes acp`, sending initialize + session/new, reading
NewSessionResponse.models, and killing the process. Its discovery cache
TTL is 60s, so an open editor re-probes continuously.
_persist() created the state.db row the moment a session was created, so
every such probe left a permanent message_count=0 shell. Measured on one
workstation: 116 accumulated, appearing on a time cadence rather than a
per-conversation one (26 real ACP user turns vs 13 shells in a day; exactly
120.0-minute spacing overnight with zero user activity).
The shells are indistinguishable from real chats in the session list and
`hermes sessions prune` cannot remove them: the rows are never ended, so
ended_at stays NULL and prune skips them by design, leaving
`hermes sessions delete <id>` one id at a time as the only remedy.
Defer row creation until the session has history. Nothing is lost for a
genuine conversation: AIAgent._ensure_db_session() creates the row on the
first turn and the post-prompt save_session() lands the ACP metadata on
top, which makes the create-time write redundant for every session except
the one case that should not be recorded at all.
Gate on state.history rather than message_count so fork_session, which
deep-copies a non-empty history into a fresh id, still persists at once.
Verified by differential execution of an identical probe against both
trees: on 693641aa8b an unprompted session/new adds a row, with this
change it adds none, and a session that does receive a prompt persists
in both. tests/acp: 138 passed. Reverting only the source change leaves
the two new assertions failing, confirming they exercise the fix.
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.
Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>
Keep parsing, contracts, gates and persisted goal mutations in one dispatcher. Adapters retain authorization, rendering and scheduling; TUI drafting resolves the target session profile off the RPC reader. Document ACP as unsupported rather than implying a goal loop exists.
The image-part membership test in _payload_chars() ran on every dict in
the request, so a tool JSON Schema with a parameter named "type" (its
value is the sub-schema dict) or a multi-type value like
["string", "null"] raised TypeError: unhashable type before the
request reached the provider. Guard the membership test with
isinstance(str): structured "type" values are payload data, not
content parts, and keep the legacy walk for them.
The identity-guarded token pop appeared twice (the runner's finally and the
new worker-start except); a future edit to one copy would silently reintroduce
the sticky-token class. Hoist into a local closure called from both sites, and
extend the _run_hook_callback_bounded docstring with the new skip reason.
Behavior-preserving follow-up on #104651.
Desktop AGENTS.md requires a compat fallback to be 'tied to an identified
older runtime' so a future cleanup knows when it can be deleted. The
original comment said only 'older backend plugins'; name the actual
boundary: backends before #35395 (May 2026) omit the attachments key.
Review follow-up on the salvage of #104529.
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
estimate_request_context_tokens (the stale-stream / non-stream watchdog and Codex TTFB floor
estimator) did len(str(payload)) // 4 over the wire payload, so one native screenshot read as
~100K tokens and selected the 1200s giant-conversation floor while the provider's real prompt
was ~2K (#63871, #76411). Image content parts on both wire shapes (Chat image_url / Responses
input_image) now cost the per-image price learned from provider usage (agent/image_token_cost,
bound per turn; worker threads inherit it via copy_context); base64-looking STRINGS stay text
and tool schemas mentioning image are not images.
A/B, one 300KB screenshot: main estimate 100,545 -> codex floor 1200s / stale 300s; now 2,025
-> floor 0s / stale 180s.
Design and boundary (structured parts only, tools/instructions opaque, plain data-URL strings
stay text) from #76471 by @crdesign8; re-authored to reuse the learned cost instead of a second
fixed constant, and trimmed to 2 invariant tests.
The Chat Completions schema has no `name` field on `role: tool` (only on
the long-removed `role: function`), but Hermes carries the tool name over
onto the result message. Permissive providers ignore it; strict ones
(aki.io) reject the whole payload with `contains item with unknown key
name`, which breaks every tool call in the session.
Follow-up to the review on #51365:
- The strip now goes through the copy-on-write `mutable_msg()` path in
`convert_messages()`, preserving the identity/copy-on-write contract
instead of mutating `msg` in place.
- `handle_max_iterations()` hand-builds its summary payload and calls
`chat.completions.create()` directly, bypassing the transport, so it
leaked `name` even with the transport fixed. It now mirrors the same
role-qualified removal, next to the existing tool_name/codex_*/timestamp
strips.
The removal is role-qualified: `name` stays on user/assistant messages,
where it is schema-valid.
Regression coverage on both paths; both tests fail without the fix.
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
Diagnosing the 1,393-agent run's cache misses took a state.db join and a live
probe because none of the three were on the one line we log per call:
write=<n> cache_creation tokens; a write costs 50x a read, so this is the
money, and "read stuck, write large" on consecutive lines is a
routing miss visible without a probe
id=<id> the provider's response id (Anthropic msg_..., OpenRouter/Nous
gen-...): what a provider needs to look a request up
upstream=<n> who actually served, when the route reports it (OpenRouter's
`provider`); how we learned GMI was not serving
Fields are appended to the existing line and omitted when absent, so every
existing parser (evals/postmortem, the two cache_prefix probes) keeps
matching; the forensics parser reads them when present.
The streamed chat path fabricated `id="stream-<uuid>"` and dropped the
chunks' id/provider on the floor; it now keeps the first chunk's id and
provider, falling back to the fabricated id only when the stream never sent
one. Nothing depended on the prefix.
Live (Nous, Fable 5.1, both wires): chat -> `write=3759
id=gen-1788727882-... upstream=Anthropic`; native -> `id=gen-1788727891-...`
(the Anthropic Message object's id was already real). Tests: fields present
and omitted, prefix unchanged; forensics parser reads new and old lines.
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
The round-2 change made check_public_surface refuse to report a clean diff
without a merge-base (exit 2). That correctly exposed that the lint job's
depth-1 checkout plus a depth-1 fetch of the base has NO merge-base, so the
advisory step had been silently reporting 0 drops on every PR. The step now
deepens both sides until a merge-base exists and carries continue-on-error so
an advisory check can never block the Windows-footguns job it rides in.
Independent review: a nonexistent base ref under --strict reported zero
drops and exited 0 (a mis-fetched CI job would look clean); the script now
verifies both refs and the merge-base and exits 2 otherwise. And a method
moved from a class into a mixin/base defined in the same module that the
class still derives from was reported as removed although the attribute
still resolves; public_methods now collects the methods REACHABLE on each
class through its in-module bases. Replay of #102117 at open: 1,703 names /
341 modules and 126 test defs / 52 files unchanged; methods 1,000 -> 951
(the 49 were in-module mixin extractions, i.e. the false positives).
Test: unresolvable ref -> exit 2 in both modes; a method extracted into an
in-module base is not reported.