Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
With recallSync on, prefetch popped only the auth notice and returned without writing the injection log. A session whose init failed for a missing user peer never told the model that memory was off, and the audit file stayed empty for every turn. The recall sync branch now pops the peer notice the way the async branch does and records each turn as injected or recall-sync-empty.
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer
and deferred-save tests that differed only in their inputs. Share the blocking
remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the
added docstrings to the what and the one non-obvious why.
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the
cached path returned none, so recall fell back to the config snapshot. Both paths now return and
store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or
flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the
writer lock, and the writer drains its queue after the join, so a put that raced shutdown is
written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining
budget instead of a fixed ten seconds.
The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop
passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a
gateway platform supplies no user id, the peer notice and tool error no longer recommend
peerName, which would merge every user of that gateway onto one peer. README documents
`injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution.
the manager under test had no peerName and no runtime identity. that
cohort now fails closed instead of minting a fallback peer, so the test
names an owner the way the auth-recovery tests do.
the dashboard login already stamps {user_id, provider} on the websocket
at upgrade time and tui_gateway keeps it as WSTransport.auth_identity,
but _make_agent never read it. every dashboard and desktop session was
built with user_id=None, so memory providers saw no runtime user and fell
back to the configured peer, mixing all logins together (#89794).
_make_agent now reads the session transport's auth_identity and passes
its user_id to AIAgent, the kwarg gateway platforms already use. the
legacy ?token= path, stdio, and the server-internal credential the PTY
child connects with carry no human and pass None, so they keep resolving
to the configured peer. the change is provider-neutral: honcho and any
other memory provider receive the id through the existing initialize()
kwargs.
a desktop or cli session with no peerName in honcho.json and no gateway
user id landed on a peer derived from the session key: user-default-<dir>
for per-directory sessions, user-<channel>-<chat> for keyed ones. every
directory got its own phantom peer, so the operator's turns and memory
never reached their real peer and the injected representation went stale
(#93326).
_resolve_user_peer_id now raises HonchoPeerUnresolvedError instead of
deriving a name. a peer is either the declared peerName or an identity
the transport supplied. the provider records the failure, tells the model
once that memory is off and which key to set, returns the same detail
from tool calls, and stops retrying init because a missing config key
does not heal mid-session. the memory-file migration gate loses its
"no owner and no runtime identity" branch: that cohort no longer has a
session to migrate into. whitespace-only peerName is treated as unset
rather than sanitized to "--".
provider shutdown joined the dialectic, sync and memwrite threads for 5s
each and then the async writer, but the session-init thread,
honcho-base-first and honcho-context-prefetch were never joined. any of
them still blocked in httpx when the interpreter finalized aborted the
process with SIGABRT 134 (#37632, #60616, #33485). the 5s join was also
shorter than the 30s http timeout a blocked call can hold (#33485).
spawn_context_thread now registers each thread under its owner (the
provider or the manager) in a weak registry, and shutdown joins every live
thread of both owners inside one deadline: at least 5s, or the configured
http timeout when longer. the manager refuses new prefetch threads once
shutdown began and flushes a late save() inline instead of respawning the
writer. threads that outlive the budget are named in a warning.
the sdk client exposes close() on its http pool. one client is shared by
every manager with the same identity in a gateway process, so a per-agent
shutdown cannot close it; close_honcho_clients() closes all pools and is
registered with atexit when the first client is built, the pattern the
hindsight and mem0 plugins use. follows #69070, #33701, #7627.
the idle sweep from #71463 left three growth paths open. _peers_cache had no
bound at all, _session_observation (from #98941) grew one entry per session
id and kept orphans when an init failed after add_peers, and a burst of
distinct sessions inside one ttl window was not bounded. the sweep also
evicted sessions whose messages had not reached honcho yet, which in
"session" write mode drops the only copy, and _flush_session re-inserted an
evicted session into the cache with no observation flags, so recall for it
routed from the config snapshot instead of its server config.
_cache and _sessions_cache now cap at 128 entries and _peers_cache at 512,
evicting least recently used first (dict order, refreshed on every hit).
a session with unsynced messages is never evicted by the sweep or the cap.
_configure_session_peers returns the synced flags and get_or_create stores
them under _cache_lock next to the cache entry, so the observation dict can
hold no id the cache does not; eviction drops both. recall reads go through
_cached_session, which stamps updated_at, so a read-only session survives
the idle sweep. follows #71461, #71463, #98936.
_flush_session read the unsynced messages, posted them, and only then set
_synced. in a short-lived run the async writer draining save()'s queue and
the exit-time flush_all() both saw the same batch unsynced, so every turn
of a one-shot run landed in honcho twice (#92458).
each HonchoSession now carries an RLock that _flush_session holds around
the whole select, send, mark sequence. the second flusher enters after the
first marked the batch and sends nothing; different sessions still flush
in parallel. the lock lives on the session object instead of a manager
dict keyed by session id (#86094, #92787): it needs no eviction, and two
flushers of one message list can never hold different locks. tests
adapted from #86094 and #92787.
HonchoSession.messages grew forever -- add_message() appended every
turn and _flush_session() only marked entries synced, never trimmed
them. HonchoSessionManager's four caches (_cache, _peers_cache,
_sessions_cache, _context_cache) had no eviction path besides an
explicit /new reset. A long-lived channel that is never manually
reset accumulates both for the gateway's entire uptime.
Honcho is the durable source of truth (get_or_create() already
re-fetches history from Honcho on a cache miss), so both bounds are
safe: evicting an idle local entry only costs one extra Honcho
round-trip next time that key is used.
- Trim already-synced messages beyond a retention window right after
a successful flush; unsynced messages are never touched.
- Add a rate-limited idle-TTL sweep triggered opportunistically from
get_or_create(), so no new background task/watcher wiring is
needed.
Fixes#71461
The four observation booleans on HonchoSessionManager were manager-wide
mutable state initialized from per-session server configs: every session
setup overwrote them, so the last session to initialize retuned recall
routing for all other sessions the manager serves. Store the server-synced
flags under each session's own id instead; the manager-level fields stay
as the config snapshot and sessions that never synced fall back to them.
injection.sessionStart: unset renders every component in table order, an
empty list renders nothing, a pinned list renders only those names in table
order (not config order), a host block pin beats root, a non-list value is
treated as unset, and initialize() reads the pin.
logging: off by default, the logging key or HONCHO_LOGGING turns it on, a
host block can turn it back off, HONCHO_INJECTION_LOG overrides the path,
each record carries reason/turn/session_key/bytes/payload, an unwritable
path never raises, and a tools-mode prefetch logs its reason. the record
holds the user's representation verbatim, which is why the default stays off.
the two session-context tests that asserted exact context() kwargs now expect
tokens= as well (adopted from #92964). config_schema declares the logging
switch so hermes memory setup shows it.
Salvage of #70951 (Willkons / Alice-Willk-bot). Two of three
session.context() sites never forwarded the configured cap, so Honcho
always returned honcho_chat_summary_long.
Adds the missing get_session_context tokens= regression the sweeper
asked for on that PR.
force_refresh_token gated the adopt-from-disk paths behind the failure cooldown.
After one of our exchanges failed transiently, a 401 within the next 30s returned
None even when a sibling process had already rotated and written a valid grant,
so the operation raised HonchoAuthError with a good token sitting on disk.
Adopting is a disk read, not an exchange; the cooldown exists to stop replaying a
single-use refresh token, so the two adopt checks now run before the gates.
_write_config parsed honcho.json twice under the lock and only wrapped the first
read into ConfigWriteRefused; _refuse_unparseable now returns the parsed dict and
the branches use it. The getattr/isinstance duck-typing collapses to one guard.
cli._read_config reuses oauth's tolerant reader (BOM-tolerant, like the strict
reader the write side uses) instead of its own utf-8 copy.
The unreadable/corrupt variants asserted the same two invariants (writers raise
and leave bytes untouched; rotation fails before the exchange) in five tests;
they are two parametrized tests now. The two lock-spy tests only asserted that
the implementation called the lock helper; the threaded save_config test is the
behavioral guard for the same property.
The CLI's _write_config took only the best-effort file lock, so an in-process
refresh thread could still interleave with a command's read-modify-write. The
dashboard's Honcho save (_write_provider_honcho) still seeded its whole-file
rewrite from a tolerant reader, so a honcho.json that exists but does not parse
was replaced by the active host's block alone from the UI - the same bug class
this PR closes on the CLI and refresh paths. `hermes profile create --clone`
swallowed the new ConfigWriteRefused as "plugin not installed".
One parametrized test covers the web writer for the corrupt and parseable cases.
_write_config applied a command's edits relative to the snapshot the read took, but never moved that snapshot after a successful write. A second write on the same object therefore compared A -> B -> A against A, saw no change, and left disk at B. After a write the snapshot and path now follow the caller's dict, so the next write applies only the edits made since.
_read_config() falls back to ~/.honcho/config.json or the default profile when honcho.json does not exist, so cfg.path never matched the write path and _write_config() wrote the whole dict. That overwrote a refresh rotation a serve child had written onto the honcho.json the setup login created moments earlier. When the local file exists at write time, the seed snapshot is now overlaid with the local file and only the command's edits are applied onto it.
save_config read the file and wrote it back without the locks the token refresh holds around its own read, exchange and write. A refresh that landed between the two steps was overwritten and its consumed refresh token was gone. The read and the write now run under both locks.
install_grant writes the login grant to disk before the wizard's later prompts. _apply_grant_to_host wrote it only into the live cfg, so _write_config saw the grant as an edit and copied it over whatever a serve child rotated onto disk meanwhile. The snapshot now takes the grant too, so the final save leaves the newer on-disk grant alone.
Near-duplicate tests become one parametrized test each: the 401 adopt
cases, the invalid_grant race, the setup apikey answers, the clone and
enable credential checks, and the plain-dict writes. _point_cli_at
replaces the per-class monkeypatch helpers in test_cli. The removed-key
check folds into the merge test, and the missing-file bootstrap check
into the file-lock test; the command-level refusal test already drives
_write_config through ConfigWriteRefused, so the unit test for it goes.
The spy helpers in the install_grant lock test and the reauth bearer
test lose their duplicated closures. Every behavior the removed tests
asserted still has an assertion.
_ReadConfig builds its snapshot in __init__; _read_config no longer
assigns the attributes after the fact. cmd_enable prints its hint
inline, since _no_credential_hint had one caller. The split
force_refresh_token signature and call fit on one line. Docstrings and
comments keep the what and the one non-obvious why. No behavior
changes; _apply_edits keeps its original body.
clone and enable accepted HONCHO_API_KEY from the environment as proof the
block could authenticate and wrote enabled: true. the variable can be absent
from the next process, leaving an enabled block with nothing behind it, which
is the cohort the previous commit set out to remove.
_resolve_api_key takes env=False at both write sites, so a block is enabled
only when honcho.json itself holds a key, an oauth grant, or a self-hosted
baseUrl. status and setup keep the environment fallback for display.
every hermes honcho command read honcho.json, changed a field, and wrote the
whole dict back with no lock. a token refresh in another process that landed
between the read and the write was overwritten with the old access and
refresh tokens, and the next refresh replayed a rotated single-use token.
_write_config now holds the same cross-process file lock the refresh path
holds, re-reads disk under it, and applies only the keys the command changed
since its _read_config(). untouched keys keep their on-disk value, so a
rotation survives; a credential the command set on purpose still wins. a
write with no prior read keeps today's whole-file behavior.
hermes memory setup writes through HonchoMemoryProvider.save_config, which
merged the new values over {} when the existing file could not be parsed
and then replaced the file. same wipe as the oauth and cli write paths,
through a different door. it now reads through _read_config_strict and
raises, so the caller sees the error and the file stays as it was.
a named profile cloned from a default profile that signed in with oauth
got a host block with enabled: true and nothing to authenticate with.
hosts.hermes holds the grant, its apiKey is not inherited (#66125), and
copying the oauth block would make two blocks replay one single-use
refresh token. 'hermes honcho enable' on a fresh profile wrote the same
shape. status then showed the profile as enabled while every honcho
call ran without memory and the plugin quietly stayed inactive.
clone_honcho_for_profile and cmd_enable now resolve a credential for the
target block (its own apiKey, the root apiKey, HONCHO_API_KEY, or a base
url) before writing enabled: true. _resolve_api_key takes the block to
check so both share one definition of "can authenticate".
cohorts:
clone from an oauth default block, no root key: block written without
enabled; the client's auto-enable rule turns it on once a credential
appears (setup apikey writes the root key, or a per-profile login)
clone from a default block with a host-level static key only: same
clone with a root apiKey, an env key, or a base url: enabled as before
enable on an empty or fresh block with no credential: refused, one
message names the profile's setup command and the hosts.<name> key,
nothing is written
legacy blocks already on disk as enabled with no credential: nothing
rewrites them; the plugin already treats them as unusable and stays
inactive; enable now prints the same message instead of "already
enabled"
env-only key: counted as a credential at write time, as the client
does at run time; if the variable later disappears the client still
refuses to initialize the block
after the token endpoint revoked a grant (invalid_grant), running
'hermes honcho setup', choosing apikey and pasting a valid key changed
nothing. the wizard wrote the key to the root apiKey only. hosts.<name>
still held the dead access token under apiKey and the grant under oauth,
and the host block wins the lookup, so status kept reporting the revoked
grant and every call kept failing (#97990).
the apikey branch now drops the host's oauth block and writes the chosen
key onto the host block as well as the root. a dead access token is no
longer shown as the current key, so a blank answer with no other key
aborts instead of keeping the grant. a static host key without a grant
is kept as before.
a login finishing while a refresh was rotating the same honcho.json ran
its read-modify-write unlocked. the two writers could interleave: the
login read the file, the refresh persisted a rotated token, the login
wrote its own copy back and the rotated token was gone.
install_grant now holds _refresh_lock and _config_refresh_lock(path)
around the strict read, the config merge and the persist, the same way
force_refresh_token does. the token response is parsed before the locks
so a malformed grant never holds them.
a truncated or hand-edited honcho.json reads as {} on the tolerant read
path. the next write then replaced the file with only the current host
block: an oauth refresh, a login, or any 'hermes honcho' command that
saves a setting wiped every other host and the root keys.
_read_config_strict now raises on a parse error the same way it raises
on a read error, and logs one sentence naming the file. _rotate_and_persist
treats that as a refresh failure and enters the cooldown without spending
the refresh token. install_grant raises into the setup flow. in cli.py
every write goes through _write_config, which runs the same check first
and raises ConfigWriteRefused; the honcho router, the setup wizard and the
profile sync print the sentence instead of a traceback. the wizard checks
before asking its questions. read paths keep the {} fallback.
follows #95860, which added the strict reader for unreadable files and
kept a .corrupt copy on a parse error. leaving the original file in place
keeps the same bytes without a second copy of the tokens on disk.
_persist_credential promises "leaving all else intact", but it seeded
its write from _read_config, which returns {} on ANY read failure, and
_atomic_write_config replaces the whole file via os.replace — which
needs only a writable parent, so a present-but-unreadable honcho.json
did not stop the overwrite. One OSError (EACCES after a root-owned
write, EIO, a stalled mount) during an automatic token refresh or a
fresh login therefore replaced the store with a single-host file,
destroying every other host's credentials and honcho.json's root
config. No user action is required to trigger the refresh path.
Same defect class as #75206 (P1, fixed for the core auth store in
Add _read_config_strict for the write paths: a missing file still
bootstraps as {}; an unreadable file raises with the store untouched;
genuine corruption still degrades but preserves a .corrupt copy first,
since a truncated store usually holds the other hosts' tokens verbatim.
_rotate_and_persist now takes its strict read BEFORE the exchange —
rotation is single-use, so an exchange whose result cannot be persisted
loses the grant — and threads the dict through to _persist_credential,
which also closes the re-read race between the locked read and the
persist. install_grant seeds its root-merge from the strict reader for
the same reason. Read paths keep their fail-open contract untouched;
both readers now use utf-8-sig so a BOM'd store is not misclassified
as corruption (the wipe vector needing no filesystem fault at all).
Adds TestPersistReadFailure: six tests, four of which fail against the
previous source; the rotate-ordering test additionally pins that no
exchange is attempted against an unreadable store.
reading the live client's api_key inside _force_reauth races the in-place
rotation: a sibling waiter's apply_token_to_client (or the proactive
refresh entered via the honcho property) can swap the bearer before the
read, so force_refresh_token receives the already-rotated token, disk
matches it, and the adopt branch never fires — reproducing the very
force-refresh burst this PR removes.
_authed_call now snapshots the bearer before invoking the operation and
passes that exact token through to force_refresh_token.
desktop spawns multiple serve processes that share one rotating honcho
refresh token. force_refresh_token treated every 401 as "rotate now" even
when a sibling had already persisted a new grant, which can replay a
single-use refresh token and revoke the whole grant.
re-read honcho.json under the existing file lock and adopt if the token
that 401'd is no longer on disk. on invalid_grant, re-read once more
before marking the grant dead.
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.
reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.
test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
- Correct the per-generation reset comment: an in-place updater restart keeps
PTB's update_queue, so old-generation dispatches can briefly exceed
received; the check already treats that as no backlog.
- Make the once-per-stall gate explicit (cap the heartbeat count) instead of
relying on `!=`.
- Drop the dead getattr in _record_updates_received (only reachable after
_record_polling_progress dereferenced the same instance state); keep the
fallbacks in the heartbeat check and group-99 handler, which sibling
watchdogs share because object.__new__ adapter doubles exist in tests.
- Split the single invariant test so a failure names the broken guard:
stall-once, re-arm-on-progress, generation-reset; drop the unused _app mock.
Follow-up to the cherry-picked #102383 commit. The check as written was neither
sensitive nor specific:
- It aged the newest received update, so a wedged PTB dispatcher was never
reported while new updates kept arriving more often than every 300s (probe:
1 update/250s for an hour -> 0 reports).
- `delivered` counted only MessageEvents reaching the gateway handler, while
`received`/`dispatched` counted every Update; a single handled callback_query,
reaction, unauthorized user or unmentioned group message produced a false
ERROR after 300s of quiet.
- `_record_updates_received` skipped the generation/teardown guard
`_record_polling_progress` applies, and the counters never reset across
polling generations, so a late response from a fenced poll or a reconnect
inflated the backlog.
Now `received` and `dispatched` count the same population (every fetched
update reaches the group-99 catch-all) and the report fires when a backlog
persists with no dispatch progress across two 90s heartbeats, once per stall,
re-armed on progress, reset per generation, at WARNING (diagnostic only;
#71240 owns recovery). The delivered counter and the `note_inbound_delivered`
facade method are dropped; the once-per-adapter "no message handler" error on
BasePlatformAdapter.handle_message stays. `_record_polling_progress` returns
whether the round-trip was accepted so the received stamp reuses its gate.
Tests trimmed to two invariants; every guard proven red by mutation.
Refs #102260
Every Telegram health probe measures the transport. A getUpdates round-trip
that returns 200 proves bytes are moving and nothing else: the stall watchdog
(#92991), the pending-update probe (#42909/#55769), the get_me() heartbeat
(#66377) and the polling-progress instrumentation all stay green while updates
arrive and then die downstream. The adapter then publishes "connected", logs
nothing at all, and is indistinguishable from a bot nobody has messaged.
That is #102260: three weeks of telegram.state "connected" plus "polling
confirmed healthy: getUpdates progressing (generation 1)" with zero inbound
reaching the agent, surviving every restart. Two of the issue's three
hypotheses do not hold on this code — _record_polling_progress fires on every
round-trip (not only at start_polling), and _send_path_degraded is cleared on
the first confirmed round-trip — and the reporter's own observation that fresh
messages are received but not processed places the failure downstream of the
transport, in the one stretch with no instrumentation at all.
Add the missing delivered side of the accounting:
- received: updates Telegram handed the process, read from the getUpdates
envelope the adapter already parses (an empty result proves the transport,
not arrival, so only non-empty results count).
- dispatched: updates PTB's dispatcher carried through the whole handler
chain, stamped in the existing group-99 catch-all before its early returns.
- delivered: inbound events that reached the gateway's message handler,
stamped in BasePlatformAdapter.handle_message for every platform.
_check_ingress_delivery_gap runs on the existing heartbeat and, when updates
arrived but nothing was delivered for 300s, names the broken hop: received >
dispatched means the dispatcher is not draining, dispatched > delivered means
Hermes is dropping what arrives. Diagnostic only — a received update
legitimately reaches no gateway turn, and reconnecting a healthy transport
cannot repair a dropped update, so this never drives recovery.
Also make the silent discard on the shared funnel speak: handle_message
returned with no log when no message handler was installed, so a mis-wired
adapter discarded 100% of inbound while connected and able to send.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018pT5hFJBRfLj8KMqFhm3qz
The composer showed "<model> · Med" in one truncating pill and the only way
to change the effort was to open the model menu, find the active model's
row, and hover it for the per-row options submenu. Users read the pill as
"this model is medium only" and never found the submenu.
- New `ReasoningPill` next to the model pill: shows the active model's live
effort (session value, else the profile default) and opens the same
Thinking / Fast / Effort rows the catalog submenu offers, for the active
model only. Hidden when the catalog reports `reasoning: false`; stays
while capabilities are unknown so it never flickers during the fetch.
Folds away with the model pill in the compact composer stages.
- `useModelMenuController` (shell sibling) now owns the session write /
preset / optimistic-store / rollback logic that lived inside
`ModelMenuPanel`; the model menu and the new `ReasoningMenuPanel` share it
so an edit from either surface is one code path. Tiles get their own
pill bound to their SessionView, primary or tile — never the globals.
- `ModelOptionsContent` (the submenu body) is exported container-free so
the pill's top-level menu renders it without a Radix Sub wrapper.
- The model pill drops the effort suffix (`formatModelPillLabel`: name +
Fast); `formatModelStatusLabel` had no other caller and is removed.
- `currentModelCapabilities()` in lib/model-options resolves the active
pick's caps through `catalogProviderMatches` (aliases, custom slugs).
Live (headless Electron + worktree `hermes serve`, CDP): before — one
pill "Deepseek V4 Flash · Low", no effort control; after — "Deepseek V4
Flash" + "Low" pill; pick High → `config.get reasoning` on the live
session returns high; a `reasoning:false` cap unmounts the pill; the
catalog row submenu still writes through and the pill mirrors it.
Credit: the dedicated-pill direction was proposed independently in
composer selector on current main with the shared-controller shape.
Two CI reds from the previous commits.
benchmark_browser_eval.py moved from scripts/ (exempt) to evals/ and its
bare shutil.which("npx") tripped test_no_unreviewed_bare_managed_runtime_
lookups. evals/ are standalone user-invoked benchmark programs, the same
class as scripts/ and skills/, and three other eval files already call
which() the same way; the guard just never saw them because the harness
that carried them lived under scripts/. Exempt the directory.
list_os_marked_tests.py used Path.rglob, which raises FileNotFoundError
when a __pycache__ directory disappears mid-scan (a sibling job in the
same workspace). The managed-runtime guard already switched to os.walk
for the identical TOCTOU; do the same here. Same output, sorted.