The r3 collapse to '_sessions.get(sid) or {}' treated a registered-but-empty session dict as
detached, writing the tour_bridge verdict to a throwaway — so an unanswered probe was re-paid
on every action. Restore the 'is None' distinction.
test_every_intree_plugin_declares_what_it_implements detects upscale support via the literal
'"upscale",' whitelist entry; the r3 tuple collapse turned it into '"upscale")'. Tool schema
(get_tool_definitions dump) is byte-identical to base 113f04616b before and after.
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog). A
profiled session with ~130 children was carrying ~1000 threads. All
of these timers now run on a single process-wide daemon thread.
- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
one Condition-driven daemon thread. schedule(fn, interval) -> handle;
handle.cancel(wait=) blocks for an in-flight run like the old join.
A callback returning False stops itself; a raising callback is logged
at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
_lease_refresh_interval; lease-lost / refresh-error interrupt paths
and the stop-event fencing are unchanged; the join(timeout=1.0) is now
cancel(wait=1.0) so the interrupt clear still runs after any in-flight
tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
schedule(); the poll body is _tick(), same sampling state machine.
Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132. At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:
1. bind_subagent_parent() stored the agent strongly in the
`hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
for its own turn, and every asyncio Handle/Future scheduled during
that turn (LSP reader loops, kernel pipe transports) snapshots the
Context — 56 live Contexts held 14 finished children after the bench.
The ContextVar now holds a weakref (non-weakrefable doubles fall back
to a closure); get_active_subagent_parent() dereferences it.
2. AIAgent.close() cleared _session_messages but not the
_db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
every successful DB flush) nor _streamed_assistant_text_parts, so the
agent — kept alive by (1) — retained every message dict. close() now
drops both.
The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.
Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.
What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.
Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).
Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
`limits=` never reached them, so mounts ran on httpx defaults with a 5 s
keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).
Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
Compaction put the 'config.yaml' literal within the config-read guard's
proximity window of a raw yaml.safe_load. Use the canonical raw reader with an
explicit path (same semantics: raw primary file, no defaults/overlay).
Compaction moved run_bang_command's Popen within the env-guard scanner's
proximity window of _bang_env's os.environ.copy(). Route through the single
factory (build_subprocess_env() == _sanitize_subprocess_env(os.environ.copy()))
and allowlist the file for the import-failure fallback copy with justification.
cmd_sessions used 'with db:' which breaks test doubles lacking the context
manager protocol (13 reds in test_sessions_pin/delete/export). Restore the
explicit try/finally db.close() (same semantics for real SessionDB). Add
get_session/count_prune_matches to the FakeDB doubles in test_sessions_delete
instead of re-adding getattr guards to production code.
readPersistedPoolLimits() runs at module evaluation and logs through
rememberLog() on every branch, but hermesLog / desktopLogBuffer /
desktopLogFlushTimer / desktopLogFlushPromise were declared ~110 lines
later. esbuild lowers const/let to var, so the packaged desktop died on
every launch with "Cannot read properties of undefined (reading 'push')"
(#101941, #101960). Moving the four declarations above the read fixes the
crash and keeps the early [pool-limits] line in desktop.log.
Salvaged from #101945 (test dropped: Desktop E2E lane is disabled in CI).
A roster click on a bot whose canonical Bot Chat is already open only
fronted the tile: the pane kept whatever transcript it last painted,
which can predate rows the bot wrote while the user was elsewhere (a
cron delivery, a teammate's message_agent, another bot's turn). The
stale snapshot persisted until the next user turn — #95600's forceResume
only covered the not-yet-open registry path.
Reuse refreshOpenBotChat (the #99393 reclaim mechanism) on the fronted
branch so forceResume re-pulls the latest transcript. Regression test
pins the behavior: fronting an open Bot Chat now requests the canonical
registry open.
The notification action (`NotificationItem`) rendered as
`variant="textStrong" size="xs"` — an 11px underlined muted-grey text link
with a ~44x20px hit target. On the data-training confirm toast raised by
`surfaceModelSwitchConfirm` / `confirmModelWarning` (e.g. picking
`muse-spark-1.2-contributor`) it read as a footnote, not the one action
the toast exists for, and users reported not being able to "press to
accept".
Promote it to the SDK's `default` variant at `size="sm"`: a filled
primary button, larger hit target, obvious affordance. No new styles.
Salvaged from #96562 (toast half only). Refs #96563.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).
Missing alias identified by @Steve-prog001 in #101905.
Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
commandcode (api.commandcode.ai) exposes authoritative
context_length via /models (muse-spark 1M, etc.) but as a
known provider it skipped the custom-endpoint probe at step 2
and has no models.dev entry, so every model fell through to the
256K DEFAULT_FALLBACK. Add a provider-aware branch mirroring
gmi/nous to resolve via _resolve_endpoint_context_length.
Fixes GOAT docs vs status-bar mismatch: muse-spark 1M was shown
as 256K.
Muse Spark 1.2 family (api.meta.ai) ships 1M context (models.dev
opencode/muse-spark-1.2 = 1048576, meta/muse-spark-1.2 = 1048576).
Zen/GO SG /v1/models only returns id (no limit.context), and
models.dev lookup via opencode was missing a hardcoded fallback, so
get_model_context_length fell back to DEFAULT_FALLBACK_CONTEXT=256k.
Banner showed Context: 256,000 for both zen and router-sg lanes.
Add longest-prefix entries 'muse-spark' and 'muse' = 1_048_576 so
all variants (1.1, 1.2, contributor, contributor-free) resolve to 1M
without network.
Add meta/muse-spark-1.3 and meta/muse-spark-1.3-contributor to the
OpenRouter curated list, the meta-ai provider fallback, the
opencode-zen / opencode-free / opencode-go floors, the setup-wizard
shortlist, and regenerate the hosted model catalog.
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.
Fixes#101668
Refs #47742
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.
Both gates now go through one _profile_rejects_tool_media() veto.
Refs #89981
(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)
The naive line scanner's open() regex matched the human-readable phrase
"next picker open (or refresh)." inside the warning string, tripping the
Windows footgun gate. Reword to "next picker open or refresh."
With the picker served from resident caches only, a cold Nous row renders
every model locked (free_tier_pending) until the background prewarm lands.
Surface why on the row's existing warning slot so the user isn't left with
an unexplained greyed-out list; never override an auth warning.
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.
Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
Conflicts in 10 tools/ files resolved as the union of both sides'
simplifications (integration3's earlier r3-33-* landings vs the parent
r3-33 reconciliation):
tools/memory_tool.py, memory_tool_store.py, microsoft_graph_auth.py,
microsoft_graph_client.py, plugin_guard.py, project_tools.py,
registry.py, schema_sanitizer.py, self_repo_guard.py,
session_search_tool.py
Tool-schema dump (/tmp/rf/tools_schema_dump.py) byte-identical to
113f04616b.
The contributor fix covers remote rows. The reporter's video shows the
sibling shape: with a remote gateway active, the LOCAL twin carries the
'default-this-device' alias, and message_agent's local resolver only
knows bare profile names / 'hermes'. Emit the same target annotation
whenever a local row's alias differs from its resolvable handle.
The Bot Mode mention middleware built message_agent targets from
botHandle(), which prefers a roster row's source-qualified UI alias
("default-vera"). Neither resolver accepts that form — the relay matches
canonical handle/profile (± @connection-id) and the local path a bare
profile name or "hermes" — so remote handoffs died with "No teammate
named" before enqueue.
Annotate the canonical form instead: profile@connection-id for remote
rows, canonical bare handle (default→hermes) for local ones. Pin the
profile@connection form on the relay side too, so the emitted target
stays inside the documented resolver contract (#97678).