Commit Graph

31579 Commits

Author SHA1 Message Date
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
545e374a68 fix(integration): _tour_request keeps the verdict on an empty-but-registered session record
The r3 collapse to '_sessions.get(sid) or {}' treated a registered-but-empty session dict as
detached, writing the tour_bridge verdict to a throwaway — so an unanswered probe was re-paid
on every action. Restore the 'is None' distinction.
2026-09-03 02:47:37 -07:00
Teknium
48b336ef7e fix(integration): fal image_gen plugin keeps trailing-comma passthrough whitelist the capability-coverage test keys on
test_every_intree_plugin_declares_what_it_implements detects upscale support via the literal
'"upscale",' whitelist entry; the r3 tuple collapse turned it into '"upscale")'. Tool schema
(get_tool_definitions dump) is byte-identical to base 113f04616b before and after.
2026-09-03 02:47:01 -07:00
Teknium
23b9ffc4fa fix(integration): restore subprocess stdin=DEVNULL / utf-8 encoding guards and windows-footgun gates dropped by round-3 compaction
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
2026-09-03 02:46:19 -07:00
Teknium
561b053f79 perf(agents): run per-child timers on one shared scheduler thread
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog).  A
profiled session with ~130 children was carrying ~1000 threads.  All
of these timers now run on a single process-wide daemon thread.

- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
  one Condition-driven daemon thread.  schedule(fn, interval) -> handle;
  handle.cancel(wait=) blocks for an in-flight run like the old join.
  A callback returning False stops itself; a raising callback is logged
  at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
  scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
  idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
  stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
  _lease_refresh_interval; lease-lost / refresh-error interrupt paths
  and the stop-event fencing are unchanged; the join(timeout=1.0) is now
  cancel(wait=1.0) so the interrupt clear still runs after any in-flight
  tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
  schedule(); the poll body is _tick(), same sampling state machine.

Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132.  At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
2026-09-03 02:44:24 -07:00
Teknium
2b55ded1ac perf(state): keep delegate-child transcripts out of the trigram FTS index (schema v30)
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.

Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.

The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
2026-09-03 02:35:37 -07:00
Teknium
c96568f66c perf(delegation): finished delegate children no longer pin their transcripts in the parent heap
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:

1. bind_subagent_parent() stored the agent strongly in the
   `hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
   for its own turn, and every asyncio Handle/Future scheduled during
   that turn (LSP reader loops, kernel pipe transports) snapshots the
   Context — 56 live Contexts held 14 finished children after the bench.
   The ContextVar now holds a weakref (non-weakrefable doubles fall back
   to a closure); get_active_subagent_parent() dereferences it.

2. AIAgent.close() cleared _session_messages but not the
   _db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
   every successful DB flush) nor _streamed_assistant_text_parts, so the
   agent — kept alive by (1) — retained every message dict. close() now
   drops both.

The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.

Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
2026-09-03 02:35:37 -07:00
Teknium
c3b411dfb7 perf(agents): share one httpx transport pool across every agent's client
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.

What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.

Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).

Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
  SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
  again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
  default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
  `limits=` never reached them, so mounts ran on httpx defaults with a 5 s
  keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
  pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).

Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
2026-09-03 02:35:21 -07:00
Teknium
80fae22bf5 fix(lsp): share one pyright process across git worktrees via workspaceFolders
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
2026-09-03 02:32:15 -07:00
Teknium
9cee679831 bench: explicit utf-8 encoding on text-mode opens (ruff PLW1514) 2026-09-03 02:31:59 -07:00
Teknium
45b0a0ae25 bench: --reply-kb payload padding, transport/live-agent gc counts, after-snapshot rows in --compare 2026-09-03 02:31:59 -07:00
Teknium
c77b9d637b bench: fan-out resource harness (threads/RSS/fds/pyright/kernels/httpx clients per N children x W worktrees) 2026-09-03 02:31:59 -07:00
Teknium
e1487be830 fix(integration): cron preflight reads primary config.yaml via read_user_config_raw()
Compaction put the 'config.yaml' literal within the config-read guard's
proximity window of a raw yaml.safe_load. Use the canonical raw reader with an
explicit path (same semantics: raw primary file, no defaults/overlay).
2026-09-03 02:08:03 -07:00
Teknium
5540860ca6 fix(integration): bang_shell builds its env via build_subprocess_env()
Compaction moved run_bang_command's Popen within the env-guard scanner's
proximity window of _bang_env's os.environ.copy(). Route through the single
factory (build_subprocess_env() == _sanitize_subprocess_env(os.environ.copy()))
and allowlist the file for the import-failure fallback copy with justification.
2026-09-03 02:07:33 -07:00
Teknium
dc130e54bf fix(integration): repoint auxiliary bridge source-inspection test to generalized field loop
gateway/run.py now bridges MODEL/BASE_URL/API_KEY via one (field, suffix)
table writing AUXILIARY_{_upper}_{_suffix}; the env keys set are unchanged.
2026-09-03 02:06:51 -07:00
Teknium
0fc204042d fix(integration): sessions CLI — close db via try/finally, complete test doubles
cmd_sessions used 'with db:' which breaks test doubles lacking the context
manager protocol (13 reds in test_sessions_pin/delete/export). Restore the
explicit try/finally db.close() (same semantics for real SessionDB). Add
get_session/count_prune_matches to the FakeDB doubles in test_sessions_delete
instead of re-adding getattr guards to production code.
2026-09-03 02:06:29 -07:00
liuhao1024
05f548f35d fix(desktop): declare rememberLog state before the top-level pool-limits read
readPersistedPoolLimits() runs at module evaluation and logs through
rememberLog() on every branch, but hermesLog / desktopLogBuffer /
desktopLogFlushTimer / desktopLogFlushPromise were declared ~110 lines
later. esbuild lowers const/let to var, so the packaged desktop died on
every launch with "Cannot read properties of undefined (reading 'push')"
(#101941, #101960). Moving the four declarations above the read fixes the
crash and keeps the early [pool-limits] line in desktop.log.

Salvaged from #101945 (test dropped: Desktop E2E lane is disabled in CI).
2026-09-03 01:14:37 -07:00
cmyyy
3ea71a47b3 fix(desktop): refresh Bot Chat transcript when a roster click fronts an already-open tab
A roster click on a bot whose canonical Bot Chat is already open only
fronted the tile: the pane kept whatever transcript it last painted,
which can predate rows the bot wrote while the user was elsewhere (a
cron delivery, a teammate's message_agent, another bot's turn). The
stale snapshot persisted until the next user turn — #95600's forceResume
only covered the not-yet-open registry path.

Reuse refreshOpenBotChat (the #99393 reclaim mechanism) on the fronted
branch so forceResume re-pulls the latest transcript. Regression test
pins the behavior: fronting an open Bot Chat now requests the canonical
registry open.
2026-09-03 01:10:16 -07:00
Edder Talmor
37fd6eea97 fix(desktop): toast action is a real button, not a hairline text link
The notification action (`NotificationItem`) rendered as
`variant="textStrong" size="xs"` — an 11px underlined muted-grey text link
with a ~44x20px hit target. On the data-training confirm toast raised by
`surfaceModelSwitchConfirm` / `confirmModelWarning` (e.g. picking
`muse-spark-1.2-contributor`) it read as a footnote, not the one action
the toast exists for, and users reported not being able to "press to
accept".

Promote it to the SDK's `default` variant at `size="sm"`: a filled
primary button, larger hit target, obvious affordance. No new styles.

Salvaged from #96562 (toast half only). Refs #96563.
2026-09-03 00:58:46 -07:00
Teknium
d0b7cec0b8 fix(prompt): Muse Spark gets tool-use enforcement + execution guidance on defaults (#96550)
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.

Co-authored-by: Edder Talmor <talmoredder@gmail.com>
2026-09-03 00:58:32 -07:00
Teknium
0b96eaf06b chore(contributors): map sosxradar@gmail.com -> GTHell 2026-09-03 00:58:16 -07:00
Teknium
4359af7705 fix(models_dev): alias opencode-free to the Zen "opencode" catalog; pin Muse Spark 1M invariant
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).

Missing alias identified by @Steve-prog001 in #101905.

Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
2026-09-03 00:58:16 -07:00
Steve-prog001
779aecb62b fix(context): resolve commandcode models via live /models
commandcode (api.commandcode.ai) exposes authoritative
context_length via /models (muse-spark 1M, etc.) but as a
known provider it skipped the custom-endpoint probe at step 2
and has no models.dev entry, so every model fell through to the
256K DEFAULT_FALLBACK. Add a provider-aware branch mirroring
gmi/nous to resolve via _resolve_endpoint_context_length.

Fixes GOAT docs vs status-bar mismatch: muse-spark 1M was shown
as 256K.
2026-09-03 00:58:16 -07:00
GTHell
bb8f4afa46 fix(context): add muse-spark 1M fallback (zen/GO SG showed 256k)
Muse Spark 1.2 family (api.meta.ai) ships 1M context (models.dev
opencode/muse-spark-1.2 = 1048576, meta/muse-spark-1.2 = 1048576).

Zen/GO SG /v1/models only returns id (no limit.context), and
models.dev lookup via opencode was missing a hardcoded fallback, so
get_model_context_length fell back to DEFAULT_FALLBACK_CONTEXT=256k.
Banner showed Context: 256,000 for both zen and router-sg lanes.

Add longest-prefix entries 'muse-spark' and 'muse' = 1_048_576 so
all variants (1.1, 1.2, contributor, contributor-free) resolve to 1M
without network.
2026-09-03 00:58:16 -07:00
Teknium
fa53e4fedd chore(contributors): map csreyes92@gmail.com -> csreyes (salvage #93073) 2026-09-03 00:57:55 -07:00
mr-r0b0t
cfa7e72c9e fix(models): correct contributor guard, 1M context, docs for muse-spark-1.3
- model_data_policy_guard: name the triggering -contributor model instead
  of hardcoded 1.2; per-version verified price tables (1.3 standard
  $1.25/$4.25 via OpenRouter live metadata; cached figures 1.2-only)
- model_metadata: muse-spark-1.3 + muse-spark family at 1048576 (OpenRouter
  verified 2026-09-02) with pre-catalog stale-cache keys so 256K-fallback
  sessions self-heal
- docs: contributor-tier notes cover 1.2 + 1.3
- tests: 1.3 guard regression, muse stale-cache guard, live-catalog mirror
  gains 1.3-contributor-free (confirmed on live relay)

143 tests pass (guard, selection guards, opencode catalog, model_metadata);
ruff clean.
2026-09-03 00:57:55 -07:00
mr-r0b0t
7e60d0c042 feat(models): add Meta Muse Spark 1.3 family to picker
Add meta/muse-spark-1.3 and meta/muse-spark-1.3-contributor to the
OpenRouter curated list, the meta-ai provider fallback, the
opencode-zen / opencode-free / opencode-go floors, the setup-wizard
shortlist, and regenerate the hosted model catalog.
2026-09-03 00:57:55 -07:00
Christian Reyes
7f2aa70add feat(models): add Muse Spark contributor to OpenRouter 2026-09-03 00:57:55 -07:00
Teknium
d6bb94a1fb fix(meta-ai): declare supports_vision_tool_messages=False — Muse Spark 400s on image tool results
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.

Fixes #101668
Refs #47742
2026-09-03 00:57:40 -07:00
liuhao1024
b462989a68 fix(vision): honor supports_vision_tool_messages=False in tool-result media gates
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.

Both gates now go through one _profile_rejects_tool_media() veto.

Refs #89981

(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)
2026-09-03 00:57:40 -07:00
Teknium
75934b89c8 Merge branch 'simp/r2-tools-b' (tools wave 2 net-deletion: approval/delegate/skills_hub/terminal/computer_use) into simp/integration3
Conflicts (union): computer_use/{cua_backend,cua_backend_daemon,doctor}.py — kept r2-tools-b's
_run_quiet/_run_driver/_cli_text helpers and the integration3 stdin=DEVNULL fix (06fb390212).
2026-09-03 00:44:12 -07:00
Teknium
6b393f70eb Merge branch 'simp/r2-tools-b-delegate2' (tools wave 2: approval/delegate/skills_hub/terminal/computer_use) into simp/integration3 2026-09-03 00:41:31 -07:00
kshitijk4poor
b9dc790332 fix(inventory): reword pending-entitlement warning to avoid Windows footgun false positive
The naive line scanner's open() regex matched the human-readable phrase
"next picker open (or refresh)." inside the warning string, tripping the
Windows footgun gate. Reword to "next picker open or refresh."
2026-09-03 13:08:31 +05:30
kshitijk4poor
fb723084be fix(inventory): explain the locked Nous list while entitlement is pending
With the picker served from resident caches only, a cold Nous row renders
every model locked (free_tier_pending) until the background prewarm lands.
Surface why on the row's existing warning slot so the user isn't left with
an unexplained greyed-out list; never override an auth warning.
2026-09-03 13:08:31 +05:30
fangliquanflq
401ac7e1f8 fix(gateway): keep first model picker open responsive on cold pricing cache
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.

Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
2026-09-03 13:08:31 +05:30
Teknium
b6041240d1 test(desktop): trim Bot Chat pane-focus regression to the two invariant cases
Salvage follow-up to #101639 (@helix4u): keep the re-adopt-after-overlay and
miss-propagation cases, drop the hidden-pane and healthy-path controls.
2026-09-03 00:36:09 -07:00
Gille
b6d549d002 fix(desktop): restore missing Bot Chat panes before claiming focus 2026-09-03 00:36:09 -07:00
Teknium
61df062534 Merge branch 'simp/r3-33' (tip ac872a405d, +r3-33-F reconciliation) into simp/integration3 2026-09-03 00:29:22 -07:00
Teknium
0d336d055a Merge branch 'simp/r3-33' into simp/integration3
Conflicts in 10 tools/ files resolved as the union of both sides'
simplifications (integration3's earlier r3-33-* landings vs the parent
r3-33 reconciliation):
  tools/memory_tool.py, memory_tool_store.py, microsoft_graph_auth.py,
  microsoft_graph_client.py, plugin_guard.py, project_tools.py,
  registry.py, schema_sanitizer.py, self_repo_guard.py,
  session_search_tool.py

Tool-schema dump (/tmp/rf/tools_schema_dump.py) byte-identical to
113f04616b.
2026-09-03 00:28:36 -07:00
Teknium
ac872a405d Merge branch 'simp/r3-33-F' into simp/r3-33 2026-09-03 00:25:07 -07:00
hermes-seaeye[bot]
e629c900a8 fmt(js): npm run fix on merge (#101952)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-03 07:24:01 +00:00
Teknium
bd6cc48b94 fix(desktop): annotate the resolvable target for aliased LOCAL rows too
The contributor fix covers remote rows. The reporter's video shows the
sibling shape: with a remote gateway active, the LOCAL twin carries the
'default-this-device' alias, and message_agent's local resolver only
knows bare profile names / 'hermes'. Emit the same target annotation
whenever a local row's alias differs from its resolvable handle.
2026-09-03 00:18:31 -07:00
liuhao1024
2e542c92e6 fix(desktop): annotate canonical relay targets for remote @mentions
The Bot Mode mention middleware built message_agent targets from
botHandle(), which prefers a roster row's source-qualified UI alias
("default-vera"). Neither resolver accepts that form — the relay matches
canonical handle/profile (± @connection-id) and the local path a bare
profile name or "hermes" — so remote handoffs died with "No teammate
named" before enqueue.

Annotate the canonical form instead: profile@connection-id for remote
rows, canonical bare handle (default→hermes) for local ones. Pin the
profile@connection form on the relay side too, so the emitted target
stays inside the documented resolver contract (#97678).
2026-09-03 00:18:31 -07:00
Teknium
fee5b9efb8 Merge branch 'simp/r3-33-E' into simp/r3-33 2026-09-03 00:17:19 -07:00
Teknium
011bbf0cd6 Merge branch 'simp/r3-22-B2-small' into simp/integration3 2026-09-03 00:12:19 -07:00
Teknium
063ca416b3 Merge branch 'simp/r3-00-mix2' into simp/integration3 2026-09-03 00:12:00 -07:00
Teknium
4b0eaf0e54 Merge branch 'simp/r3-00-cli2' into simp/integration3 2026-09-03 00:11:59 -07:00
Teknium
6a1d96c346 Merge branch 'simp/r3-12-w2b' into simp/integration3 2026-09-03 00:11:57 -07:00
Teknium
ea6083dbfc Merge branch 'simp/r3-12-w2a' into simp/integration3 2026-09-03 00:11:56 -07:00
Teknium
5f9a4d3f02 Merge branch 'simp/r3-00-st2' into simp/integration3 2026-09-03 00:11:54 -07:00