jev_cycles_report.py renders the cycles/freed/floor/end-state table from
jev_cycles.py outputs; README carries the exact commands, thresholds and
cost so the diminishing-returns and stuck/fallback findings can be
reproduced. Fallback records now count the real tool calls.
Closed-book arms scored the summary with its session_search pointer unused
(43% vs 79% on the same banks) and made external compactors look like wins.
Bare policy names stay available as an explicit opt-in floor.
Against summary + one session_search round-trip (78.9% @ 55K) Jev's default
arm (75.5% @ 115K) loses on recall and retains 2.1x the tokens; the earlier
closed-book comparison scored our summary with its recovery pointer unused.
Programmatic tool-result removal as the primary compaction leaves ~2x the
context billed every turn and tightens compaction cadence every cycle (each
one a prompt-cache break); the summary's one-time ~90% reduction keeps a
long conversation cache-warm. The recall gap points at summariser retention
of identifiers in assistant text, not at a cadence change.
jev_cycles.py compacts a lineage with Jev every time the estimate crosses a
threshold and records freed tokens, text floor and fitting stage per cycle.
Six runs show the floor (never-removed user/assistant text) is the binding
limit: freed-per-cycle decays to 0-8% and a 200K-window host is stuck after
0.42M tokens; >~500 tool calls between compactions cannot fit the 32K state.
Scorecard gains a decisive verdict section.
Adds an eval-only budget selector to the Jev arm (rank tool pairs by Jev
keep_result or by recency, keep until a token budget) so Jev's judgment can
be separated from the effect of keeping text verbatim, and records the
3-transcript scorecard: default Jev +32 pts recall at 2.1x retained tokens,
1/9 the compaction cost; at matched budget Jev ranking ties recency.
Adds a Python port of tamaratran/fast-jev-compaction as an eval policy
(`engine: jev`): TypeSafe's Jev decision model scores every tool call and
result over the whole history and stale ones are dropped or truncated, with
no summary and no rewriting of user/assistant text. Transport is OpenRouter's
Decisions API (~typesafe/jev-latest). A state that cannot fit the plugin's
25K-token ceiling is recorded as a fallback (the plugin's own behaviour)
rather than scored.
Every arm now reports what its compaction step cost (calls, tokens, USD —
Jev reports cost directly; summary calls are metered through the
compressor's call_llm binding and priced at OpenRouter list), and the run
writes the harness's own answer/judge token bill so the eval's spend is
visible. The question cache key includes --cap-tokens so different caps on
one transcript no longer share a bank.
Why: measure remaining tokens, compaction cost and recall accuracy of the
"decide, don't summarize" approach against current/lean on real Hermes
lineages before deciding whether any of it belongs in the compressor.
Eval probes wrote fixtures, receipts and evidence to hard-coded /tmp paths and
two live A/B tasks literally instructed the model to work in /tmp. They now
derive locations from tempfile.gettempdir() / os.tmpdir() (env overrides kept),
usage examples use relative output names, and the desktop e2e screenshot dirs,
perf scripts and the short-session repro fixture stop naming /tmp. Also adds
the explicit encoding= the windows-footgun check wants in the touched files.
* refactor(desktop): Capabilities is a module at /capabilities
The Capabilities page lived in `app/skills/`, was exported as `SkillsView`
and was routed at `/skills`, although Skills is only one of its four tabs.
The name sent every reader to the wrong place.
- `app/skills/` becomes `app/capabilities/`, with one folder per topic:
`skills/`, `plugins/`, `mcp/`, and `catalog/` for the browser that the
skills and plugins tabs share.
- `SkillsView` becomes `CapabilitiesView`, in the plugin SDK too. The
bundled bots plugin reads the new export name.
- The route, its id and its view become `/capabilities` and
`capabilities`. The keybind action becomes `nav.capabilities` and the
sidebar item id follows, with its i18n keys in every locale.
- The `hermes://open/...` link that the plugin notice fires follows.
No alias is kept for the old route. A remembered `/skills` route no longer
matches a page; the restore path already drops a route it cannot validate
and opens the last session.
No behaviour change otherwise.
* refactor(desktop): each Capabilities tab owns its list, detail and writes
`CapabilitiesView` was one 770-line function. It held the tab routing, the
whole Skills tab and the whole Tools tab, while the Plugins and MCP tabs
already lived in their own files.
- `capabilities/index.tsx` is the page shell: tab selection, the search
header, the profile and connection scope, the refresh hotkey, and one
table that maps a tab to its component (1190 lines down to 230).
- `skills/` and `toolsets/` each hold their tab, detail pane, data hooks and
helpers. `scope-selector.tsx` holds the scope hook and its selector.
`primitives.tsx` holds what two tabs share.
- The shell still fetches the two installed lists, because the tab pills
count them for the tab the user is not on. Query keys are unchanged;
`store/hub-actions.ts` imports the skills key instead of copying it.
- Every tab is keyed on the scope, so a profile or connection switch mounts
a fresh tab. This replaces the epoch counters and manual resets. One small
change follows: a return to a scope seen before selects the first row, not
the row selected last time.
hermes_cli/AGENTS.md: the update pipeline no longer runs pulled code in the pre-pull
interpreter; the receipt crosses the hand-off; purge/reload/isolate fixes are not to be
reintroduced. Stale comments in update_receipt / update_abort_recovery / update_cmd_maint
that explained the old shape are rewritten to say why the code is the way it is now.
evals/update_pipeline/post_swap_handoff_ab.sh: installs a real clone at a given SHA,
publishes a target commit whose new module-level import (a symbol added to a module the
old updater kept cached) and tagged completion line show which interpreter ran the tail,
runs `hermes update` against a disposable HOME with no npm and --no-gateway-restart, and
prints one VERDICT line. base (6005aa1f): exit 1, "Could not check config version", old
completion line. fix: exit 0, config check runs, "[completion printed by: pulled-code]",
one receipt carrying post_swap_pid.
Use hermes-tools for the probe's managed server configuration and tool-call target so it exercises the endpoint receiving worker overrides. Leave intentional user-defined hermes-mcp fixtures unchanged.
Scope-risk: narrow
Tested: Three complete isolated probe runs on macOS with Codex CLI 0.137.0; related regression selection 202 passed, 2 existing skips; Ruff and repository static checks
Not-tested: Full repository suite, native Linux/Windows, hosted CI, cloud-model turns
Extend the real-wire device fixture with a `multi_issuer` mode whose
protected-resource metadata lists an issuer-mismatching server before the
valid one, and run the production CLI login through it. Red on main
(`Authorization server metadata issuer mismatch`), green with the scan.
Ported from PR #112068.
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.
Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.
Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.
Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
Retain executable Electron evidence for the named-unowned and live-owner
paths; stale PATH fails on base and passes with the contributor fix.
Clarify that --in selects cwd rather than the profile database.
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema
One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.
MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.
Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.
`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.
Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.
`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.
The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.
* wip(desktop): connection.request store, resume restore, card routing for MCP targets
Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.
* fix(config): hermes update turns on the connections toolset for saved toolset lists
`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.
Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.
`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.
* refactor: anti-slop pass on the desktop slice; shorten added comments
Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.
* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed
The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.
session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.
lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.
vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.
* chore: drop __pycache__ files swept in by an over-broad git add
* fix(desktop): correlate the connection.request row with the model's tool call by reason
The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.
* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds
* fix(connections): settle reason derives from target state, never from the renderer
A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.
* fix(desktop): a pending connection card re-arms on resume and activate
The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.
* style: literal wording in added comments, docstrings and docs
* fix: shared gateway-event contract and config-schema category for the connection events
connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.
* style: import order (perfectionist) in the desktop and shared files this PR touches
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
The cookie gate (middleware._attempt_refresh) never coalesced concurrent requests
carrying the same stale refresh token, so a browser/desktop burst after access-token
expiry replayed a just-rotated RT into the provider's reuse detection and the whole
session was revoked (#55712). Both refresh paths also ran the synchronous provider
HTTP call on the ASGI event loop, which wedged /api/status behind a slow IdP.
Generalise Doud-FR's native-route single-flight (#71548) into
refresh_singleflight.refresh_session_coalesced and use it from both paths, each in
run_in_threadpool. The replay key is the RT alone: a burst that straddles a network
change must still coalesce, and whoever presents the RT already owns the session.
Middleware keeps its refresh_expired / provider_unreachable audit events via callbacks.
Live E2E (evals/dashboard_auth/refresh_singleflight_live_e2e.py, real uvicorn, stub
rotating IdP with reuse detection): base 1/4 requests survive each burst, 4 provider
calls, /api/status 2.7 s behind one refresh; fixed 4/4, 1 call, 10 ms.
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".
- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
-> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.
A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.
Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
Found by using the feature as a user (natural prompts, real sites, CLI PTY + native Electron), not by naming tools:
- browser_vault_save_login was registered but never offered: toolsets.py is a hand-maintained list. Added, with an
invariant test that every registered browser_vault_* tool is in the browser toolset.
- Vault tools were absent on the DEFAULT backend (Browser Use): the gate deferred to check_browser_requirements(),
which is False by design there. Gate = is_browser_use_cli_mode() or check_browser_requirements().
- The model typed a page-shown demo password with browser_type and offered to take one in chat: the vault rules
lived only on the vault tools. browser_type/browser_exec now carry a vault note when the vault tools are
present ("call browser_vault_list first … never type a password with this tool, never accept one in chat, even
if the page shows it"); the browser_exec login-wall line points at the vault instead of "ask the user".
- On Browser Use the saved item was bound to chrome://new-tab-page: the supervisor's default page session is the
daemon's blank tab. browser_vault_save_login now focuses the tab holding a password field before reading its
origin (focus_page("", accept=probe); about:/chrome: pages are never candidates). Live E2E leg added.
- Settings row: "identifier · Added <date>", origin omitted when it duplicates the label.
Live (real model): CLI on Browser Use — first visit prompts, signs in, saves; second visit fills silently; GitHub
decline (Enter or ESC) stops the agent, which refuses chat passwords. CLI on the built-in stack — same three
scenarios pass. Desktop native Electron — same three scenarios plus Settings list/remove pass. Password never in
a transcript, UI, or a file outside vault/.
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could
fill them: a dead surface holding real card numbers. browser_vault_fill now handles all
three kinds through the same origin-bound, supervisor-only, redacted path:
- classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens
(cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code,
country-name) with label/name heuristics as backup; a combined "MM/YY" control gets
exp_month+exp_year and suppresses the split fills; inspection now covers <select>
(country, state, expiry month) and the fill script picks an option by value or text.
- Every payment fill goes through request_elicitation_consent (gateway button round-trip
or CLI panel) before a byte is written; declined → payment_declined, headless sessions
are refused. A prompt injection that reaches a checkout can ask, not spend. Card values
join the redaction registry like passwords; the result lists targeted field tokens only.
- Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the
field) because a card without a bound origin is unfillable.
- The tool descriptions, docs and CLI copy drop "Phase 1 / login only".
Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes
nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the
country <select> untouched, and neither the card number nor the CVC appears in any result.
browser_vault_tool also: focuses the tab on the bound origin holding the right form before the
origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's
input tool" (rewritten per session by model_tools); _check_vault_available is registered
uncached because its answer is per profile (vault dir + config) and the probe is a file stat.
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.
Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
Adapt the command-position, bounded-candidate and launcher-option work from
embwl0x's #76063 to the current detection owner, then add executable basename
projection from Rohith Pariki's #104338. Parse raw quote state before applying
existing text normalization so quoted arguments do not become commands.
Cover shell payloads and literal env split-string carriers, retain path-specific
rules and whole-command globs, and document the supported normalization rather
than claiming an OS capability sandbox. Related: #104308, #76037, #76063,
#104338, #78521, #86711. No automatic closing directives: the older carriers
also contain broader case syntax and git-option work not included here.
Co-authored-by: embwl0x <embwl0x@users.noreply.github.com>
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Salvage #82561's idle watch consumer through the existing drain, retaining
notify-off suppression and retrying failed injections. Preserve #102647's
secondary-profile ownership for reconstructed sources, preflight, synthetic
completion turns, and raw process progress/final notices. Unavailable
transports defer rather than borrowing the primary bot or losing the event.
Keep retained transport provenance and alias-aware relay resolution. Add
three invariant tests and an offline real-process/Discord lifecycle eval.
Co-authored-by: trwpang <134544126+trwpang@users.noreply.github.com>
Co-authored-by: holny <holny@foxmail.com>
Bind settlement to the exact adapter task and event. Rejected or cancelled preparation refunds the existing claim; cancellation after the agent runner starts remains counted. Keep profile-scoped callback context and the manager replacement guard. A fire count is not outbound delivery proof. Drop departed routes rather than executing their stale schedule.
Slim accounting-invariant salvage of #93174; preserve current direct adapter dispatch instead of reviving its FIFO/inflight implementation. Prior art #92858.
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.
Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
Stress run with a real orchestrator subagent showed the per-task keys buried in results[] were not relayed
upward; the same prose lines now sit at the top of the sync delegate_task result. evals/subagent_process_handoff/
stress_handoff_live.py runs 12 real parent+child scenarios (handoff, orphan, unread, clean read, cap, exited refusal,
mixed fan-out, parent controlling the inherited process, sibling theft incl. adversarial, nested orchestrator,
5-way burst) and scores them against runtime accounting, never model prose.
Five test files and five evals scripts merged since this branch's base still import the
moved names via gateway.platforms.base. They work (base.py imports the names for its own
use) but the PR's invariant is that in-tree code imports from the defining module.
Based on the earliest batching proposal by BrunoBza (#104686) and JoaoMarcos44 structured correction (#104703). Slim redo into topical siblings and shared TUI poller/post-turn routing. Reported by Xipong (#104671). Local shell-to-loopback probes demonstrate 12 to 1 turn dispatches; suites await the campaign test lock.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Capture the durable parent session before output readers start, including CLI
and non-notifying spawns. Require that parent or its compression continuation
for retained reads; exact and prefix handles alone do not authorize access.
Live Linux terminal/one-shot linger/fresh-reader A/B: base loses results;
updated owner recovers both streams and exit 7. Unbound, foreign session,
delegated child, and other profile cannot recover the receipt. No notifications
are replayed. Full tools suite is queued behind the campaign test lock.
Follow-up to contributor salvage #104805 for #104511.
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.
Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.
Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.
Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>