scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".
- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
-> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.
A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.
Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
Found by using the feature as a user (natural prompts, real sites, CLI PTY + native Electron), not by naming tools:
- browser_vault_save_login was registered but never offered: toolsets.py is a hand-maintained list. Added, with an
invariant test that every registered browser_vault_* tool is in the browser toolset.
- Vault tools were absent on the DEFAULT backend (Browser Use): the gate deferred to check_browser_requirements(),
which is False by design there. Gate = is_browser_use_cli_mode() or check_browser_requirements().
- The model typed a page-shown demo password with browser_type and offered to take one in chat: the vault rules
lived only on the vault tools. browser_type/browser_exec now carry a vault note when the vault tools are
present ("call browser_vault_list first … never type a password with this tool, never accept one in chat, even
if the page shows it"); the browser_exec login-wall line points at the vault instead of "ask the user".
- On Browser Use the saved item was bound to chrome://new-tab-page: the supervisor's default page session is the
daemon's blank tab. browser_vault_save_login now focuses the tab holding a password field before reading its
origin (focus_page("", accept=probe); about:/chrome: pages are never candidates). Live E2E leg added.
- Settings row: "identifier · Added <date>", origin omitted when it duplicates the label.
Live (real model): CLI on Browser Use — first visit prompts, signs in, saves; second visit fills silently; GitHub
decline (Enter or ESC) stops the agent, which refuses chat passwords. CLI on the built-in stack — same three
scenarios pass. Desktop native Electron — same three scenarios plus Settings list/remove pass. Password never in
a transcript, UI, or a file outside vault/.
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could
fill them: a dead surface holding real card numbers. browser_vault_fill now handles all
three kinds through the same origin-bound, supervisor-only, redacted path:
- classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens
(cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code,
country-name) with label/name heuristics as backup; a combined "MM/YY" control gets
exp_month+exp_year and suppresses the split fills; inspection now covers <select>
(country, state, expiry month) and the fill script picks an option by value or text.
- Every payment fill goes through request_elicitation_consent (gateway button round-trip
or CLI panel) before a byte is written; declined → payment_declined, headless sessions
are refused. A prompt injection that reaches a checkout can ask, not spend. Card values
join the redaction registry like passwords; the result lists targeted field tokens only.
- Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the
field) because a card without a bound origin is unfillable.
- The tool descriptions, docs and CLI copy drop "Phase 1 / login only".
Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes
nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the
country <select> untouched, and neither the card number nor the CVC appears in any result.
browser_vault_tool also: focuses the tab on the bound origin holding the right form before the
origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's
input tool" (rewritten per session by model_tools); _check_vault_available is registered
uncached because its answer is per profile (vault dir + config) and the probe is a file stat.
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.
Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
Adapt the command-position, bounded-candidate and launcher-option work from
embwl0x's #76063 to the current detection owner, then add executable basename
projection from Rohith Pariki's #104338. Parse raw quote state before applying
existing text normalization so quoted arguments do not become commands.
Cover shell payloads and literal env split-string carriers, retain path-specific
rules and whole-command globs, and document the supported normalization rather
than claiming an OS capability sandbox. Related: #104308, #76037, #76063,
#104338, #78521, #86711. No automatic closing directives: the older carriers
also contain broader case syntax and git-option work not included here.
Co-authored-by: embwl0x <embwl0x@users.noreply.github.com>
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Salvage #82561's idle watch consumer through the existing drain, retaining
notify-off suppression and retrying failed injections. Preserve #102647's
secondary-profile ownership for reconstructed sources, preflight, synthetic
completion turns, and raw process progress/final notices. Unavailable
transports defer rather than borrowing the primary bot or losing the event.
Keep retained transport provenance and alias-aware relay resolution. Add
three invariant tests and an offline real-process/Discord lifecycle eval.
Co-authored-by: trwpang <134544126+trwpang@users.noreply.github.com>
Co-authored-by: holny <holny@foxmail.com>
Bind settlement to the exact adapter task and event. Rejected or cancelled preparation refunds the existing claim; cancellation after the agent runner starts remains counted. Keep profile-scoped callback context and the manager replacement guard. A fire count is not outbound delivery proof. Drop departed routes rather than executing their stale schedule.
Slim accounting-invariant salvage of #93174; preserve current direct adapter dispatch instead of reviving its FIFO/inflight implementation. Prior art #92858.
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.
Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
Stress run with a real orchestrator subagent showed the per-task keys buried in results[] were not relayed
upward; the same prose lines now sit at the top of the sync delegate_task result. evals/subagent_process_handoff/
stress_handoff_live.py runs 12 real parent+child scenarios (handoff, orphan, unread, clean read, cap, exited refusal,
mixed fan-out, parent controlling the inherited process, sibling theft incl. adversarial, nested orchestrator,
5-way burst) and scores them against runtime accounting, never model prose.
Five test files and five evals scripts merged since this branch's base still import the
moved names via gateway.platforms.base. They work (base.py imports the names for its own
use) but the PR's invariant is that in-tree code imports from the defining module.
Based on the earliest batching proposal by BrunoBza (#104686) and JoaoMarcos44 structured correction (#104703). Slim redo into topical siblings and shared TUI poller/post-turn routing. Reported by Xipong (#104671). Local shell-to-loopback probes demonstrate 12 to 1 turn dispatches; suites await the campaign test lock.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Capture the durable parent session before output readers start, including CLI
and non-notifying spawns. Require that parent or its compression continuation
for retained reads; exact and prefix handles alone do not authorize access.
Live Linux terminal/one-shot linger/fresh-reader A/B: base loses results;
updated owner recovers both streams and exit 7. Unbound, foreign session,
delegated child, and other profile cannot recover the receipt. No notifications
are replayed. Full tools suite is queued behind the campaign test lock.
Follow-up to contributor salvage #104805 for #104511.
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.
Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.
Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.
Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Add explicit RFC 8628 device login and oauth.flow selection while keeping
browser PKCE and the SDK runtime refresh path. Reuse issuer/resource
validation, configured client authentication and profile-scoped storage.
Only persist an approved, validated grant; never echo endpoint error bodies.
Slim redo of #104752 by @wjorgensen, replacing duplicate HTTP/storage
wrappers with the existing SDK and two real-wire invariant tests.
Refs #104742
Co-authored-by: Wes Hermes <weshermes@Wess-Mac-mini.localdomain>
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Main now snapshots systemd unit listings before stopping old processes
and passes them into _restart_systemd_gateway_units_best_effort; the
catch-up budget test and eval call the new two-argument shape while
still asserting the unit's stop+start budget reaches the client timeout.
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.
Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.
Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698
Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904