Bundled by #83063 into the always-shipped set, but it serves only
multi-agent kanban campaigns — too niche for every install's skills
index (~20 tok/call for all users). Moved to the optional catalog
(installable via /skills search + hub, official source) and renamed so
the trigger is legible at a glance: 'merge-reconciler' read like a
generic git helper; 'agent-merge-conflict-arbiter' says who it is for.
- skills/autonomous-ai-agents/merge-reconciler -> optional-skills/autonomous-ai-agents/agent-merge-conflict-arbiter
- frontmatter name + description updated (description within the 60-char hardline its own contract test enforces)
- contract test moved/renamed, 9/9 green
- kanban docs (en + zh-Hans) repointed; zero merge-reconciler refs remain
The supporting-files block emitted every file twice per line
(relative -> absolute), duplicating an identical directory prefix
hundreds of times on reference-heavy skills. The absolute base is
already stated once in the [Skill directory: ...] header and the
footer example, so each line now carries only the relative path.
On hermes-agent-dev (462 supporting files) the activation message
drops from 42,145 to 29,042 o200k tokens (-13,103, -31%) — paid on
every session that preloads or invokes the skill.
* feat(compaction): always rebuild the system prompt at the commit boundary — keep-prompt now gated on byte equality of the LIVE builder output; plugin sections re-render with fail-open to last good bytes
* feat(clock): 'Conversation started' resolves through the session-lineage ROOT — a compacted/rotated session keeps its original birth date (Bot Mode forever-chats know when they were first born)
* test: retire old-contract pins — plugin sections re-render at invalidate (freeze stays restore-only), commit boundary always runs the live builder, byte-equal keep preserves object identity
Four gateway test files build module-level fake telegram.ext trees; a fake
missing InlineQueryHandler makes the adapter's lazy-install import fail
mid-suite, leaving ParseMode/ChatType bound to None for every subsequent
telegram test in the same worker (the CI-only MARKDOWN_V2/SUPERGROUP
NoneType cascade). Swept all fakes via grep, not just the one CI named.
The lazy-install placeholder test failing mid-run also left the adapter
module's ParseMode as None, cascading into test_telegram_thread_fallback
failures in the same CI process — all downstream of the same root.
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
Resume was attaching an unused store as {todos: [], revision: 0} and the desktop rejected tool.start updates that have no revision. A merge:true start after reconnect never patched the list until complete.
Skip unused empty snapshots. Apply unversioned updates without moving the watermark so a later todo.updated can still win.
Real-profile browsing (browser.use_real_profile) is meant to drive a COPY of
the user's profile headlessly in the background so they can keep working while
the agent acts on their behalf. Instead, on any host with a display it launched
the user's real browser binary HEADED, popping a window that grabbed focus on
every turn.
Root cause: the native launch (which bypasses agent-browser's own launcher to
avoid --use-mock-keychain dropping keychain-encrypted cookies) only added
--headless=new when Linux had no DISPLAY/WAYLAND_DISPLAY. On a normal desktop
the guard was false, so Chrome opened a visible window.
Fix: launch headless by default on every platform. Chrome's NEW headless mode
shares the profile's normal cookie store (unlike legacy --headless), and the
cookie drop we guard against comes from --use-mock-keychain, not from headless
— so real-profile auth still loads. Users who want to watch can opt in via the
existing browser.headed / AGENT_BROWSER_HEADED toggle (honored for real-profile
now, same as the rest of the browser stack); display-less hosts stay headless
regardless so the launch doesn't die at startup.
Live A/B on a real X seat: old argv mapped a Chrome window (focus steal),
--headless=new mapped zero windows while still exposing a working CDP port.
Updated the stale test that pinned 'never passes --headless' (a legacy-headless
premise) to positively assert the chrome launch is --headless=new by default.
Anchor env var name matches with \b to avoid matching legitimate
env vars that contain KEY/TOKEN/API as substrings (e.g.,
$TRILLIUM_ETAPI_URL). The patterns now require KEY/TOKEN/SECRET/PASSWORD
to appear at the END of the env var name, reducing false positives on
common API-usage documentation in SOUL.md while still catching actual
exfiltration attempts.
Fixes#63977
Follow-up reconciling the four cherry-picked contributor fixes with the
full-dir GitHubSource.fetch() that landed in #98246:
- Missing SKILL.md-linked support paths now warn and install without the
file at all three sources (GitHub full-dir, GitHub fallback, UrlSource) —
dangling links are prose over-matches or repo-only dev tools, not install
blockers (#66760/#90081). A referenced path present in the tree as a
SYMLINK stays a hard rejection.
- Extension requirement dropped from the glob/placeholder filter:
references/LICENSE is a legitimate support file (82236's tests pin this).
Truncated prose placeholders (references/type-<name>.md -> 'type-') are
still rejected via the trailing-separator shape.
- percent-quoted Contents-API path (82236) merged with revision pinning
(96336) in _fetch_file_bytes.
- Fixture typo fix: four cherry-picked test strings used '\---' where
'\n---' was meant (DeprecationWarning + frontmatter never parsed).
Validation: 167/167 across tests/tools/{skills_hub,skills_guard,
skill_bundle_provenance} + tests/hermes_cli/test_skills_hub.py; live GitHub
fetches (impeccable 163 files rev-pinned; anthropics frontend-design).
Addresses the review on #96336:
- Every GitHub byte fetch in an install (SKILL.md included) now carries
the resolved tree's SHA as ?ref=, closing the pre-existing TOCTOU
where /contents floated to the default-branch HEAD and bytes could
come from a newer revision than the tree the paths were validated
against. The tree is resolved first (idempotent + cached) so the pin
covers the whole install.
- Same-dir link targets are canonicalized before validation: query/
fragment stripped via urlsplit, percent-decoded, leading ./ removed —
the same normalization the support-dir branch applies.
- A case-variant link to skill.md never ships as a bundle entry, and a
case-folded collision among accepted siblings (A.md + a.md) drops the
pair — both would overwrite/collide on case-insensitive filesystems.
_referenced_support_paths only kept links whose first path segment was
one of the five support directories (references/templates/scripts/
assets/examples), so a SKILL.md linking same-directory siblings —
mattpocock/skills' domain-modeling links ./CONTEXT-FORMAT.md and
ADR-FORMAT.md — installed 'successfully' with those files silently
omitted: the bundle came out semantically incomplete with unresolved
links.
A second pass now collects same-directory markdown-link targets
(](./FILE.ext) or ](FILE.ext)) that name an extension-bearing file,
carry no internal slash, and are not SKILL.md itself; a leading '..'
is rejected fail-closed exactly like the support-dir traversal branch,
external URLs/anchors/mailto/site-absolute targets are left to their own
resolution, and every accepted name still runs the bundle path
validator. Unlinked siblings remain excluded — the fetch-minimization
contract is unchanged; only files the document explicitly links ship.
A SKILL.md that mentions a support-path token which doesn't exist in the
repo (a prose glob like `references/type-*.md`, a placeholder truncated
by the regex, or a repo-only dev script) previously made GitHubSource
and UrlSource fetch fail with 'Could not fetch ... from any source',
even though every real file was present.
- _referenced_support_paths: skip glob/placeholder tokens and bare
prefix matches that can't name an actual file.
- GitHubSource/UrlSource fetch: warn and continue when a referenced
support file is missing or unfetchable, bundling what exists, instead
of aborting the whole install. Non-regular tree entries (e.g. git
symlinks) are still rejected.
Regression tests for both syntax filtering and the fetch loops.
Covers the reviewer-requested end-to-end path: a referenced support file
that 404s is skipped, and the URL skill still installs through quarantine,
scan, install, and lock provenance with only the reachable files on disk
and in the lock file.
UrlSource.fetch() aborted the whole SKILL.md URL install if any single
referenced support file (references/templates/scripts/assets) 404'd or
was otherwise unreachable, even when the SKILL.md itself and most
companion files were fine. Skip the missing file with a warning and
keep the ones that were fetched successfully.
Fixes#66760
The create-path coercion (salvaged from #78928) left update_job and the
cronjob tool's update handler comparing/storing raw strings: repeat=
'forever' via update raised TypeError in the tool path and stored the raw
string via update_job, breaking the next mark_job_run ('str' has no
.get). Extract normalize_repeat_value as the shared chokepoint (shape
from #77366 by @andrexibiza, with garbage-rejection semantics) and route
create_job, update_job, and the tool update handler through it.
Completed counters are preserved across repeat updates.
Class: #66824#64520#7142#71987#95706, update half of #77366.
The snapshot dirs were 0700 but every file inside landed umask-wide:
shutil.copy2 preserves Chrome's own 0644 profile-file modes and
sqlite3.connect creates the online-backup destinations as plain umask
files — so the copied Cookies / Login Data / Web Data (the user's live
session credentials) sat 0644. The 0700 parents contain it by default,
but the documented HERMES_HOME_MODE traversal hatch makes group/world-
readable children a real exposure.
snapshot_real_profile now reconciles every file (0600) and nested dir
(0700) inside the snapshot through the house helpers (_secure_file /
_secure_dir — managed-mode and container carve-outs included) at the
end of every pass, so snapshots written by older builds heal on their
next launch. Best-effort, never blocks a launch.
Tests: owner-only walk under umask 022 (fails on the pre-fix code —
sabotage-verified) + heal-on-refresh for a pre-existing 0644 Cookies.
The issue's other two findings are already fixed on main: mock-keychain
flags eliminated by the direct native-binary launch (#98249, salvage of
#96763); the 'Device not configured' TTY failure is superseded by the
same launch-path rework.
Sibling tests outside the salvaged PR's files created one-shots via bare
'30m', which is now a recurring interval per the corrected contract.
Fixtures whose assertions depend on kind='once' (run-claim clearing,
terminal-record rearm, web-server completed-snapshot) now use 'in 30m';
sites indifferent to kind keep the bare form.
Contract bug (2026-08-04): the cronjob tool schema documents '30m' as
'(every 30 minutes)' — recurring — but parse_schedule returned
kind='once' for bare durations, silently creating a one-shot job for a
recurring request (agent passed '30m' for 'every 30 min', job ran once
and died). Bare durations ('30m','2h','1d') now parse as recurring
intervals matching the documented contract; explicit one-shot by
duration is 'in 30m'/'in 2h' (fires once that far from now). ISO
timestamps stay one-shot.
Also fixes the sibling repeat-coercion class (#66824/#64520/#7142):
repeat='forever'/'once'/'N' strings now coerce in create_job instead of
raising "'<=' not supported between instances of 'str' and 'int'".
Tool description rewritten to teach the corrected contract and steer
relative requests to 'in Nm' (no more hand-computed ISO timestamps).
Supersedes the doc-only direction of #53739 while keeping its goal
(relative one-shots must be expressible) via the 'in X' form.
Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
The gateway's session-backed MCP OAuth flow (mcp.servers.oauth.start) binds
its browser-callback listener on the BACKEND machine's 127.0.0.1. When the
Desktop app connects to a remote backend (SSH/Tailscale), the user's browser
resolves that loopback to the user's machine, the redirect dies, and every
OAuth catalog server (ClickUp, Hospitable, ...) fails in-app with no working
path — the exact topology from the 'MCP Recurring erros' support thread.
Fix mirrors the Desktop's native gateway login (native-oauth-login.ts):
- gateway: mcp.servers.oauth.start accepts client_redirect_uri (loopback-only,
RFC 8252-style validation); when supplied no gateway listener is bound and
the OAuth redirect_uri pins to the client's listener.
- gateway: new mcp.servers.oauth.callback RPC relays the client-captured
code/state into the flow; state verification stays in
DashboardOAuthFlow.deliver_callback (constant-time compare, replay-safe).
- desktop: mcp-oauth-callback-ipc.ts hosts a one-shot 127.0.0.1 listener in
the main process (hermes:mcp-oauth:listen/wait/cancel via preload bridge).
- desktop: hermes-bots mcp-setup.tsx prefers the client listener for local
AND remote backends, falling back to the legacy gateway-listener flow on
older gateways (feature-detect via start rejection).
- docs: remote-host MCP OAuth section documents the automatic Desktop path.
Validation: 19 new gateway tests (validator allowlist, listener skip, relay
accept/reject/replay) — sabotage-verified; 5 new desktop tests against a real
ephemeral listener; E2E through the real session registry + flow bridge with
a stubbed provider probe; tsc electron+renderer builds clean.
hermes skills install impeccable (and the docs-page install button) now
installs the impeccable frontend-design skill as an official optional-skills
entry. The local optional-skills/creative/impeccable/ dir is a catalog STUB:
its frontmatter declares metadata.hermes.upstream (repo + path), and
OptionalSkillSource.fetch() pulls the real 163-file bundle live from
pbakaus/impeccable:.hermes/skills/impeccable — the Hermes-native bundle
upstream maintains and verifies. Nothing vendored, never stale.
New mechanism (generic, not impeccable-specific):
- OptionalSkillSource._upstream_pointer(): parses/validates the upstream
pointer (owner/name repo, clean relative path, traversal rejected).
- _fetch_from_upstream(): delegates to GitHubSource.fetch(), relabels the
bundle official/<rel> at trust 'trusted' (curated endorsement, but
third-party content — dangerous scan verdicts still block).
- The live-repo fallback path redirects stubs the same way, so stale local
checkouts behave identically.
Three real gaps this surfaced, all fixed:
- GitHubSource.fetch() only downloaded SKILL.md plus paths linked from a
canonical support dir (references/, scripts/, ...). Impeccable keeps its
playbooks under reference/ (singular) and links scripts only from
reference files, so fetch shipped 1 of 163 files. fetch() now downloads
the full skill directory via the git tree (same approach as the
optional-skills live fetch), still rejecting symlinks/hidden/unsafe paths
and still failing on a missing SKILL.md-linked references/ path.
- The five env_exfil_* scanner patterns flagged loopback requests as
critical exfiltration: impeccable's live mode polls
http://localhost:PORT/status?token=TOKEN and scored two CRITICALs.
Scheme-anchored loopback exemption added; evil.com/?u=localhost decoys
still fire (10-case regex matrix in tests).
- unified_search() truncated to limit before ranking, so official catalog
entries got crowded out by skills.sh mirrors and bare-name installs
stalled on an ambiguity table. Results now stable-sort by trust rank
before the cut, and _resolve_short_name prefers a sole official exact
match over community mirrors.
Also fixes pre-existing test pollution: TestInstallPathSafety's fixture
monkeypatched the PEP 562 dynamic SKILLS_DIR, permanently shadowing dynamic
resolution and breaking the served_repo E2E tests in any combined run
(reproducible on main).
Validation: live E2E do_install("impeccable") against real GitHub —
resolves to official/creative/impeccable, verdict SAFE, 163 files on disk,
skill loads, /impeccable slash command registers. 128/128 targeted tests;
full-dir fetch test sabotage-verified. Docs: optional-skills catalog row,
generated skill page, sidebar.
Widens _natural_every_to_cron to consume comma/'and'-separated weekday
lists ('Monday, Wednesday at 9am' -> '0 9 * * 1,3') and applies the same
helper to schedules without the 'every' prefix, matching the exact forms
the Desktop dialog advertises in the #51975 repro.
`parse_schedule()` detected cron expressions with a digit-only field pattern
(`^[\d\*\-,/]+$`), so any field using named months or weekdays — `MON`, `JAN`,
and common ranges/lists like `MON-FRI` or `MON,WED,FRI` — failed detection and
fell through to a confusing "Invalid schedule" error, even though croniter
supports them and they're standard cron.
Allow letters in the field pattern so these route to croniter for validation.
Truly-invalid expressions (`0 9 * * FUNDAY`, `99 9 * * MON`) are still rejected
there with a clear "Invalid cron expression" message; duration/interval/ISO
parsing is unchanged.
Adds tests for named weekdays/months (incl. ranges and lists) and that an
invalid named field is still rejected.
parse_schedule's "every " branch passed everything after the prefix
straight to parse_duration(), so documented natural-language schedules
like "every monday 9am" and "every day at 9am" (AGENTS.md, SKILL.md,
cron docs) were rejected with "Invalid duration". Convert weekday and
daily/weekday/weekend phrases to cron expressions before the duration
fallback; "every 30m"/"every 2h" interval parsing is unchanged.
The bundled plan skill's auto-generated slash command fell off the capped
Telegram/Discord command menus for most installs (skills are the only tier
trimmed at the platform caps, alphabetically — 'plan' sat past the cutoff at
index 57 of 82 bundled skills). Converting it to a first-class CommandDef
gives it a guaranteed core-tier menu slot on every platform.
- agent/plan_prompt.py: build_plan_prompt() — plan-mode rules + authoring
craft distilled from the retired skill; prompt-injection pattern like
/learn and /init (no engine, no model-tool footprint, cache-safe).
- CLI: _handle_plan_command mixin handler (pending-input injection).
- Gateway: /plan branch rewrites event.text and falls through (role
alternation preserved).
- TUI: command.dispatch branch ('plan' was already in
_PENDING_INPUT_COMMANDS).
- Removed skills/software-development/plan/ + docs pages (EN + zh-Hans),
catalog rows, sidebar entry, related_skills references.
- PROTECTED_BUILTIN_SKILLS is now empty (mechanism kept); dependent
curator/usage tests moved to monkeypatched sentinels.
Salvages #67292 by @webtecnica (credit: first /plan command submission,
issue #67264); reworked from inline planning prompt to the prompt-injection
pattern with workspace-saved plans. Closes#67264, closes#36821 (empty
/plan infers task from conversation context).
Completes the #90953 salvage on post-#98237 main:
- New _merge_request_overrides helper defines the precedence contract:
explicit delegation.request_overrides merges OVER runtime/parent-derived
overrides — explicit top-level keys win; extra_body is deep-merged one
level so runtime extra_body keys survive unless redefined. Inputs are
copy.deepcopy'd so transport-side mutation can't leak into config or the
provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
provider-alongside-base_url runtime overrides instead of being a separate
return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
(previously only when override_provider was set), enabling the inherit
branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
proofs, explicit-over-runtime precedence on the provider-alongside-base_url
path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
The direct base_url branch of _resolve_delegation_credentials returned no
request_overrides key, so a direct OpenRouter delegation (provider=custom,
base_url=openrouter.ai/api/v1) could not pass routing hints to its children.
The named-provider branch already forwards runtime request_overrides; this
gives the direct branch the same contract, honouring delegation.request_overrides
from config (dict → forwarded, anything else → None).
Primary use: extra_body.provider = {"sort": "throughput"} so delegation
children route to the fastest OpenRouter provider for their model, per the
fab-swarm throughput work (#901).
Follow-up hardening on the two cherry-picked contributor commits:
- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
with KEY-SCOPED removal — only keys the OLD provider's custom_providers
entry contributed (value unchanged since the init-time merge) are dropped.
Caller/profile-provided extra_body keys survive the swap, matching the
caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
request_overrides, so a post-switch transport recovery or fallback restore
reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
(test_provider_fallback.py), switch-then-recover / switch-then-restore
(test_primary_runtime_restore.py).
Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.
Fixes#75091
Port the PR #52432 regression (fast turn then normal turn on a cached
gateway agent) onto the current _run_agent harness: the original test's
host file context no longer exists on main after the TurnRunner
extraction, so the scenario is re-expressed with the existing
_CapturingAgent fixture. Asserts init-time custom-provider extra_body
survives both a /fast turn (service_tier layered on top) and the
following normal turn (only the stale fast-mode key drops).
Salvaged-from: #52432
Co-authored-by: Heng Cai <abtion@outlook.com>
Addresses the hermes-sweeper review on #53765. The in-place /model switch
helper (_apply_switched_provider_request_overrides) derived a custom
provider's extra_body by provider *name* only, while build-time matching in
agent_init._merge_custom_provider_extra_body matches by provider key, base_url,
AND model. So a different model selected at the same named endpoint could
inherit an extra_body configured for another model.
Reuse the shared agent_init._custom_provider_extra_body_for_agent matcher
(provider key + base_url + model), sourcing custom_providers from the
init-time agent._custom_providers cache (fresh-load fallback if absent). A
stale extra_body is always cleared when no entry matches; non-provider
overrides (service_tier / speed from /fast) are preserved.
Tests: add nonmatching-model and endpoint-mismatch regressions; update the
existing switch tests onto the model/base_url-aware matcher.
Follow-up to the previous commit (which fixed the default/fallback
provider path). A mid-session `/model` switch stores a per-session
override bundle in `_session_model_overrides` that omitted
`request_overrides`, and the two consumers
(`_resolve_session_agent_runtime` fast path and
`_apply_session_model_override`) only copied
provider/api_key/base_url/api_mode. So switching *to* a custom provider
via `/model` did not apply its `extra_body`.
- `ModelSwitchResult` gains a `request_overrides` field, derived for the
switched provider via `_get_named_custom_provider` /
`_custom_provider_request_overrides` (the same overrides
`resolve_runtime_provider` surfaces for the default path).
- Both `/model` override-storage sites in slash_commands.py persist it.
- Both consumers apply it; `_apply_session_model_override` also clears a
stale value when switching to a provider that has none.
Extends tests/gateway/test_turn_request_overrides.py (3 new cases).
A `custom_providers` entry can carry an `extra_body` (e.g.
`chat_template_kwargs` to toggle a local vLLM model's thinking).
`resolve_runtime_provider()` correctly surfaces it as `request_overrides`
on the resolved runtime dict, but the gateway never plumbed it through to
the per-turn agent:
- `_resolve_runtime_agent_kwargs()` rebuilt the runtime dict from a fixed
key whitelist that omitted `request_overrides`.
- `_resolve_turn_agent_config()` rebuilt `runtime` from the same whitelist
and set `route["request_overrides"]` solely from `/fast` service-tier
overrides (`{}` otherwise).
- The per-turn `agent.request_overrides = turn_route.get(...)` assignment
then clobbered the value `_merge_custom_provider_extra_body()` applied at
agent construction.
Net: on the gateway, a custom provider's configured `extra_body` never
reached the model -- only `/fast` overrides survived. The CLI/TUI path
(which does not go through `_resolve_turn_agent_config`) and the auxiliary
client (which sends `extra_body` directly) were unaffected.
Fix: carry `request_overrides` through the runtime resolvers
(`_resolve_runtime_agent_kwargs`, `_try_resolve_fallback_provider`) and
merge the provider overrides into the per-turn route, layering any `/fast`
service-tier overrides on top (top-level keys, no collision with
`extra_body`).
Adds tests/gateway/test_turn_request_overrides.py.
Known follow-up: the mid-session `/model`-switch override path
(`_session_model_overrides` / `ModelSwitchResult`) does not yet carry
`request_overrides`.
Third in the series. The gateway rebuild path (previous two commits)
carries a custom provider's `request_overrides` (`extra_body`, e.g.
`chat_template_kwargs`) into the agent, but the *in-place* live switch used
by the TUI dashboard and the CLI — `agent.switch_model()` ->
`agent_runtime_helpers.switch_model()` — swapped
model/provider/base_url/api_key without ever updating `request_overrides`.
So a `/model` switch to a thinking-enabled custom provider in the TUI/CLI
kept the previous provider's `extra_body`.
`switch_model()` now re-derives the switched-to provider's
`request_overrides` (via `_get_named_custom_provider`) and applies it in
place, preserving non-provider overrides (`service_tier`/`speed` from
`/fast`). Logic factored into `_apply_switched_provider_request_overrides`
for testability.
Adds tests/agent/test_switch_model_request_overrides.py.
Carry provider-derived request_overrides through runtime resolution,
fallback projection, session /model state, restart rehydration, and
turn-route merge so named custom providers keep extra_body and related
overrides.
Salvaged from #56876 (cron half only; the delegation half is superseded
by #98237). run_job's ephemeral AIAgent constructor passed api_key /
base_url / provider / api_mode from the resolved runtime but dropped
request_overrides, so cron jobs on custom providers silently lost
extra_body / extra_headers request settings.
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
avg_latency_s, avg_tps (reads the same per-call deque history from
agent/conversation_loop.py; keys omitted when no data — Codex
app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
segments (breakpoints 96/104/110 cols, lowest priority — they shed
first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
filters TUI segments too: cache_hit, latency, tps, duration,
compressions, bg_tasks, bg_subagents, voice, battery, title,
context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
Profiles can now grant narrowly scoped tools to the background review runtime whitelist while unrelated tools remain denied. Document the configuration and cover it with a real-config regression test.
Agent: codex
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions
The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
Four fixes for real-profile browsing (browser.use_real_profile), found and
verified end-to-end on macOS with a live Chrome:
1. _copy_auth_file: sqlite3.connect('file:...?mode=ro') on a live Chrome
auth DB can block indefinitely inside lock negotiation - the busy
timeout never fires, so the 'fail fast' path hangs the launch forever.
Try immutable=1 first (reads instantly, correct for a committed
snapshot of a file another process owns); mode=ro stays as fallback.
2. Launch shape: agent-browser's own launch injects --use-mock-keychain /
--password-store=basic / --headless=new. On macOS the mock keychain
makes Chrome treat every keychain-encrypted cookie as undecryptable
and drop it - the copied profile launches signed out (~3 anonymous
cookies instead of the full jar). Launch the user's real browser
binary directly on the copy (no mock-keychain switches), wait for
DevToolsActivePort, then attach agent-browser via --cdp.
3. Snapshot copy: Local State was copied verbatim, still naming the
SOURCE profile (last_used='Profile 2', info_cache listing several)
while the copy only contains Default. Chrome opens the missing profile
dir and starts signed out. Normalize the copy's Local State to
Default-only.
4. CDP resolution: the agent-browser daemon may report the endpoint of a
browser IT spawned (throwaway temp profile) instead of the real
browser we launched on the copy. Trust the port our browser wrote to
DevToolsActivePort.
Also adds browser.real_profile_pin (optional): pin which source Chromium
profile dir is snapshotted instead of following profile.last_used - on a
machine with a work profile and a personal one, last-used roulette can
silently give the agent the wrong identity. A pin naming a missing dir
fails closed (signed out) rather than falling back to last_used.
Tests: 4 new pin tests + 3 launch tests reshaped to the direct-launch
contract (Popen the real binary, agent-browser attaches). 77 passing.
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
baseline resets, title badge gating
Add a ◎ XX% indicator to the CLI status bar showing the prompt cache
hit rate (cache_read / prompt_tokens). This helps users monitor how
effectively their provider's prompt caching is working.
Features:
- Color-coded: green (≥70%), yellow (40-70%), red (<40%)
- Adaptive precision: integer on narrow terminals, one decimal on wide
- Only shown when cache data is available (provider supports it)
- Compatible with OpenAI, Anthropic, DeepSeek, xiaomi, and other
providers that return prompt_tokens_details.cached_tokens
Tests: 6 new test cases, 43/43 passing
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.
Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.
When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.
total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.
Closes#41909