Commit Graph

34668 Commits

Author SHA1 Message Date
kshitijk4poor
53cfca9766 fix(agent): run the outbound payload sanitizers on the summary request
Building the summary through `_build_api_kwargs` means it now carries
`tools`, but it still skipped the `_sanitize_structure_surrogates` /
`_force_ascii_payload` chokepoint the main loop applies after building
(turn_api_request). On cache-planned routes the main loop scrubs a deep
copy of the tools, so `agent.tools` can still hold a lone surrogate that
providers reject with a non-retryable 400 (#50959 class) — a request the
tool-less summary never used to make. Apply the same two passes here.
2026-09-14 20:35:28 +05:30
kshitijk4poor
488f2fc86d refactor(agent): drop the orphaned hand-rolled summary kwargs builder
`_chat_summary_attempt` now builds through `_build_api_kwargs`, leaving
`_iteration_summary_chat_kwargs` (56 lines mirroring the transport by hand)
and its only consumer `AIAgent._resolve_lmstudio_summary_reasoning_effort`
without a caller. The transport already owns every quirk they re-derived
(fixed temperature, LM Studio `reasoning_effort`, portal tags, provider
preferences, pareto router plugin), so there is nothing to keep in sync.
2026-09-14 20:35:28 +05:30
ywatanabe
419a050427 fix(agent): preserve tool cache in iteration summary 2026-09-14 20:35:28 +05:30
kshitijk4poor
2edc1249f9 chore: map ywatanabe@scitex.ai -> @ywatanabe1989 (PR #110480 salvage) 2026-09-14 20:35:28 +05:30
kshitijk4poor
9b199246e5 fix(desktop): cancelAndWait resolves only once the composed drain settles
Every caller tears the SSH connection down right after `await
cancelAndWait(scope)`, so resolving on this call's own barrier alone let a
connection-apply teardown overlap the pool stop's still-running afterStop
teardown for the same key. Wait for the composed barrier instead; chain the
barriers (drain promises never reject, so this equals allSettled). The test
now waits a macrotask and asserts that neither the apply nor the new
bootstrap runs before the first teardown completes.
2026-09-14 20:22:49 +05:30
kshitijk4poor
2aca48b416 fix(desktop): same-scope SSH drains compose instead of replacing the barrier
A pool stop blocked in teardownSshConnection() holds drains[scope]; a
concurrent connection apply calling cancelAndWait() for the same scope
replaced that barrier with its own and, having nothing to drain, cleared the
map entry as soon as it finished. start() then saw no drain and began a new
bootstrap while the first SSH teardown was still running.

cancelAndWait() now composes with the drain already in flight and the entry
is cleared only when the composed barrier settles, so start() waits for
every active teardown. Regression test reproduces the pool-stop / apply
race; it fails on the previous coordinator.

Reported by ehz0ah on #110025 (follow-up to #106935).
2026-09-14 20:22:49 +05:30
teknium1
4748caff76 fix(gateway): explicit tool_progress new/all keeps text progress in un-cardable Slack chats
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.

Test proven red on the salvaged head (adapter.sent == [] with `all`).
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
e9c037f65a docs(slack): clarify progress resolution and fallback lifetime 2026-09-14 07:46:51 -07:00
Victor Kyriazakos
a4e2a82a6d fix(gateway): preflight task-card destination before transport fallback 2026-09-14 07:46:51 -07:00
Victor Kyriazakos
dab7eebf47 fix(gateway): resolve progress mode and intent from the same source 2026-09-14 07:46:51 -07:00
Victor Kyriazakos
bc125d59d5 fix(gateway): null tool_progress inherits; name the task-card suppression latch
Review findings (Salt, adversarial pass on the two preceding commits):

- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
  counted as an explicit mode because the gate tested key presence, while
  the display resolver skips None and inherits. Null resolved to Slack's
  tier default `off` and disabled cards, which is the default-off trap the
  change exists to avoid. Explicit intent is now a non-None value (or the
  env bridge). Tests cover null at each level plus null-over-global-all;
  mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
  destinations, so the name no longer described the field. Renamed to
  `publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
  described the opt-in as independent of tool_progress. Rewritten: cards
  follow an operator-written off (including /verbose), null inherits, an
  un-threaded chat with the card lane active shows no tool progress, other
  native failures keep the editable fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
3412490ad1 fix(gateway): no text tool progress when a Slack chat cannot host a task card
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.

Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
ed25a40917 fix(gateway): explicit tool_progress off disables Slack task cards
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.

Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.

Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
2026-09-14 07:46:51 -07:00
kshitijk4poor
62e5f46656 test(streaming): one Anthropic event-stream fake for the three parse-error tests
Three inline context-manager classes shared the same __enter__/__exit__
boilerplate and differed only in the events yielded before the raise.
Also drop the unreachable 'or agent.base_url' fallback: every
anthropic_messages init path sets _anthropic_base_url, and the two
sibling call sites read it bare.
2026-09-14 20:01:19 +05:30
kshitijk4poor
32a0d2d2e1 refactor(agent): one home for the stream parse-error markers
The SDK's malformed-frame ValueError substrings were hand-copied into the
retry classifier and the error summary; a third SDK message would need
both remembered. The mid-tool retry branch no longer clears
partial_tool_names itself — _start_stream_attempt resets it for every
attempt. buffer_anthropic_tool_input's docstring states the trade: the
knob stays on the turn's kwargs and the changed tools block costs one
prompt-cache miss.
2026-09-14 20:01:19 +05:30
kshitijk4poor
982e504262 fix(anthropic): partial tool names are reset per stream attempt
A tool_use block name recorded by a stream attempt that died before any
visible text survived into the next attempt: only the deltas_were_sent
mid-tool branch cleared result["partial_tool_names"]. When the retry then
streamed plain text and dropped, the partial stub blamed the stale tool
("Stream stalled mid tool-call (old_tool)") and the stale name could make
a later attempt look mid-tool-call when deciding whether the drop is
retryable. Reset it in _start_stream_attempt alongside
provider_tool_in_flight, which already has attempt-local semantics.

Regression: two-attempt stream (tool_use start + parse error, then text +
drop) — stub content and emitted deltas carry no stale tool name.
2026-09-14 20:01:19 +05:30
kshitijk4poor
3d88259483 fix(anthropic): retry a malformed tool-JSON stream with buffered tool input
Follow-up to the cherry-picked fix: keep the classifier widening
("expected value at line" is now a transient stream parse error on the
main turn) but replace the messages.create() fallback with a retry on the
same stream wire.

Why not create(): the fallback ran outside Relay (lost request rewrites),
outside _handle_stream_error (could replace text already shown to the
user with a different generation), ticked no liveness events for the
whole buffered payload, and a bare identical retry still re-emits the same
malformed JSON.

Why not drop the fine-grained-tool-streaming beta (#108583/#109056): live
probe on claude-sonnet-4-5, ~500-line tool call - beta on: max inter-event
gap 1.6 s; beta off: 139 s zero-event gap while Anthropic buffers the
args, which the 180/240 s stale-stream detector kills on larger payloads
(the regression 80a899a8e2 fixed).

Instead, on a parse error the retry sets `eager_input_streaming: false`
on every tool for that request only (the SDK/API per-tool field overrides
the legacy beta header), so Anthropic returns buffered, server-validated
args while the happy path keeps fine-grained streaming. A tool_use that
started streaming is registered in partial_tool_names so the mid-tool
transient retry fires the same way it does on the chat_completions wire.
2026-09-14 20:01:19 +05:30
joaomarcos
dfb4caf4b7 fix(anthropic): recover malformed streamed tool JSON 2026-09-14 20:01:19 +05:30
kshitijk4poor
39ef1c1f3d refactor(desktop): keepalive redials with capped backoff; staleness derived, not counted
A dead tunnel was redialled every 2 s until the scope was torn down;
delays now double from the base to a 30 s cap and reset on open. The
generation counter duplicated what the entry map and entry.socket
already say (stop() removes the entry, connect() replaces the socket),
so abandonIfStale reads those instead. afterStop's unused entry
parameter is dropped.
2026-09-14 19:56:15 +05:30
kshitijk4poor
e24b344b6c fix(desktop): a failed SSH teardown in the pool-stop fence is logged, not thrown
afterStop awaited cancelAndWait(teardownSshConnection) with no catch; the
idle reaper calls stopPoolBackend un-awaited, so a teardown rejection
became an unhandled rejection on the main process. Log it through the SSH
log and let the fence release.
2026-09-14 19:56:15 +05:30
kshitijk4poor
fa67a4c9ea test(desktop): trim keepalive suite to two behaviour cases
Four registry tests overlapped: sibling independence is a two-line
assertion inside hold-until-stop, and the reject-missing-target case is
the setup half of the empty-scope primary case. Folding them keeps every
invariant covered (hold, no-reconnect-after-stop, sibling isolation,
''-is-a-real-key, baseUrl/token required) with fewer fixtures to keep in
sync. The pool-stop teardown fence is already covered by
pool-stop.test.ts::afterStop and ssh-bootstrap-coordinator.test.ts.
2026-09-14 19:56:15 +05:30
kshitijk4poor
109fbe097f refactor(desktop): drop the fail-open branches when WebSocket is not a function
Electron bundles a global WebSocket in the main process; a missing
constructor is a build/environment error, not a runtime state to route
around. The two silent `typeof WebSocketImpl !== 'function'` returns in
start()/connect() would turn that build error into an armed-looking
registry that never opens a socket, letting web_server_idle_exit retire
an owned SSH-isolated sibling with no log line — exactly the bug this
module exists to prevent. Let `new WebSocketImpl(url)` throw into the
existing catch, which logs and schedules a reconnect.
2026-09-14 19:56:15 +05:30
KoNit-K
5350f9021a test(desktop): trim SSH keepalive suite to four behaviour cases
Review banned source-reading main.ts greps and asked the remaining registry
cases down to ~4: hold-until-stop with no reconnect, empty-string primary,
reject missing target, and sibling isolation. Prettier the helper.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 19:56:15 +05:30
KoNit-K
134c5226af fix(desktop): hold the pool-stop fence through SSH teardown
Process-less SSH entries finish child exit immediately. Keep inFlight
and the bootstrap drain up until keepalive teardown completes, and
stop reading main.ts from tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 19:56:15 +05:30
KoNit-K
be69a8d043 fix(desktop): tear down SSH lifecycle when a pooled backend is stopped
Idle-reaper and LRU eviction only called stopPoolBackend, which left sshConnections and the keep-alive WS armed on process-less remote descriptors.
2026-09-14 19:56:15 +05:30
KoNit-K
e785128934 fix(desktop): hold SSH-isolated keepalive WS so idle-exit won't retire owned siblings
Desktop keeps live sockets on the profile-less backend while the --profile
sibling only sees short RPC sockets; idle-exit then retires the sibling and
the app restarts broadly. Arm a main-process /api/ws keep-alive for every
published sshConnections scope (including '') until teardown, without
treating spawn artifacts as liveness (#101626).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 19:56:15 +05:30
teknium1
d57c28a554 fix(tests): compression stall-fallback tests stop racing the 0.2s ceiling
Under a loaded runner the primary stall plus the fallback retry overran the
0.2s total ceiling, so the retry never started and attempts==1 failed
intermittently (seen once in a 40-worker tests/agent run). Idle stays 0.05s;
the ceiling moves to 2s per the >=2s wall-clock rule in AGENTS.md.
2026-09-14 07:21:18 -07:00
teknium1
d28938d3da fix(computer-use): screenshot dedup forgets its last frame at a compaction boundary
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
2026-09-14 07:21:18 -07:00
teknium1
31964ff4c6 fix(computer-use): dedup keyed by the scoped session, cleared on release
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
2026-09-14 07:21:18 -07:00
Teknium
682b973b32 feat(computer-use): stop resending unchanged screenshots
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.

Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
2026-09-14 07:21:18 -07:00
kshitijk4poor
66ddd5f83c fix(auth): only a Codex token refresh writes through to root
Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.

Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
2026-09-14 19:49:36 +05:30
kshitijk4poor
0ff20dc98a test: trim Codex write-through tests to two invariants on the real profile layout
The picked tests monkeypatched _auth_file_path/_global_auth_file_path
directly and leaned on a HOME override to dodge the pytest seat belt.
Isolate the way the rest of tests/hermes_cli does instead: Path.home ->
tmp_path and HERMES_HOME -> <root>/profiles/<name>, so the fixture drives
the same get_default_hermes_root() resolution production uses. Drop the
classic-mode test (no new behaviour: source == active store is the
pre-existing save path). Two invariants remain: root-borrowed refresh
lands in root (singleton + pool) with no profile shadow; profile-owned
grant stays local with root untouched.
2026-09-14 19:49:36 +05:30
liuhao1024
6bd29f26f6 fix(auth): write profile-refreshed Codex tokens through to the global store
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).

Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes #87503
2026-09-14 19:49:36 +05:30
kshitijk4poor
b8c474d7b5 refactor(cli): the launcher guard is linux_desktop_entry._needs_interpreter
The shebang read-and-classify wrapper was a byte-for-byte twin of
_needs_interpreter, which already delegates to _shebang_escapes_running_env
and carries its own edge-case tests; import the whole predicate instead of
half of it. Lazy import: linux_desktop_entry lazily imports resolve_hermes_bin.
2026-09-14 19:49:12 +05:30
kshitijk4poor
329257c060 fix(cli): keep venv-pinned console scripts exec-able on relaunch
The launcher guard rejected every python shebang, so a pip/uv console
script pinned to the running venv (#!<venv>/bin/python) was also
discarded in favour of `python -m hermes_cli.main`. Reuse
linux_desktop_entry._shebang_escapes_running_env, which already knows
that `env` shebangs escape and a shebang inside the running
interpreter's directory does not; only the escaping launcher loses the
venv. Also drops the second shebang classifier the fix had introduced.
2026-09-14 19:49:12 +05:30
frozen
5c2ddeb55a fix(cli): preserve venv across self-relaunch 2026-09-14 19:49:12 +05:30
kshitijk4poor
b9da06af4b chore: map contributor email for frozen 2026-09-14 19:49:12 +05:30
kshitijk4poor
7e7641561b refactor(gateway-windows): one death predicate, no second process scan
attested_gateway_died() re-ran find_gateway_pids() (current profile only)
although both callers had just proven the process table empty with
all_profiles=True, and it re-implemented check_start_attestation's
liveness rule. Callers now pass the liveness they hold (current_pids=[])
and both probes share _attested_dead(), so the consuming and read-only
twins cannot drift.
2026-09-14 19:47:09 +05:30
kshitijk4poor
da382a413e test(update): trim #109538 coverage to two invariant tests
Four new tests overlapped: plan-time and spawn-time attested-death overrides
both exercised the same predicate via monkeypatched lambdas. Collapse to:
- one end-to-end test using a real attestation marker in a tmp home: dead
  attested gateway keeps the plan under Desktop ownership, survives the
  spawn-time re-check, and the marker is consumed by the spawn;
- one probe test: no marker / null or non-list pids / non-dict / non-JSON all
  read False (fail closed), alive and clean-exit read False, read-only when
  it does read True.
Existing #76129 tests keep their attested_gateway_died=False pins unchanged.
2026-09-14 19:47:09 +05:30
kshitijk4poor
674e3f3cd1 fix(update): consume the start attestation once the cold-start spawns
attested_gateway_died() is deliberately read-only so the CLI-start warning
still fires, but that left the dead marker in place after the update path
acted on it. If the restored gateway never became ready (or died again
before the next CLI start consumed the marker), the same stale crash marker
would re-authorize another cold start against Desktop ownership on the
next update. Clear it via the existing _clear_start_attestation() path
right after _spawn_detached() succeeds - the marker has done its job at
that point; a new one is written once the spawn is confirmed ready.
2026-09-14 19:47:09 +05:30
kshitijk4poor
6a26556e1a fix(gateway-windows): fail closed on a null or non-list pids attestation field
_attested_pids_from now returns [] when "pids" is missing, null, or any
non-list value. The previous `data.get("pids", [])` only guarded a missing
key: `{"pids": null}` raised TypeError on iteration and a scalar would too.
attested_gateway_died() drives a Desktop-ownership override for the update
cold-start, so a malformed marker must never be able to authorize a spawn.
2026-09-14 19:47:09 +05:30
ennheng
830c8f443d fix(update): keep the Windows cold-start plan for a dead attested gateway
A Desktop self-update hand-off exits the app before the updater runs and can
kill the messaging gateway in those same seconds (#109538), so the updater's
discovery finds no live PID while the one-shot start attestation still
vouches for the dead one. Both Desktop-ownership checks then read "nothing
running" as "nothing to restore" and the bot stayed down until a manual
start.

Consult the attestation non-destructively before Desktop-owned lifecycle
suppresses a cold-start: a vouched-for PID gone without a clean ledger exit
keeps the plan and is restored; no attested death preserves the #76129 skip
unchanged.
2026-09-14 19:47:09 +05:30
kshitijk4poor
5d810f318b chore: map contributor email for ennheng 2026-09-14 19:47:09 +05:30
joaomarcos
6bc0e9e6df fix(agent): same-model review fork keeps the parent's affinity header and Portal conversation root (#109964)
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:

- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
  carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
  key falls back to the conversation root); the fork now publishes exactly the same instead
  of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
  value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
  `_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
  parent's PHYSICAL id, so after a compression rotation the review's Portal
  `conversation=` tag fragmented usage attribution across one logical conversation.

Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
2026-09-14 06:55:54 -07:00
salch-cred
a4b620f17c fix(agent): same-model review fork inherits the parent's resolved cache scope (#109964)
build_cache_parity_fork gives the same-model fork the parent's session_id,
cached system prompt, tools[] and session_start — but with
_persist_disabled=True and _session_db=None, BOTH cache-identity resolvers
diverged from the parent on their own: declared_conversation_scope failed
closed on _persist_disabled, and the lineage walk skipped on the missing
DB. The fork's affinity header (set_affinity_scope) and body
prompt_cache_key (cache_scope_id on the OpenAI-wire transports) therefore
keyed a different bucket than the gateway parent, costing one cold
~full-context request per review. Not gateway-only: any parent whose
lineage root != current physical id diverges too (teknium1's triage table).

Fix, per the triage's suggested direction: on the not-routed branch only,
the fork stamps _inherited_cache_scope = resolve_prompt_cache_scope_safe
(parent) — the parent's ALREADY-RESOLVED scope, no DB access from the fork,
persistence fully detached. Both declared_conversation_scope and
resolve_prompt_cache_scope return the inherited scope first when set, so
the header path and the body path are fixed together (fixing only one
leaves the other divergent — Vivamisu's header/body split observation).
Routed (different-model) forks, /branch children, delegate/tool children
and fresh sessions set nothing; the fail-closed default stands untouched.
/btw shares build_cache_parity_fork and gets the repair for free.
2026-09-14 06:55:54 -07:00
teknium1
1782bf79c8 test(desktop): contract-skew tests read REQUIRED_BACKEND_CONTRACT instead of a frozen 6 2026-09-14 06:54:32 -07:00
teknium1
d3a44784b1 chore(desktop): backend contract v7 — blocking prompts are JSON-RPC server->client requests
A renderer built after d9834a3e86 listens for srq- request frames; a v6 backend
still emits <kind>.request notifications, so every approval/clarify card would
silently never render. The skew toast now points the user at the backend update.
2026-09-14 06:54:32 -07:00
teknium1
274fd56dca fix(state): WAL lock guard follows the handle's lifecycle
Three gaps in the #110544 guard, all reported in its review and reproduced:

- A writer reopened by _reopen_after_close_locked (teardown/worker race,
  #94736) came back with no guard: the next stray close + foreign close
  deleted its WAL again.
- _try_wal_checkpoint refreshed the guard outside self._lock; landing after
  close() it pinned an OFD lock with no connection behind it, so a foreign
  `PRAGMA journal_mode=DELETE` saw `database is locked` forever.
- Refcounts keyed on (fd, inode) treated a recycled fd number as a surviving
  lock: A+B live, close A, C reuses A's fd, close B left C recorded as guarded
  while a foreign EXCLUSIVE succeeded.

The guard now counts handles per inode, re-locks every matching descriptor on
each hold (OFD re-lock is idempotent), and unlocks on the last handle only;
the reopen path holds it; the checkpoint refresh runs under self._lock and
skips a closed handle. The macOS holder scan folds case so a case-only alias
of the sidecar path on APFS still matches.
2026-09-14 06:54:07 -07:00
teknium
743140cd82 feat(gemini): send full JSON Schema tool parameters via parametersJsonSchema
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).

Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.

Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
2026-09-14 06:44:54 -07:00
teknium1
49c6d4a9e0 test(contracts): tests mirror tui_gateway/; the runtime-artifact spoof test asserts the new 4000
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
2026-09-14 06:12:19 -07:00