Commit Graph

25576 Commits

Author SHA1 Message Date
Andrex Ibiza, MBA
82a702bf67 feat(update): bake authoritative image provenance into the Docker image
Cherry-picked core of #92545: the image build writes a versioned,
non-secret marker (/etc/hermes/image-provenance.json) outside both the
bind-mountable checkout and the HERMES_HOME volume, and
hermes_cli/image_provenance.py reads it fail-closed — absence means
'not image-managed', any present-but-malformed marker still means
image-managed (an integrity defect is never permission to mutate the
image in place).

(#91277 Phase 3; salvaged from #92545 by @andrexibiza — marker bake +
reader only, the scoped carve-out.)
2026-08-26 11:41:04 -07:00
Jack Lau
2552579912 fix(tui_gateway): ask before queueing a guarded model picked mid-turn (#91043)
* fix(tui_gateway): ask before queueing a guarded model picked mid-turn

config.set model on a running session cannot swap the agent in place, so it
stashes the pick in session["pending_model_switch"] and applies it at the
next turn start. That branch answered confirm_required=False without ever
running the selection guards.

A client that implements the confirm round-trip was therefore told no
consent was needed and never prompted. One turn later
_apply_pending_model_switch ran the guards with the stashed (unconfirmed)
flag, saw the warning, and dropped the switch by design. The model reverted
with no confirm ever offered, because the only moment a round-trip was
possible had already passed.

Evaluate the guards before stashing, where the client still has a live
response to turn into a prompt. Nothing is queued for an unconfirmed
guarded pick, so the session is left exactly as it was and the re-send
carrying confirm_expensive_model queues it for real. The apply-time check
stays as the backstop for guards that can only decide after resolution.

The data-policy guard keys on the model id alone, which is all this branch
can see before resolution. The cost guard returns None when pricing is
unknown and its models.dev lookup is allow_network=False, so calling it
early can only under-fire and never blocks the RPC thread.

* test(tui_gateway): pin provider forwarding, name the canonical confirm field

Two review follow-ups, no behavior change.

_pending_switch_selection_warning forwards `provider=provider or None`, but
nothing asserted it: a guarded model id fires the data-policy guard on the
model alone, so the existing tests passed with `provider` dropped entirely.
Record the kwargs instead. Dropping the argument fails the first test;
removing the `or None` normalization fails the second.

The confirm responses carry `warning` and `confirm_message` with identical
text, which reads like an accident. Name which one clients should read
(`confirm_message`; `warning` is the pre-confirm-era alias that
_apply_pending_model_switch already treats as a fallback) so the two do not
drift apart later.

Both raised by @Enough1122 in review.
2026-08-26 13:38:53 -05:00
xxxigm
b519ce29ad fix(desktop): keep attachment close and code copy icons visible (#95611)
* fix(desktop): keep attachment close and code copy icons visible

Hover-only opacity-0 hid the composer remove control and code-block copy button, so they stayed clickable but invisible on Windows and other no-hover surfaces.

* test(desktop): pin attachment close and code copy visibility at rest

The remove chip and code-block copy control must stay in the tree without a hover class, so Windows and no-hover surfaces cannot hide them again.
2026-08-26 13:38:19 -05:00
Gille
d0fdbfd655 fix(desktop): show repo-root-only sessions in project drill-in (#94552) 2026-08-26 13:37:43 -05:00
Pedro Fontana
b0dbf72f76 Merge pull request #91716 from NousResearch/fix/scale-to-zero-no-pointless-quiesce
fix(gateway): don't quiesce for a suspend this platform can't schedule
2026-08-26 14:58:38 -03:00
Nikita Barkov
2e80d7fa05 fix(slack): keep the resolved proxy on bolt's per-request client
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.

The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.

Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.

Co-authored-by: Junie <junie@jetbrains.com>
2026-08-26 10:35:05 -07:00
Nikita Barkov
bac960e23d fix(slack): stop injecting thread roots as reply context 2026-08-26 10:26:00 -07:00
Nikita Barkov
1f92c5d4ce feat(kanban): carry the review handoff summary into the wake turn
`completed` already puts the worker's summary inside the synthetic wake
turn, so the woken creator sees what was done. `review_requested` did
not: the summary rode the passive ping only, and the wake turn said just
"handed off for review", forcing the woken reviewer to re-read the board
(and losing the PR link the worker had already written).

Reuse the same first-line handoff the `completed` branch builds, so the
existing `gateway.kanban.wake.handoff` string renders it — no new locale
keys, no change to the passive message.
2026-08-26 10:25:33 -07:00
Nikita Barkov
7700d3a011 fix(kanban): wake the origin on review handoffs and triage escalations
`review_requested` and `block_loop_detected` are terminal event kinds that
hand a decision back to the origin subscriber, but neither was listed in the
gateway notifier's `_WAKE_KINDS`. A `notify+wake` subscription therefore got
the passive ping only and the origin agent never took a turn — so an agent
that delegated implementation work slept through the "ready for review"
handoff and through a task being routed to triage, while the equivalent
`blocked` event woke it.

Add both kinds to the wake set, add their status strings to the synthetic
wake message in every locale, and document which events wake.
2026-08-26 10:25:33 -07:00
Nikita Barkov
5538bd1f93 fix(slack): prevent duplicate rich-text message content
Slack sends an authored message twice: flat in `event.text` and structurally
in `event.blocks`. The blocks are rendered so quoted and forwarded content is
not lost, and whatever the render carries beyond the flat text is appended to
the message. That comparison had several ways to fail on the *same* sentence,
each of which showed the author their own words a second time:

1. HTML entities — the flat copy escapes `&`/`<`/`>` while `blocks[].link.url`
   stays raw, so any link with query parameters (every "Copy link" on a
   thread) mismatched.
2. Permalink unfurls — the live inbound path skips `is_msg_unfurl`
   attachments, thread/parent hydration did not, so the linked message's body
   was appended again.
3. The Block Kit dump — it serialized the authored `rich_text` alongside the
   UI blocks it exists for, and its allowlist drops `url`, so the sentence
   reappeared with every link removed.
4. Unknown inline elements — the renderer knew eight types and silently
   dropped the rest. A pasted message permalink arrives as `message_mention`,
   so the link vanished from the render and the sides stopped comparing equal.
5. `message_mention` without a url — `url` is optional on that element while
   `channel_id` and `message_ts` are not, so the element rendered as nothing
   and the sentence came back with a blank in the link's place.
6. `date` elements — `fallback` and `url` are both optional, and the flat
   `<!date^…>` form was never read down to what the rich text renders.
7. Labelled mentions — Slack may attach a label (`<@U…|name>`,
   `<#C…|general>`, `<!subteam^S…|@marketing>`, `<!here|@here>`) in the flat
   text while the blocks carry the bare id. The bot's own mention is one of
   these, and stripping only its bare form left it in the flat copy.
8. Autolink schemes — only `https` and `mailto` were matched, so a `tel:` link
   kept its angle brackets and mismatched too.

Unknown inline types are now read by their `url`/`text`/`fallback` so a type
Slack adds later still renders, and `team`, `color` and a fallback-less `date`
render into the flat form Slack sends. Every field is read as a string or
not at all: Block Kit carries text as an object in many places, and a
non-string one reaches the renderer's `str.join` and raises there, which
costs the whole message. `channel_id` and `message_ts` are the
permalink's own components, so a url-less `message_mention` renders the
permalink's tail; the workspace host and the thread query cannot be rebuilt
from the element, so a permalink on either side is reduced to that same tail.
Canonicalization is used for matching only -- the authored text still reaches
the agent verbatim, so a mistake here can cost an unrendered element, never an
altered or missing message.

An element carrying neither a url nor a label still renders as nothing, and a
message containing one is still appended twice. Suppressing such a render was
tried and is worse: an app message whose body lives only in the blocks
disappears, and a forwarded quote is dropped. Genuinely additional content --
quotes, lists, code blocks, attachments, interactive bot blocks -- is
unaffected throughout.

Tests cover both merge sites (live inbound and thread hydration) and the
negative cases.
2026-08-26 10:10:51 -07:00
Teknium
3a4c309cfb chore: map contributor email for salvage of #88470 2026-08-26 10:10:11 -07:00
Alexander Prendota
613164dadc fix(acp): key the ACP runtime exclusions on the scheme, not on one vendor
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.

Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
2026-08-26 10:10:11 -07:00
Alexander Prendota
37fd61d13b fix(background-review): skip the fork when the provider can't emit tool calls
The review fork's entire job is to emit `memory` / `skill_manage` tool calls,
and by default it inherits the parent's live runtime. A provider that IS an
autonomous agent reaches Hermes through a client shim; if that shim cannot
carry Hermes tool calls back, the fork is a guaranteed no-op that still pays
for a full agent spawn — a whole CLI process, sometimes a JVM — on every
review cadence.

A client declares `SUPPORTS_HERMES_TOOL_CALLS = False` and the fork is
skipped with a warning naming the `auxiliary.background_review.{provider,model}`
override that routes the review to a normal model instead. Anything that says
nothing is assumed capable, so ordinary providers are untouched.

The check runs before the thread-scoped silence so the warning is not
swallowed, and only resolves the review runtime once the cheap capability
test has already failed, so the normal path does not resolve it twice.
2026-08-26 10:10:11 -07:00
Alexander Prendota
07200e9cd6 feat(agent): fold an agent-as-provider's own tool work back into the turn
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.

Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:

- the self-improvement loop, which distils memories and skills by replaying
  `messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
  Hermes tool iterations, of which there are none.

So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.

The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
2026-08-26 10:10:11 -07:00
Alexander Prendota
083c5920e5 refactor(acp): share one OpenAI bridge between the ACP clients
ACP has no OpenAI `tools`/`tool_calls` channel: a prompt is text and a
response is text plus the agent's own tool notifications. Hermes' agentic
surface — memory, todo, skill_manage — is dispatched from OpenAI-shaped
tool_calls, so on an ACP provider it only works if the schemas travel into
the prompt as text and the calls are parsed back out of the response text.

copilot-acp already carried that bridge as private module-level helpers.
Lift it verbatim into `agent/acp_openai_bridge.py` so every ACP client
shares one implementation of the wire contract instead of re-deriving it —
`agent/claude_code_acp_client.py` (#81375) is currently a third copy of the
same four functions, and each copy is a place the `<tool_call>` contract can
drift.

copilot-acp is migrated onto it as the in-tree consumer and loses 176 lines
of duplication; its prompt shape is unchanged, which the new tests pin.

Two things the shared version adds over the copy:

- `render_tool_bridge_sections(..., allowlist=)`. A CLI with no tools of its
  own forwards Hermes' whole toolset (copilot, unchanged: no allowlist). A
  CLI that *is* an autonomous agent must forward only Hermes' agent-level
  tools — re-offering the overlapping read/edit/execute ones makes Hermes
  re-run work the agent already finished.
- `StreamChunks`, a list subclass that keeps response-level attributes.
  Hermes reads provider extras off the object returned by
  `chat.completions.create`; the old plain-list return silently dropped them
  whenever a caller asked for `stream=True`.
2026-08-26 10:10:11 -07:00
Teknium
03537d69dc feat(gateway): updaters pause gateways over the control socket instead of tree-killing them (#92091 step 2)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.

- gateway/run.py: pause-for-update verb handler registered on the
  existing control server; marshals onto the loop thread, ACKs with
  {pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
  no-answer (older gateway / no socket), so every caller keeps the
  legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
  per mapped profile gateway before the drain wait; positive ACKs extend
  the wait to the gateway's own declared drain budget (+ teardown grace)
  so a mid-turn gateway isn't force-killed at the end of a too-short
  local default. Marker write + force-kill ladder retained verbatim as
  the fallback.

Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
2026-08-26 09:59:17 -07:00
unsupportedpastels
c4f376c19a fix(config): block generic Copilot ACP controls 2026-08-26 09:54:31 -07:00
unsupportedpastels
5425ba14f2 fix(config): harden MCP env policy on Windows 2026-08-26 09:54:31 -07:00
unsupportedpastels
08cf4fea5d fix(mcp): restrict catalog environment writes 2026-08-26 09:54:31 -07:00
Teknium
1a5547c5c5 fix(desktop): guard setPrimaryGatewayConnectionId against non-primary active scopes
Hardening follow-up to #95396 (#95628): ignore primary connection-id writes
while the active key is a composite secondary scope, so future
presentation-layer writes cannot relabel the primary socket and poison
new-session routing. Regression test is sabotage-proven (fails with the
guard reverted).
2026-08-26 09:52:27 -07:00
David Bartoš
10746a53ba fix(desktop): preserve primary gateway identity across source switches 2026-08-26 09:52:27 -07:00
Teknium
4e8e5f84bc chore: contributor mappings for #95633/#95652 salvage 2026-08-26 09:51:55 -07:00
Thawatchai
7eb11cabfc fix(desktop): route clarify responses through session owner 2026-08-26 09:51:55 -07:00
Luke Roberts
6fdaef6a9f fix(desktop): resolve session owners from the cron and messaging sidebar slices
The owner ladder's row rung (tile route -> hint -> session row) only searched
$sessions (recents). Cron- and messaging-sourced sessions are fetched as their
own sidebar slices ($cronSessions / $messagingSessions), so a scheduler-minted
cron session had no tile, no hint, and no row the rung could see. On a
registry-topology install every session-scoped RPC for such a session then
failed closed:

  Session owner could not be resolved for "xxxx" (approval.respond): no owner
  route, hint, connection-tagged row or profile probe named the backend ...

which made command approvals raised inside a cron chat impossible to answer
(the approval bar errors on every choice), even on a single-local-backend
install whose cron row - with its profile stamp - was already loaded for the
sidebar's cron section. The approval-bar path (requestForOwnedSession) has no
async REST-probe rung, so nothing downstream recovered.

Fix: ownerLookupSessionRows() returns the union of the three source-scoped
slices for OWNER lookups, keeping $sessions' array identity in the common
recents-only case so reference-keyed memo caches still hit. All four row-rung
call sites (session-rpc-dispatcher, knownOwnerForSession, tile delegate, tile
actions) now search it.

Repro: any cron session that raises approval.request (open the cron chat,
reply with a prompt that triggers a gated command, click Run).
2026-08-26 09:51:55 -07:00
Teknium
778feee37a chore: eslint --fix on salvaged e2e spec 2026-08-26 09:47:23 -07:00
686f6c61
bdbb41b62e fix: adapt salvaged staleness-probe tests to the kickoff option-object signature
createCanonicalChat gained { kickoff } on main (#95326) after #91227 was
filed with a bare positional openingStillCurrent; the salvage merges both
into one options object.
2026-08-26 09:47:23 -07:00
686f6c61
0ecbf1ce91 fix(desktop): snapshot Close All pane ids before persist-close
closeSessionTile can rewrite the layout tree. Iterating the live
group array then skips tiles. Copy the list first.
2026-08-26 09:47:23 -07:00
686f6c61
f1755cc1c5 fix(desktop): persist Bot Mode Close All session tiles
Close All only dismissed layout-tree panes. Bot tiles live in the
shared __bots_workspace__ bucket, so clicking a bot or swapping
profiles rehydrated the closed tabs. Persist-close those tiles
before dismissing the rest of the strip.
2026-08-26 09:47:23 -07:00
funky-xamarin
9ea7a37cc1 fix(desktop): restore live group after bot open failure 2026-08-26 09:47:23 -07:00
funky-xamarin
fbd149035b test(desktop): verify group-to-local-bot handoff in Electron 2026-08-26 09:47:23 -07:00
funky-xamarin
35ee27b1c6 test(desktop): add RED group-to-local-bot E2E 2026-08-26 09:47:23 -07:00
funky-xamarin
7c255dbbbf test(desktop): focus group-to-local-bot handoff on close-callback RED
Rewrite the handoff proof to ≤150 lines on the hermes-bots VM seam so
base fails because the registered group closer is not invoked, not a
missing helper marker. Cover BotRow/Active Now local close-before-open,
remote no-dismiss, and no-group/old-host safety.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin
2fa39fac8c test(desktop): make group-to-local-bot handoff proof behavioral
Exercise real BotRow and Active Now open handlers so close-before-open
and remote stay-put are load-bearing, not source-regex false greens.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin
f72786f764 fix(desktop): close group main tab when opening a local bot chat
Local BotRow and Active Now opens now share a handoff that retires the
selected group workspace before canonical chat open, so the bot chat can
own the main surface after leaving a group tab.
2026-08-26 09:47:23 -07:00
pefontana
b2c493c173 Assert the skip branch ran in the no-quiesce watcher test
The three existing assertions are absence checks, so the test also
passed when the loop got no iteration inside the sleep window (with
interval=5.0 it passes without the gate ever executing). Checking
_scale_to_zero_no_suspend_logged proves the branch was taken.
2026-08-26 13:02:04 -03:00
pefontana
ae0e3bb12b Fix docstrings that still describe off-Fly dormancy
The watcher now skips the quiesce when self_suspend_available() is
False, but its docstring and the one on self_suspend_available() still
said the suspend step is skipped and dormancy still happens. Say what
happens instead: the gateway stays connected until the platform freezes
it.
2026-08-26 13:02:04 -03:00
pefontana
85816595d2 Merge remote-tracking branch 'origin/main' into fix/scale-to-zero-no-pointless-quiesce 2026-08-26 13:02:03 -03:00
Teknium
9aa7530f7b style(desktop): sort oauth-partition import (lint) 2026-08-26 08:40:07 -07:00
Teknium
697087f2eb fix(desktop): SIGKILL-escalate the owned SSH backend when it survives the graceful quit wait (#91668 remainder)
The #95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
2026-08-26 08:40:07 -07:00
Teknium
8497edb2ac test(desktop): pin the activation-epoch guard against mid-handshake profile switches (#92434 close-candidate)
Reproduces the reported Bot↔Default switch shape at the gateway.ts
activation seam: a switch-back that lands while the outgoing switch's
WS handshake is still pending keeps the route, the late-completing dial
neither steals the foreground nor breaks its socket, and re-activating
the bot works without an app restart. The guard (activation epochs +
open-socket-publish, landed via #89622/#92265/#81094) already prevents
the reported permanent break; this pins it so it cannot regress.
2026-08-26 08:40:07 -07:00
Teknium
1e9a12a71f fix(desktop): stop spawning loopback serve children when the registry primary is remote (#91564, #90316)
'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
2026-08-26 08:40:07 -07:00
Teknium
62e2d6e1e4 fix(desktop): quarantine malformed connections.json entries per-entry instead of dropping or nuking the registry (#94246)
- normalizeRegistry now preserves every malformed entry (unknown kind,
  url-less remote/cloud, host-less ssh, mangled non-object items, and
  any entry whose normalization throws) under a capped 'quarantined'
  key that survives write cycles — healthy entries keep loading and
  user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
  connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
  a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
  surfaces reason+label summaries (never raw entries/token envelopes).
2026-08-26 08:40:07 -07:00
Teknium
31250da505 fix(desktop): key cookie-auth session partitions on connection identity, not auth mode (#92183)
Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
2026-08-26 08:40:07 -07:00
Teknium
30749ed9dc test(tools): guard _ever_connected set in reconnect regression mock so it bites on pre-fix code (#94671 hardening) 2026-08-26 08:40:07 -07:00
chelsealong
8f517f5ca6 chore: address AI-review nits on _ever_connected fix
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
2026-08-26 08:40:07 -07:00
chelsealong
e8dc0af5b1 fix(tests): set _ever_connected in reconnect-scenario test mocks
These pre-existing tests fake a successful first connect by calling
only _ready.set(), which is what the real code did before this PR.
Now that run() gates the initial-vs-reconnect branch on the new sticky
_ever_connected flag instead, their later simulated reconnect failures
were misclassified as never-connected and hit the 3-attempt ladder,
failing test_reconnect_counter_resets_after_successful_session,
test_parked_server_self_probes_and_revives, and
test_retry_attempts_log_debug_transitions_warn in CI. Set the flag
alongside _ready.set() to mirror the real success sites, same as the
new test added in tools/mcp_tool.py's own PR.
2026-08-26 08:40:07 -07:00
chelsealong
c7673f322b fix(tools): stop treating a post-registration reconnect drop as an initial-connect failure
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.

Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.

Fixes #94654
2026-08-26 08:40:07 -07:00
Zeus-Deus
fd565c80e9 feat(desktop): fleet profile rail — every registered gateway's agents on one strip
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.

- Every registered gateway's profiles sit on the one strip, in registry order
  (This device first, then by label), each group headed by that gateway's
  kind glyph. The active gateway's squares are unchanged; the others are
  "at rest" (dimmed) with tooltips/accessible names qualified by machine
  (`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
  the statusbar switcher, landing on that exact (gateway, profile):
  `selectConnection(id, { profile })`. The spinner sits on the clicked
  square; the previous source stays painted until the target answers.
  Groups keep their slots whichever gateway is active, so a square never
  moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
  SOUL.md / Delete, executed on the owning gateway (renameProfile,
  getProfileSoul and updateProfileSoul accept the same scope deleteProfile
  already had); the delete confirmation names the machine. The legacy
  per-profile "Connect to a remote host…" item is hidden on multi-gateway
  setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
  two registrations of one backend collapse to one group; past thirteen
  squares across the fleet the strip condenses into a menu sectioned by
  gateway. Roster is fetched on mount / focus / registry change only — no
  periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.

Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.

Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.

Docs: multi-connection-desktop.md describes the fleet rail.

Refs #89304, #92384, #91047, #94724

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 08:39:45 -07:00
Teknium
dd0aae4173 fix(desktop): paint stored Bot Chat history immediately instead of stranding the wake on an unsatisfiable profile gate
Fixes #89843. On a shared-remote connection every profile is served
through the primary socket, so waitForFocusedSessionHydration's
profileMatches gate could never become true — a bot chat whose stored
transcript painted within seconds still burned the whole 20s hydration
budget and then stranded the pane with 'Timed out loading <bot>'s
session history'.

The wake now resolves paint-first: once the stored transcript is painted
on exactly the target session, the content is its own proof — the pane
opens immediately and a subtle 'Syncing…' badge (new $hydrationSyncProfile
atom + ChatSyncBadge) shows until the profile gate catches up in the
background. Fail-closed everywhere content is not its own proof: a
superseded/conflicting concurrent wake still rejects, and an
expected-empty chat still waits for the full runtime gate.
2026-08-26 08:39:45 -07:00
etzelvon
dbbed456c1 style(desktop): satisfy curly + padding-line lint for guarded switch
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
2026-08-26 08:39:45 -07:00