Commit Graph

5514 Commits

Author SHA1 Message Date
brooklyn!
accb82032c feat(desktop): let each chat hide its composer status stack 2026-09-21 12:35:58 -05:00
brooklyn!
2a6debe5eb fix(desktop): leave Space activation to focused controls 2026-09-21 12:35:58 -05:00
teknium1
25c74f731c fix(desktop): unscopable mutations keep a process-scoped backend; stamp shared-primary event provenance
Review round 2 on the host-backend collapse. Two defects the collapse
introduced, both fixed on top of #118275's server-side `profile` plumbing
(which this PR now requires: it must merge first).

1. The collapse ate the invariant "a route the server cannot profile-scope
   must keep a process-scoped backend". Routing case 6 sent EVERY local
   request at the shared primary, so a mutating handler that reads no
   `profile` executed against the primary's HERMES_HOME. Restored as a
   mechanical gate, not a route list: `unscopableMutatingRequest()` is true
   only when the method is unsafe AND `localPrimaryRequestScope()` returns
   null, and `sharesHostBackend()` refuses to collapse that request. The day
   a handler learns `profile`, it joins the scoped table and the gate stops
   seeing it.

   `ensureBackend()` now re-resolves with the ORIGINATING request, so its own
   route decision matches the caller's `resolveProfileApiRequest` one instead
   of silently collapsing a request routing deliberately kept pooled, and the
   spawn guard is told when a pooled spawn is the legitimate outcome.

   After #118275 the residual unscopable-mutating set is the host-path
   families (`POST /api/files/upload`, `DELETE /api/files/managed`) plus any
   unlisted mutating method (`POST /api/skills`). Session writes are scopable
   by `body.profile`, so they share the backend with their path left alone.

2. Events from non-primary local profiles arrived unowned. `sharedPrimaryRoute`
   returns the primary gateway without `createSecondary`, so nothing ran
   `stampSecondaryProfileOwner`, and use-gateway-boot captured the source
   profile ONCE in a `useEffect(..., [])` — every later profile's events were
   stamped with the BOOT-time profile. The profile is now read PER EVENT and
   the owner marker is applied whenever the active descriptor is
   `sharedPrimary`, so `runtimeSessionOwner()` resolves again and the live
   sessions/cron sync stops degrading to slow polling. The renderer is the
   right place: `_event_frame` carries no `profile`, so a server echo would be
   a cross-surface wire-contract change.

   Session PATCHes (rename/pin/archive/mark-read) now always name the owning
   profile in the body — falling back to the ACTIVE profile when the caller
   knows none — because the handler resolves its state.db from `body.profile`
   alone and no per-profile process remains to stand in for it.

Also: the routing-table docstring entries 4-6 and the `hermes:api` comment now
describe the behaviour they gate, and the spawn guard reads the same
`profileRouteOptions()` the router does so the two cannot drift (it was
dropping the per-profile SSH-override term).

(cherry picked from commit 1e2354fb15866cc24d98e8e552a59240efe97a25)
2026-09-21 10:07:03 -07:00
teknium1
58eb248d08 refactor(desktop): collapse the per-profile backend pool onto one host backend
Multiplex-only says exactly ONE `hermes serve` runs per host and multiplexes
every profile. Desktop already ATTACHes its primary backend to a running one
(25d88ad0c4); the backend POOL was the other producer of `hermes serve`
children -- one per profile, up to `maxBackends` (default 3) on top of the
primary.

A local profile now resolves onto the same host backend and carries its own
`profile` on the wire. The server already binds per connection-of-work rather
than per process: `session.create {profile}` stores `profile_home` on the
session and every turn re-binds it, and a sessionless RPC takes the explicit
`profile` argument (`tui_gateway/server.py::@_profile_scoped`).

- `host-backend-singleton.ts`: the pure collapse decision plus
  `assertNoSecondLocalBackend`, the single enforcement point for "never spawn
  a second backend", placed at the one site that starts a local child.
- `resolveProfileBackendRoute` case 6 (local profile, no remote) now returns
  the shared primary with `scopePath`, so REST names the profile instead of
  relying on a per-profile process HERMES_HOME.

Escape hatches unchanged: `HERMES_DESKTOP_ISOLATED_BACKEND=1` keeps a private
backend, and remote/SSH/Cloud profiles stay pooled -- a different host is out
of scope for this host's singleton.

(cherry picked from commit 13bb57327e0d168816c7b0b9141516af3891142c)
2026-09-21 10:07:03 -07:00
Gille
996417c385 fix(desktop): wake retired bots for group turns 2026-09-21 10:02:49 -07:00
teknium1
2e497bead8 fix(dashboard): close the indirection holes and the unnamed-profile 400s
Round 2 of the REST profile-scope pass. A call-graph audit (AST over every
router plus their imported helpers, following functools.partial bindings,
closures and callbacks handed to to_thread/_spawn_job) found the writes a
signature grep cannot see, and the SPA/Desktop callers that never named a
profile at all.

Handlers that reached a config write through indirection:
* PUT /api/dashboard/plugin-providers wrote memory.provider + context.engine
  through functools.partial(_write_config_value, ...) — the SAME key
  PUT /api/memory/provider scopes — into the launch profile's config.yaml.
* POST /api/local-models/quickstart is `activate` plus a download; only
  `activate` had been scoped, so _set_runtime_enabled/_assign_default landed
  in the launch profile.
* POST /api/local-models/runtime/install regenerates launch presets under
  get_hermes_home().
* GET /api/model/recommended-default lazily PERSISTS discovered custom-provider
  models via build_models_payload -> _save_discovered_models_to_config.

Policy gaps:
* POST /api/ops/hooks is now gated like DELETE: writing an arbitrary command
  into `hooks:` and, with approve, into the consent allowlist is strictly more
  privileged than removing one.
* POST /api/sessions/prune (non-dry-run), DELETE /api/sessions/empty and
  POST /api/sessions/bulk-delete join the same destructive class.
* PUT /api/memory/provider ran its readiness check OUTSIDE the scope, so it
  judged the launch profile and could write a broken setting into another.

Callers:
* /api/credentials was missing from the SPA's PROFILE_SCOPED_PREFIXES, so the
  credential-pool delete button 400'd unconditionally; so were
  /api/dashboard/plugin-providers, /api/local-models and the model route above.
* The management scope was empty until the switcher resolved, so every
  destructive route 400'd in working UI on any host with a second profile
  directory. The backend now injects the profile it itself serves
  (__HERMES_DASHBOARD_PROFILE__) — and only when that name provably resolves
  back to its own home, so the fallback can never retarget another profile.
* authedFetch went around the scope entirely; the Desktop /api/ops callers
  (doctor, security-audit, backup, debug-share) sent no profile while Electron
  pins that family to the shared primary backend.
2026-09-21 09:59:11 -07:00
teknium1
6a3dcc39ab fix(dashboard): every REST route honours the profile it is given
One backend process now serves every profile, but a handful of REST handlers
resolved get_hermes_home() directly and ignored the profile the request named:
POST /api/memory/reset wiped the LAUNCH profile's MEMORY.md/USER.md no matter
which profile the switcher pointed at, and the same held for the curator state
file, webhook subscriptions, shell hooks + their consent allowlist, checkpoints,
backups/imports, the credential pool and the dashboard's own theme/font/plugin
preferences. Per-profile backend processes used to mask it.

* Every one of those handlers now takes `?profile=` and runs inside the
  existing `_config_profile_scope` / `_profile_scope` seam; action-spawning
  routes pass `-p <profile>` to the child via `_profile_cli_args`.
* Destructive routes (memory reset, webhook delete, hook delete, checkpoint
  prune, ops import/import-upload, credential-pool delete, curator run) refuse
  an UNNAMED profile with 400 while the process hosts more than one profile
  (`is_multiplex_active()`, decided at boot). A genuinely single-profile host
  keeps the old meaning, so plain `hermes serve` + curl is unchanged.
* Callers updated in the same change: the Desktop main-process routing table
  (`LOCAL_PRIMARY_SCOPED_ROUTES` + the /api/webhooks and /api/ops families) so
  a fixed handler reaches the shared backend WITH the profile query instead of
  a per-profile backend that no longer exists, and the dashboard SPA's
  `PROFILE_SCOPED_PREFIXES`.
2026-09-21 09:59:11 -07:00
teknium1
2632229bcf fix(auth): named profiles read the root auth.json again (revert #111724)
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.

Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.

Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
2026-09-21 09:55:40 -07:00
kshitijk4poor
6bde794b8b fix(desktop): type the known-origin WeakSet over object
resolveConfigWriteScope receives the explicit scope as ProfileScope (nullable
fields), so WeakSet<ConfigReadOrigin>.has() and the spread failed tsc. Key the
set by object identity and cast the known origin on the way out.
2026-09-21 21:35:01 +05:30
kshitijk4poor
6aef2e3bdd test(desktop): give the toolset-config-panel @/hermes mock the config-origin helpers
VoiceProviderFields now reads useHermesConfigRecord, which pulls
peekConfigReadOrigin/retainConfigReadOrigin from the '@/hermes' barrel. The
panel test's hand-rolled '@/hermes' mock predates those exports, so the two
TTS-row tests threw "No peekConfigReadOrigin export is defined on the mock"
and timed out. Provide inert stubs (no origin, identity retain); the panel
tests are not about write routing. Blank-line padding fixed by eslint --fix.
2026-09-21 21:35:01 +05:30
kshitijk4poor
8403fd1166 fix(desktop): give a captured read origin the same write tag on both paths
resolveConfigWriteScope had two ways to reach the same captured read origin:
the hook's `writeScope` passed explicitly (→ capabilityScoped(origin)) and the
record-only WeakMap lookup (→ spread of the origin). For a scope-selector pin
both produce the same object, because capabilityScoped already stamps
`priority: 'foreground'` on the GET. For the AMBIENT app-wide origin they
differed: the GET origin is `{ connectionId?, profile? }` with no priority
(profileScoped(undefined)), and re-running capabilityScoped on it added
`priority: 'foreground'` — a spawn-priority change origin/main's bare
`saveHermesConfig(patch)` never made.

Remember every bound origin in a WeakSet and, when the explicit scope is one
of them, spread it exactly like the WeakMap branch. Fresh selector pins still
go through capabilityScoped.

Probe (vitest, deleted): pin explicit == pin WeakMap ==
{connectionId:'connection-b', priority:'foreground', profile:'worker'};
ambient explicit == ambient WeakMap == {} (was {priority:'foreground'} vs {}).
2026-09-21 21:35:01 +05:30
kshitijk4poor
a8624999e6 refactor(desktop): read ProfileScope/profileScopeKey through the @/hermes barrel
use-config-record.ts pulled `ProfileScope` and `profileScopeKey` from
'@/api/client' while every other config helper came from '@/hermes'.
origin/main imported all of them from the barrel, and the split import made
the file's mocks and dependency graph inconsistent with its siblings for no
gain — the barrel already re-exports both. Restore the single import.
2026-09-21 21:35:01 +05:30
kshitijk4poor
1aeb913cd6 refactor(desktop): share one bound-record fetch for config GET and defaults
getHermesConfigRecord and getHermesConfigDefaults carried the same six lines
(capabilityScoped → window.hermesDesktop.api → bindConfigReadOrigin → return),
so a change to how a config read is stamped with its serving origin had to be
made twice. Fold both into a private fetchBoundConfigRecord(profile, request).

getHermesConfigDefaults loses its `profile` parameter: no caller passes one
(settings/index.tsx reset-to-defaults and use-hermes-config.ts both call it
bare), and the ambient-scope path is what the app-wide reset should use.
2026-09-21 21:35:01 +05:30
kshitijk4poor
74ecb8301b fix(desktop): let saveHermesConfig resolve the i18n record's own read origin
i18n/context passed `peekConfigReadOrigin(config)` as the explicit scope,
duplicating what saveHermesConfig already does: resolveConfigWriteScope's
second branch reads the record's captured origin from the WeakMap when no
object scope is given, and withConfigDisplayLanguage retains that origin onto
the derived record. Passing the origin explicitly also routed it through the
capabilityScoped() object branch instead of the captured-origin spread.

Pass `undefined` and drop the now-unused import.
2026-09-21 21:35:01 +05:30
kshitijk4poor
a2ac754a56 perf(desktop): restore structural sharing for the hermes-config record
`structuralSharing: false` made every refetch — each consumer mount at
staleTime 0 and every invalidate — hand back a fresh object even when the
payload was unchanged, so every consumer re-rendered, the MCP tab's useMemo
recomputed and the chat/terminal-font autosave effects re-armed.

Use replaceEqualDeep as before and re-stamp the surviving object with the read
origin of the NEW fetch, so a retained record still routes its next write to
the gateway that served the latest GET (the 'writeScope flips to connection-b
after refetch' invariant). setQueryData also runs this pass, but it stamps the
origin of the incoming value (which optimistic patches lack) and only once the
observer has built the query, so writeHermesConfigCache keeps its explicit
carry-over from the previous record.
2026-09-21 21:35:01 +05:30
kshitijk4poor
14bdc01b8f perf(desktop): expose writeScope as a getter instead of spreading the query
useQuery returns a tracked-props Proxy (query-core observer.trackResult):
only the props a consumer actually reads are compared when deciding to
re-render. `{ ...query, writeScope }` enumerated every key through the proxy,
so every consumer of useHermesConfigRecord re-rendered on fetchStatus /
dataUpdatedAt / isFetching churn — a regression against origin/main, which
returned `query` directly.

Attach `writeScope` with Object.defineProperty (non-enumerable, lazy getter
over `query.data`) so the proxy is returned intact and only `data` is tracked.
2026-09-21 21:35:01 +05:30
kshitijk4poor
1dd8afedb7 style(desktop): eslint --fix the stack-touched files and drop stray blank lines
Run `eslint --fix` over the files this stack changed so the
padding-line-between-statements warnings it left in config.ts and
config-settings.tsx are gone, and remove the blank lines the stack
inserted between a `saveHermesConfig(...)` call and its `.then(`/`.catch(`
chain (appearance-settings, terminal-font-setting, voice-provider-fields)
and before the `catch` in model-settings. Those hunks were pure noise in
the diff against main and split promise chains visually from their call.
2026-09-21 21:35:01 +05:30
kshitijk4poor
0024eee068 refactor(desktop): read config helpers through the @/hermes barrel
use-config-record.ts was the only non-test module importing '@/api/config'
directly, and its queryFn still wrapped getHermesConfigRecord in an
async/await/return shell left over from an earlier debugging pass. Go back
to the '@/hermes' barrel like every other consumer and pass the fetch
through unchanged.

Tests that bare-mock '@/hermes' and render a consumer of the hook would now
hit a missing `peekConfigReadOrigin` export. The hook test keeps the real
barrel via importOriginal and overrides only the fetch; config-settings and
model-settings spread the real '@/api/config' helpers into their existing
minimal mocks so the WeakMap peek/bind stays live without pulling the whole
barrel (and its store stack) into those suites.
2026-09-21 21:35:01 +05:30
kshitijk4poor
078af0a490 refactor(desktop): add retainConfigReadOrigin for derived config records
The "peek the read origin from the source record, bind it onto the derived
record" sequence was hand-written three times: the config cache writer in
use-config-record.ts, withConfigDisplayLanguage in i18n/context.tsx and the
auto-archive persist in sessions-settings.tsx. Each copy had to remember
the same null-guard, and a fourth site would likely have forgotten it.

Fold the sequence into one exported helper next to bind/peek in
api/config.ts and call it at all three sites. Behaviour is unchanged: a
source without a captured origin leaves the derived record unbound.
2026-09-21 21:35:01 +05:30
kshitijk4poor
44d492b98d fix(desktop): include writeScope in the voice-provider autosave effect deps
The autosave effect in voice-provider-fields.tsx reads `writeScope` to pick
the config write route, but its deps array omitted it behind the
exhaustive-deps disable comment. A scope change would keep a stale route
until another dep changed. List `writeScope` like the sibling autosave
effects in chat-font-setting, terminal-font-setting and config-settings
already do; the disable comment stays for `profile`/`baseline` (keyed by
scopeKey) and the locale string.
2026-09-21 21:35:01 +05:30
kshitijk4poor
b6ad556367 fix(desktop): include writeScope in the terminal-font autosave effect deps
The autosave effect captures writeScope for saveHermesConfig; without it in the
dependency list a route change after mount would keep saving through the stale
scope. chat-font-setting.tsx already lists it.
2026-09-21 21:35:01 +05:30
kshitijk4poor
5485d17e9e test(desktop): trim config route-binding tests to the two invariants
Keep only the behaviour the salvage guarantees: a record read from
connection A is never written to connection B after the primary changes,
and the hook's writeScope follows the record a refetch swaps in. The
dropped cases (explicit-pin precedence, unbound ambient path, null
scope, defaults round-trip) re-tested capabilityScoped/profileScoped
plumbing already covered elsewhere. The kept hook test also asserts
writeScope is `undefined`, not `null`, before data loads.
2026-09-21 21:35:01 +05:30
kshitijk4poor
d578fafcc7 refactor(desktop): read the auto-archive write route from the record
AutoArchiveSetting mirrored the config read origin into a writeScopeRef
that had exactly one writer and one reader. Read peekConfigReadOrigin
(config) at save time instead, and rebind the origin onto the replaced
state snapshot so a second toggle still routes to the gateway that
served the GET (the spread copy is a new object the WeakMap doesn't know).
2026-09-21 21:35:01 +05:30
kshitijk4poor
f0eed37de4 fix(desktop): spread the captured config read origin on write
resolveConfigWriteScope rebuilt `{ profile, connectionId }` from the
captured origin, silently dropping `priority: 'foreground'` that
capabilityScoped() attaches to explicit scopes (#111651). Spread the
captured origin instead, and give an explicit profile string the same
foreground priority profileScoped(string) grants, so route-bound saves
keep their scheduling priority.
2026-09-21 21:35:01 +05:30
kshitijk4poor
744634fd94 fix(desktop): default config writeScope to undefined, not null
useHermesConfigRecord().writeScope is handed straight to saveHermesConfig
alongside sparse `setNested({}, …)` patches, so the read-origin WeakMap
misses and resolution falls through to capabilityScoped(writeScope) →
profileScoped(writeScope). profileScoped treats the two nullish values
differently: `undefined` keeps the app-wide `_apiProfile`, `null` drops
it. With `?? null`, any save fired before the first GET resolved wrote
the PRIMARY profile instead of the active one. Use `?? undefined`.
2026-09-21 21:35:01 +05:30
KoNit-K
389e89bd47 fix(desktop): retain config routes through cache refresh
(cherry picked from commit 283262339152b960c9e4c66d67d76f08a3f1c4e4)
2026-09-21 21:35:01 +05:30
KoNit-K
f63c7c803d fix(desktop): retain config routes for font and reset
(cherry picked from commit 6b79c046dd5f9fa4ca622b4d9714a4562c2d43ba)
2026-09-21 21:35:01 +05:30
KoNit-K
8961499ff6 fix(desktop): bind config writes to the config-read route
A config record read from gateway A must be written back to gateway A,
not to whatever registry primary is current at PUT time. Snapshot the
(connectionId, profile) that served each config GET in a WeakMap and
route saveHermesConfig / saveHermesConfigRecord through it; callers pass
the record (or its captured origin) as the write scope.

Salvaged from PR #109363. The Electron-side dispatcher
(dispatchApiRequestForSelectedConnection, remoteRegistryPrimaryConnectionId,
ipc wiring and its tests) was left out: it duplicated the registry-primary
rung already in main.ts::handleHermesApiRequest and threw for every
untagged /api/config/* write during the boot window.

(cherry picked from commit acc2781223433d689e52dd102d72cc13f8bb9d02)
2026-09-21 21:35:01 +05:30
teknium1
75b083e939 fix(ci): eslint ignores *.generated.ts; restore the contract file's header
The on-merge `npm run fix` bot (bc655bfb40, #118250) ran eslint --fix over
apps/shared/src/gateway-contract.generated.ts and stripped its `/* eslint-disable */`
header. scripts/gen_gateway_contracts.py emits that header, so
tests/tui_gateway/contracts/test_generated.py::test_generated_files_are_current has been
red on main since, and on every PR opened after it. Generated files are not lint targets:
ignore the pattern in the shared config and regenerate the file.
2026-09-21 08:19:52 -07:00
hermes-seaeye[bot]
bc655bfb40 fmt(js): npm run fix on merge (#118250)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-21 14:41:52 +00:00
Austin Pickett
782886f87d fix(desktop): bound the Chromium log while the shell is running, not only at launch
The startup reclaim caps the log a previous run left behind, but a shell
that stays up for days writing Chromium ERRORs is unbounded until it
restarts — the case @ehz0ah flagged in review.

Chromium owns that descriptor in append mode for the life of the
process, so rotation is the wrong primitive: renaming leaves the writer
on the renamed inode and the cap silently stops applying. Poll the live
size and truncate in place instead; an O_APPEND writer resumes at offset
0, so a noisy process stays bounded at ~cap plus one interval's output.
The timer is unref'd and every tick is best-effort.

Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>

Refs #100573
2026-09-21 10:23:04 -04:00
Austin Pickett
9a21794450 fix(desktop): optional Linux crash diagnostics must not be fatal or unbounded
Two follow-ups to #117851, both raised in review by @ehz0ah:

1. The diagnostics block created HERMES_HOME/logs with an unguarded
   module-level mkdirSync, before app readiness. A read-only or invalid
   logs path terminated the Linux desktop at startup — optional
   diagnostics killing the app they exist to diagnose. Wiring now runs
   through enableLinuxCrashDiagnostics(), where every step is
   best-effort: no writable logs dir degrades to "no Chromium log" (and
   keeps the crash reporter, which writes elsewhere), and a Crashpad
   handler that refuses to start is not a startup failure either.

2. Electron opens an explicit --log-file with APPEND_TO_OLD_LOG_FILE, so
   desktop-chromium.log accumulated ERROR/FATAL output across launches
   with no bound. desktop.log is capped at 10 MiB x 3 backups precisely
   because it once reached ~326 GB and exhausted the disk; the new file
   now goes through the same planner before Chromium appends to it.

Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>

Refs #100573
2026-09-21 10:23:04 -04:00
Austin Pickett
5c57bfdcaa refactor(desktop): a shared, path-parameterized log rotation planner
desktop.log's cap (10 MiB x 3 backups, plus the discard ceiling that
reclaims a boot-loop log outright instead of renaming it to .1) lived
inline in main.ts, keyed to one hardcoded path, and main.ts cannot be
imported by a test.

Move the planner to its own module, parameterized by base path, so the
Chromium diagnostic log added for #100573 gets the same bound and the
behaviour is provable without booting Electron.

Refs #100573
2026-09-21 10:23:04 -04:00
kshitijk4poor
8f7b3b48fc fix(desktop): close the Linux notification bus on will-quit
The session-bus socket was created lazily and never closed. main.ts
already tears down its pooled keep-alive sockets in will-quit so nothing
holds the event loop open or leaks an fd past app teardown; the new
long-lived socket joins that policy. Disposal goes through the existing
close path, so live notifications are failed the same way a daemon crash
fails them.
2026-09-21 19:27:10 +05:30
kshitijk4poor
a327f1ec14 test(desktop): cover both notification transports from any host
The transport was picked from process.platform, so the ipc test mocked
the Linux module and the only vitest lane (ubuntu) never exercised the
Electron Notification branch that Windows and macOS run, while the Linux
tests skipped everywhere else. Make the platform an injectable host
parameter: the ipc test drives the Electron branch as darwin, the Linux
tests pass linux explicitly and lose their skipIf.
2026-09-21 19:27:10 +05:30
kshitijk4poor
c93b157249 fix(desktop): no notification cooldown after a benign daemon swap on Linux
Every failure in show() armed the global 10 s cooldown, including the
ones caused by the notification daemon being replaced mid-call: the
owner-changed fences, and the bus's NameHasNoOwner error for a call
already addressed to the vanished unique name. Linux has no Electron
fallback, so every notification in that window was silently dropped
exactly when the new daemon was healthy.

Classify by generation rather than by error: a failure against an owner
that is still current is the daemon's fault and cools down; one against
an owner that has since been replaced does not. The fixture now answers
calls to a stale unique name with NameHasNoOwner like the real bus, and
the race test asserts the next notification reaches the new owner
immediately.
2026-09-21 19:27:10 +05:30
kshitijk4poor
f09af639ac refactor(desktop): drop unreachable fences in the Linux notification transport
`finished` cannot be observed inside show(): close() aborts the signal
first, and dbus-native settles an aborted invoke synchronously, so the
awaiting code throws instead of resuming past the check. `connection !==
state` is implied by the generation check (disconnect bumps the
generation before clearing the connection), `connection !== target.state`
in close() is implied by `delivered` (disconnect fails every live entry,
which clears it), and receive() can never run after release() deleted its
map entry.

A malformed Notify id now gets its own error instead of being reported as
an owner change.
2026-09-21 19:27:10 +05:30
BearHuddleston
ab4455d965 fix(desktop): release naturally closed Linux notifications
(cherry picked from commit 5b0a1e3ab1982c2db8dbbceadda1499d72827631)
2026-09-21 19:27:10 +05:30
BearHuddleston
326ce363a3 fix(desktop): prevent Linux native notification freezes
(cherry picked from commit 3dd43eb59480e279835ce0f6409567ee864d47fc)
2026-09-21 19:27:10 +05:30
Austin Pickett
00e08cec30 fix(desktop): a Linux SIGTRAP leaves its FATAL line and a minidump behind
Every Chromium CHECK/LOG(FATAL) traps at the same instruction (the
ImmediateCrash tail of logging::LogMessage::HandleFatal), so the cores
collected for #100573 all share one address and none say which check
fired. The message goes to stderr, which the .desktop entry and the
Omarchy wrapper both discard, and the app never started Crashpad.

On Linux, route Chromium's log (ERROR and above) to
HERMES_HOME/logs/desktop-chromium.log and keep local minidumps; nothing
is uploaded. Child processes inherit the switches, so zygote/GPU checks
land in the same file.

Refs #100573
2026-09-21 09:00:05 -04:00
kshitijk4poor
6668ac14a0 fix(desktop): count the answer row in a page-boundary fold
A hydrated page can start mid-turn: the first row is the tool-only
assistant, the second its tool result, and the answer arrives third with
no active bubble to append to. The bubble that fold creates stands for
all three backend rows, but only the two pending rows were counted
(serverRowSpan: 2), so the release rewound the older-page offset short
of what it gave up and 'Show earlier' skipped rows the reader could no
longer reach.

Count the current answer row too, and pin it with a regression whose
page starts on the tool-only assistant/result pair.

Found in review of 72a5aa3fd2 by @ehz0ah.
2026-09-21 17:14:45 +05:30
kshitijk4poor
3f985c84cd fix(desktop): count released transcript rows in backend rows
The hydration fold merges a turn's tool rows into the bubble they belong
to, so a ChatMessage is not one backend row (a ten-tool-call turn is one
message over eleven rows). The older-page offset is counted in backend
rows, so rewinding it by the store's message count skipped history: the
next 'Show earlier' page started past rows the reader could never reach.

Carry the count out of the fold as ChatMessage.serverRowSpan and rewind
by it, which makes the rewind exact instead of approximate. Drop the
'has a rowId' release guard for the same reason it was wrong before: a
folded bubble can have no single durable id while every row behind it is
persisted and re-fetchable; what must not be released is a row still in
flight (pending).
2026-09-21 17:14:45 +05:30
kshitijk4poor
7c799c26ce fix(desktop): rewind the older-page offset in the backend's own units
The backend pages display history (include_compacted reads group by display
order; inactive rows are not counted), so a locally counted row total is not
guaranteed to be the offset's unit. Decrement the offset the backend reported by
the rows the store gave up, clamped at zero: a mismatch can only overlap a page
already in memory (which the merge dedupes), where an over-counted absolute
offset would skip rows the reader could then never reach.
2026-09-21 17:14:45 +05:30
kshitijk4poor
71a0ce913c refactor(desktop): share the tail-entry lookup, and plan nothing when nothing is released
Three copies of the tail-entry resolution (record the page, rewind, read the
state) now go through one resolveTailEntry; the retention planner reports
released:false without walking the weights or counting persisted rows, and the
hook checks for a fetch route before planning instead of discarding the plan
afterwards.
2026-09-21 17:14:45 +05:30
kshitijk4poor
14b4327445 fix(desktop): the store releases paged-through history instead of holding it for the window's lifetime (#77311)
The transcript window bounds what reaches assistant-ui, but the session store
underneath keeps every message it ever materialized: the tail hydration, every
older page "Show earlier" fetched, and every turn the window streamed. Renderer
retention therefore grows with session content and never comes back down — the
remaining half of #77311 after #117681.

`boundRetainedTranscript` keeps the live window plus one window page of slack
(so the first "Show earlier" still pages through memory) and releases the rows
older than that, which are off-screen and already persisted. Two invariants
keep the release safe: nothing is released without a durable `rowId` (a row the
backend cannot serve again would be content lost), and a branch group is never
split at the cut, for the same reason the window cut does not.

`rewindTranscriptTail` points the session's tail bookkeeping at the retained
prefix, so the released rows stay reachable: the existing older-page backfill
fetches them back on the next "Show earlier", and the rewind is refused outright
when the session has no recorded page route to fetch from.

`useTranscriptRetention` runs on the window's own cut moves — streaming grows
the tail, "Show earlier" prepends, a re-cut moves the anchor — so the weight walk
never happens per token. It computes the cut inside the store updater, from the
array it is about to replace: a view snapshot can be a flush behind the store,
and trimming a stale copy back into it would drop the newest rows of a live turn.
2026-09-21 17:14:45 +05:30
teknium1
73f3f64eb9 fix(desktop): jump-button message count settles after a timeline rail jump
countMessagesBelow measured the turn straddling the fold by its messages'
boxes. After a programmatic jump the straddling turn is still
content-visibility skipped for one frame — its messages report empty rects —
so they were dropped as "above the fold" and the count published short. The
ResizeObserver only re-measured when the turn's remembered size differed from
its real one, which is why repeat clicks on the same rail mark disagreed.

An unsettled measure (a straddling message with no box) re-schedules one frame
later; the retry publishes whatever is measurable so an empty turn never stalls
the count.

Part of #117298 (atom 2).
2026-09-21 01:07:06 -07:00
teknium1
f71939403c fix(desktop): timeline rail index follows turns created after the session was opened
A complete timeline index was a snapshot of the prompts that existed when it
was read: use-timeline-history fetched it once per (session, scope) and
timeline-index answered every later call from the cache, so the rail and
previousPromptRowId (the "Show earlier" anchor from a history page) never saw
turns added afterwards.

fetchTimelineIndex takes the newest row the caller has seen; a complete index
that does not reach it pages forward from its own cursor (one bounded
metadata request, merged by row id) instead of answering from the cache.
useTimelineHistory derives the newest persisted prompt row from the live store
as a scalar and pages once per new prompt — no timer, no polling.

Part of #117298 (atom 3).
2026-09-21 01:07:06 -07:00
teknium1
d90c3ad40a fix(desktop): group-chat member failure row names the error's first line (#117366)
A member turn that failed without a typed gateway reason rendered as a bare
"X hit an error" — a stopped backend, a dead IPC bridge and a provider refusal
all looked identical and gave the user nothing to act on.

groupFailureReason now falls back to the error message's first non-empty
line (trimmed, secret-shaped spans redacted, capped at 200 chars), so the
Activity row reads "X hit an error — <cause>". Typed data.reason and the
slot-wait classification keep precedence; the roster badge still classifies
from the same text.

The plugin fence keeps the Electron-side redactSecrets out of reach, so the
same shapes (bearer header, ?token=/?api_key= params, vendor key prefixes,
user@host:password) live next to the label.

Fixes #117366
2026-09-21 01:05:32 -07:00
teknium1
25d88ad0c4 feat(desktop): attach to the running host backend instead of spawning a second one
Multiplex-only (Teknium ruling): exactly ONE `hermes serve` per HOST,
multiplexing every profile. Desktop spawned unconditionally — every launch,
and every profile, added another backend — because startup had no discovery at
all and `dashboard-token.ts::isForeignBackendToken` actively REFUSES a backend
it did not spawn.

Startup now discovers first and spawns only when the host has none:

- `backend-discovery.ts` (pure): parse the machine-root `spawn-ledger.json` the
  CLI already writes after its socket binds (`process_identity.register_self`
  with `detail={host, port, profile}`), and decide attach-vs-spawn.
- `host-backend-attach.ts`: validate a candidate at the boundary that matters —
  served session token (`GET /` -> `window.__HERMES_SESSION_TOKEN__`), HTTP
  readiness, then `/api/ws` auth. A record is discovery only; only a validated
  candidate is ever attached. It also holds a host-level spawn gate so two apps
  starting at once produce one backend, not two.
- `primary-backend-startup.ts`: the attach step runs after update exclusion and
  before runtime resolution / the first-run gate — a host with a live backend is
  already set up.
- An attached backend has no child, so `child.exit` cannot drive recovery; a
  readiness poll invalidates the connection and hands the respawn to the same
  supervisor path.

A backend registered by ANOTHER profile is still THE backend and is attached to.
`HERMES_DESKTOP_ISOLATED_BACKEND=1` is the deliberate escape hatch.
2026-09-21 00:53:38 -07:00
teknium1
dec236b214 feat(desktop): Uninstall for standalone desktop plugins via Electron IPC
Rows whose only half is a folder in <HERMES_HOME>/desktop-plugins (no agent
package behind it) get the same trash button + destructive confirm as the
agent rows. The renderer names the FOLDER, never a path; Electron
(`hermes:plugin:removeDesktop` -> desktop-plugin-remove.ts) resolves it under
the app-level root, refuses anything that is not a single contained segment,
refuses unified-package halves (the reconcile would re-copy them; the agent
uninstall prunes them), removes a symlinked folder as the link, and deletes
the tree. The loader then retires the entry (unload, drop rows, stop watch)
so the pane/commands vanish without waiting for the next scan.

Why an IPC: these plugins have no agent half, so `plugins.manage remove` on
the gateway cannot reach them — they live on this computer, in every
connection mode.
2026-09-21 00:09:50 -07:00