Address review on the persisted unread dots, plus a latent data-loss bug
in the shared persistence helper that the restart e2e exposed.
Review findings:
- Session ids are caller-supplied and each profile backend is its own
namespace, while the desktop's lists routinely mix profiles (cron and
messaging slices are always cross-profile; recents are too in
all-profiles mode). Both persisted records are now bucketed per
profile - nested records keyed by the ROW's own profile
(normalizeProfileKey, absent -> default), never the live gateway's,
except the live busy->idle edge with no loaded row, which can only
come from the active gateway. Same-id sessions in different profiles
no longer share watermarks or markers.
- Markers are now bounded (200 per profile, oldest evicted) and cleaned
up when a session leaves the user's world: forgetSessionUnread() is
wired into removeSession, archiveSession, and the settings
permanent-delete path (which bypasses the other two).
Cold-boot clobber (found by the restart e2e after the refactor):
- persistentAtom wrote its value back to storage immediately at
creation. On a cold boot the bundle can evaluate against a storage
snapshot that has not caught up yet, so that echo overwrote real
records with the fallback. Creation is now read-only; only actual
changes persist. Regression-tested in persisted.test.ts.
- The read side of the same race is handled in session-unread.ts: the
first list arrival re-reads both records from storage (readable by
then) and merges them under the in-memory state, so a boot that
seeded empty atoms adopts the disk state instead of re-seeding every
row and burying the unread gap. Unread listeners are also disabled in
secondary windows - their partial list view must not write the
primary's whole-record state (same isolation rule as session tiles).
Tests: cross-profile same-id regression, live-edge profile fallback,
forgetSessionUnread cleanup, marker cap, persistentAtom creation
read-only; the restart e2e passes again end to end.
#76185 removed both defaults; only the bare shift+n chord hijacks normal
typing (uppercase N / IME input outside an input field). Cmd/Ctrl+N is a
deliberate two-key chord that matches every browser and chat app, so it
stays. Follow-up narrowing of the salvaged fix.
Adds a rebindable 'session.archive' keybind action (shipped unbound, like
session.togglePin) plus an ⌥+⇧-click gesture on sidebar session rows,
extracted into a pure, unit-tested click resolver so modifier precedence
(⌥⇧ archive vs ⇧ pin vs ⌘/⌃⇧ new window) stays correct.
Salvaged from #59759. Closes#59308.
The session.new default keybindings included 'mod+n' and 'shift+n',
which fired when users typed uppercase N (Shift+N) or accidentally
pressed Ctrl+N while using an IME to type Chinese — silently creating
new sessions mid-conversation.
Remove both combos so New Session is only reachable via the sidebar
button. Ctrl+Shift+N (session.newWindow) is unaffected.
Addresses review feedback on #68744: downscaling inside
attachmentPreviewDataUrl made previewUrl the 2048px PNG, but current main
(d8cb73b4ab) feeds that field to ImageLightbox and useImageDownload, so
large attachments would open and download at reduced resolution.
- attachmentPreviewDataUrl returns the full-resolution data URL again
- ComposerAttachment gains thumbnailUrl?: string (store/composer.ts)
- attachImagePath stores previewUrl (full-res) + thumbnailUrl (downscaled)
- Attachment pill renders thumbnailUrl ?? previewUrl; lightbox/download
keep the full-res previewUrl
- optimisticAttachmentRef prefers thumbnailUrl for the in-flight bubble
display ref (same main-thread decode freeze at send time)
- Attachment-level regression test in use-composer-actions.test.ts:
full-res previewUrl preserved while thumbnailUrl is a separate
downscaled value (4000x3000 -> 2048x1536)
Chinese/Japanese/Korean IMEs emit keydown events during composition that
carry preedit keystrokes and the commit keypress (Enter/Space/Shift for
candidate selection). Treating them as combos fires unrelated keybinds —
e.g. typing 你 with a CJK IME could dispatch session.new and silently
open a new session.
Guard comboFromEvent():
- Bail out entirely while composing (event.isComposing or key === 'Process')
- Ignore keydowns whose event.key is a bare modifier name but whose code
is a regular key — legacy IMEs that synthesize keystrokes (Q9 2002 sends
key="Control" with code="KeyW") would otherwise canonicalize into
phantom combos like mod+w that close the active tab.
Tested with Q9 (九方) legacy IME on Windows.
First slice of multi-source agent support: the desktop can now persist ANY
number of named backends (local runtime, remote gateways, Hermes Cloud
instances, SSH hosts) side by side instead of one global connection plus
per-profile overrides.
- electron/connection-registry.ts: pure v2 registry module — required
case-insensitively-unique labels (device names), @name-device handle rule
for duplicate profile names across sources (agentHandle), defensive
normalizeRegistry for corrupt files, one-time v1→v2 migration that imports
the global block + per-profile overrides (deduped by URL/host) and leaves
connection.json untouched for older builds.
- main.ts: connections.json storage beside connection.json (same secret
posture: safeStorage-encrypted tokens, 0600, tighten-before-parse, mtime
cache) + hermes:connections:* IPC (list/save/remove/set-primary/test).
Test maps registry entries onto the existing testDesktopConnectionConfig
probe stack — no new probe code.
- Settings → Connections: manage the registry (add/edit/remove/test/make
primary) with forced naming; local entry is non-removable; removing the
primary retargets to local. en + zh locales.
Storage-level only by design: routing/pool generalization to composite
(connection, profile) keys, the multi-source roster, plugin SDK surface, and
fan-out updates land as follow-up PRs.
Sibling site of the idle-resume rule from the stale-fold fix: the
assistant-tail append exit (user row persisted, no projection row) still
carried the journal's streamId onto a not-running resume, which kept the
journal entry alive (persistInFlightTurnState only clears when streamId is
null) and re-folded the same tail on every open. Apply the same
keepPending gate and pin it with a regression assertion.
Refs #85308
The inflight-turn journal can outlive the turn it recorded (reclaim,
reconnect or restart races skip the settle that clears it). On session
resume the fold then re-appends journaled assistant rows to a transcript
that already holds the committed replies, so the conversation ends with
duplicate answers in scrambled order. The fold also carried the stale
entry's streamId onto the resumed state on an idle resume, which kept the
journal entry alive (persistInFlightTurnState only clears when streamId is
null) and re-folded the same tail on every open.
Detect text-level staleness before the append path: when every recoverable
journaled assistant row already exists as committed text in the base
transcript, treat the entry as caught up and clear it. Only keep a stream
target when the resumed session is genuinely running (keepPending), so an
idle resume self-heals instead of re-folding.
The inflight journal regression test always spied on Storage.prototype,
but Node 26's jsdom setup can provide a plain in-memory localStorage fallback.
That left the test unable to observe the setItem call in CI even though the
per-session journal write was correct.
Select the native window.Storage prototype when available and otherwise spy
on the active localStorage object, preserving the assertion across both
storage implementations.
Refs #82832
The desktop journal synchronously read, parsed, cloned, and rewrote one
aggregate localStorage value while streamed turns were repainting. Large
tool results and multi-session state could therefore block the renderer and
leave the app unresponsive, while the existing macOS diagnostic path lacked
a real native hide/restore regression check.
Store bounded recovery projections under per-session keys, migrate legacy v1
data once, isolate quota and storage failures, and preserve the newest
recoverable tail without allowing oversized writes to replace valid state.
Add a real Electron/CDP macOS-arm64 A/B harness with native visibility control,
renderer heartbeat and Settings/composer/transcript checks, plus focused
regressions. Keep bulk tool payloads, diagnostics, and the existing recovery
merge behavior out of the hot path.
Fixes#63047
Three safeguards keep finished chats from looking busy: tool rows seal on turn settle, vanished runtimes clear awaiting state and open tool parts, and late stream events no longer land in a freshly opened session.
`refreshSessions` swaps the session page into `$sessions` only when
`sameCronSignature` reports a change, and that signature compared row
content — id, lineage root, title, source, profile, preview,
message_count, last_active, ended_at — but not row state. A page whose
only delta was `pinned` was judged identical and discarded, so the row
cached in the atom kept its old flag indefinitely. An idle conversation
never moves any of the compared fields again, which is exactly the kind
a user goes and unpins.
`session-pin-sync` treats that row as authoritative. Its write guard
(daeedf67c) is released by a page that CONFIRMS the value it wrote, and
falls back to letting the server win once WRITE_GUARD_MS elapses with no
confirmation. Because the confirming page was filtered out one layer up,
the fallback was the only branch that ever ran: ~10s after an unpin the
next reconcile read the frozen `pinned: true` row and called
pinSession() again. Adoption marks the id `mirrored`, so the push pass
never corrected the backend either — the local pin set and
sessions.pinned drifted apart permanently, which is why four of five
pins rendered in the sidebar read pinned=0 in state.db.
Compare both flags so a pin-only page reaches the atom. That restores
the guard's confirm path and makes WRITE_GUARD_MS a backstop again
rather than the load-bearing branch. `archived` is included for the same
reason: it is row state a consumer reads. Neither flag moves outside a
deliberate user action, so the churn the gate exists to prevent is
unaffected.
The existing `releases the guard once a page confirms the written value`
test passes on main because it hands `$sessions` the confirming page
directly — the gap was in the pipeline that decides whether such a page
is ever delivered.
Fixes#76919
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Clarify prompts share the same emitted-while-detached failure class the
pending-approval replay fixed: `clarify.request` rides `_block()`'s pending
registry, so a client whose transport was down when the event fired never
sees the question and the agent thread stays parked until timeout.
Widen the resume snapshot the same way:
- tui_gateway/server.py: `_live_session_payload` now carries
`pending_clarify` — a read-only snapshot of the clarify prompt still
blocking the session, scoped to the owning runtime sid. The registry stays
authoritative; the embedded request_id resolves via clarify.respond.
- Desktop resume paths (`use-session-actions`) restore the parked clarify
into the clarify store (multi_select preserved) and flag needsInput,
mirroring restorePendingApproval on both the activate and resume paths.
- pending_approval replay now also forwards the queue-injected request_id so
the restored prompt responds with exact-request correlation.
- Tests: server-side replay + scoping test; harmonized the #82087 replay
test with the request_id `_ApprovalEntry` now injects.
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).
Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.
Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.
Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
post-turn tick completion, and a supervised loop_wakeup_watcher that
injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
notification-poller wakeup driver + post-turn completion in the turn
dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
ticks defer until it finishes, pauses, or parks; real user input always
wins over both
Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
The honesty half (no fabricated counts) leaves shallow installs permanently
count-less. The compare API knows the full graph regardless of local clone
depth: GET /repos/<o>/<r>/compare/<current>...<target> returns ahead_by —
exactly the behind count the shallow boundary lost.
- hermes_cli/banner.py: _github_compare_behind() (bounded, unauthenticated,
best-effort); wired into _check_via_rev and the shallow branch of
_check_via_local_git. ahead_by==0 with differing tips = local-ahead => 0.
- hermes_cli/update_cmd.py: hermes update --check shallow path prints the
exact count when recoverable, presence-only wording otherwise.
- apps/desktop/electron/update-count.ts: compareApiUrl() +
parseCompareBehindCount() pure helpers; main.ts fetches the count when
resolveBehindCount() returns null, and the SSH-official passive path stops
fabricating behind:1 (uses compare API + updateAvailable flag).
- apps/desktop/src/lib/version-status.ts: updateAvailable now applies to the
client target too, so a shallow desktop install shows '(update)' instead of
nothing (or the old frozen '(+1)').
Fixes#84591; CLI siblings of #78253 / #53479 behavior.
E2E: live compare API returned 61/62 for real 61/62-commit gaps and 0 for the
reversed (local-ahead) pair; real shallow-clone fixture (depth-1 clone +
depth-1 fetch, merge-base broken) recovers the exact count with the API and
falls back to the honest sentinel offline.
Vercel, Supabase, Netlify, Hugging Face, Asana, Intercom, Airtable,
Webflow, PayPal, and Square join the directory. Every entry is a
vendor-operated remote with its docs page linked, same URL-only rule
as the founding eight. Trigger notes where words are ambiguous:
'square' the English word never fires (squareup only), and
vercel.app/netlify.app deploy-preview hosts are deliberately absent
(a pasted preview link is about the site, not the platform). Brand
glyphs wired for all newcomers.
Streamdown parses assistant markdown with allowDangerousHtml, so every
`<tag>` run in a message goes to parse5 and then through
hast-util-from-parse5, which recurses once per level of unclosed
nesting. Past roughly 1,750 consecutive unclosed tags that overflows the
call stack and throws RangeError out of the middle of a React render.
Nemotron-3-ultra degenerates into exactly that: thousands of `<unk>`
tokens emitted as reasoning, every one of them an element parse5 opens
and never closes. The payload is persisted to the session, so the throw
comes back on every reload.
Clamp the depth of unclosed elements in the prose path and escape the
opening `<` past the cap, leaving the text visible as the literal
`<unk>` it always was. The bound is on depth, not size: 20,000 balanced
`<b>x</b>` pairs and 20,000 void `<br>` tags parse fine because neither
drives the tree deeper, so only unclosed elements are counted and normal
markup is returned by identity.
A renderer-local directory of official hosted MCP remotes (URL-only,
vendor-documented endpoints — deliberately not the reviewed install
catalog) powers keyword and pasted-link suggestions: typing jira or
pasting a *.atlassian.net URL floats an 'Add Atlassian' pill in the
composer's micro-action strip. Matching is whole-word/phrase (unicode
boundaries) plus strict host-suffix on links, host hits outrank
keywords, capped at two, debounced 600ms, and excludes servers already
in mcp_servers. Pills are session-scoped like the micro-action badges
and self-limiting rather than dismissible — they exist only while a
trigger is in the draft. A click drafts the setup request; the agent's
setup_mcp card carries the consent. Brand glyphs extracted from the
mcp-tab into lib/mcp-brands (shared, monochrome marks follow the theme
so GitHub/Notion/Vercel survive dark mode).
The card follows the approval bar's consent vocabulary (primary-tinted
action + ghost decline, ⌘⏎/Esc with clarify's focus-stand-down rule) on
clarify's widget shell. Install prefers the reviewed catalog entry (env
prompts inline, background installs polled to completion) and falls back
to the desktop suggestion directory via the validated add-server POST +
OAuth; success reloads live MCP tools before unblocking the agent so it
resumes with the tools it was just promised. Esc stays live mid-flight
as cancel — the abandoned flow aborts at its next poll and a post-write
cancel rolls the config entry back. Typing while the card is pending
declines it and sends normally (skipClarifyRequest's pattern), and the
request/tool.start rows merge on the server arg so reconnects can't
double-render the card.
A pasted GitHub PR comment deep link (#discussion_r… / #issuecomment-…)
now lands as a typed review attachment instead of a bare url chip. The
card attaches optimistically and resolves through gh in the background —
author, file:line anchor, body, and the diff hunk — expanding at send
into an anchored fenced block, so "address this" carries exactly what
"this" is. When gh can't answer (offline, unauthenticated, foreign repo,
remote gateway) the card downgrades to the plain url ref and nothing is
lost.
A new "Inbox style" toggle in the sidebar filter menu renders the flat
recents list as cards: a workspace header line (project when it resolves,
else the cwd leaf, else Home) with the age at its right edge, the title
grouped with a one-line last-message preview, and a model + size footer.
The preview line ships on by default and has its own Show-menu toggle,
offered only while Inbox style is active — the one-line row has nowhere
to put it.
A render variant, deliberately not a grouping — it composes with whichever
grouping is active and only the flat recents list opts in; pinned, project,
and messaging surfaces keep the one-line row. Spacing hangs off a single
--card-gap variable; the title/preview pair is one grouped cell with its own
tighter internal gap. The age and kebab sit in flow inside the header line
rather than a full-height side column, so title, preview, and footer span
the card's entire width.
The card's project label reads through a selector that resolves the label
string, so tree polls with fresh atom identity repaint only rows whose label
actually changed.
* fix(desktop): keep config/structured code blocks fenced instead of unwrapping to prose
The desktop markdown preprocessor has a "prose fence" heuristic that
strips the fence off blocks it thinks are wrapped prose. Its
`proseLines >= 3 && codeSignals === 0` rule fires on ANY 3+ line
plaintext block with no JS/SQL tokens -- which is exactly what an SSH
config, a .env dump, or any INI/key-value listing looks like. The result
was that a fenced ```-block of SSH config rendered as a flat paragraph
instead of a code block.
Add isLikelyStructuredText() and use it as a veto in both
isLikelyProseFence() and isLikelyProseCodeBlock(): a block is treated as
structured (and kept fenced) when it has indented continuation lines, or
when it has no sentence-ending punctuation and a majority of lines are
`Key value` / `Key: value` directives. Real wrapped prose has
sentence-shaped lines and no per-line indentation, so it still unwraps as
before. The bullet-prose case in isLikelyProseCodeBlock is checked first
so markdown bullet lists remain prose.
Tests: markdown-code.test.ts gains SSH-config / flat-config / .env
regression cases for both functions, plus direct isLikelyStructuredText
coverage, and re-asserts that genuine paragraph prose still unwraps.
* fix(desktop): tighten config-line detection to not match punctuation-less prose
The first CONFIG_LINE_RE matched any 'word word' line, so a wrapped prose
fragment with no sentence punctuation (e.g. 'the quick brown fox jumps')
was misread as a config directive and its fence kept. Split into an
explicit-separator form (Key: value / Key = value) plus a short 2-3 token
'Key value' directive form; a real sentence line has more tokens, so
punctuation-less prose is no longer treated as config.
* fix(desktop): evict settled session states nothing on screen references
Closing a tile never removed its runtime's entry from $sessionStates, so
every tile ever closed parked its full transcript in the map for the life
of the process. Each leftover entry taxes every subsequent stream flush —
the map is spread-copied per delta and the busy/attention/draft projections
walk every entry per publish — so the app got slower the longer it ran,
which users read as "I need to clean my sessions/dbs".
Publish now evicts a settling state when no tile and not the primary view
holds its runtime (transition side effects still fire, so the settle keeps
its unread dot), and closing a tile drops an already-settled state on the
spot. Busy and needs-input states stay: background turns feed the sidebar
dots, and a first publish always lands because a resume can publish a beat
before the surface binds the runtime.
16 tiles streaming in a 2x2 grid with a day's worth of closed-tile residue:
worst-second 34 -> 58 fps, p99 frame 90 -> 28 ms, longtasks 37 -> 0.
* perf(desktop): index lineage aliases per sessions-list reference
lineageAliases scanned the whole recents list per call, and it is called
per cached session state per status projection per message delta — with a
populated sessions DB and a few busy sessions that multiplied out to
millions of row checks a second during streaming. Build the alias index
once per list reference (the list is replaced wholesale, never mutated)
and look aliases up in O(1).
* perf(desktop): journal each in-flight turn under its own storage key
The v1 journal kept every session's tail in one localStorage key, so each
throttled write re-parsed and re-stringified EVERY busy session's snapshot
— a grid of concurrent streams turned that into a whole-store JSON round
trip dozens of times a second, all on the main thread. Per-session keys
make a write O(own tail) no matter how many other sessions are streaming.
A v1 store migrates on first touch; expired/overflow crash residue is
pruned once per renderer.
* perf(desktop): stress the multitab scenario across grid/streaming/DB axes
The one-stack multitab run hid every cost this round of fixes removed: it
drove hook.publish (store only — no journal, no wiring cache), with an
empty recents list and no closed-tile residue. Streaming now routes through
hook.update (the real gateway write path), and the scenario grows axes for
the workloads users actually hit: --zones splits tiles across visible grid
zones, --streaming caps how many sessions are mid-turn (zone leaders
first), --sessions seeds a lived-in recents list, --dead models settled
sessions no surface references. launch.mjs pins HERMES_DESKTOP_CDP_PORT so
a non-default --port survives the app's own dev-CDP flag.