web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.
- tools/browser_tool.py: remove _extract_relevant_content and
_get_extraction_model; oversized snapshots always truncate at line
boundaries, store the full tree to cache/web, and append a read_file
pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
pattern as session_search/PR #27590), cli.py defaults + env bridge,
gateway/run.py bridged keys, hermes config display, hermes model picker,
dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
LLM path is gone and stored files are secret-redacted
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.
pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.
Fixes#93518.
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.
- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
_AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
of fallback-providers and the delegation /review section missed in
#93339
The OpenCode Zen wire slug for the Ox Alpha stealth model is opaque
(x-preview-f-free); users searching the picker for 'ox' or 'ox-alpha'
found nothing. Adds the search alias across all four synced alias
tables (CLI, desktop, web, TUI) plus tests. Wire id is unchanged and
still what renders and gets sent to the provider, matching the k3 →
kimi-k3 precedent. No canonical-dedup collision with opencode-go's
keyed ox-alpha-free slug.
Wire babel-plugin-react-compiler through @vitejs/plugin-react v6's
reactCompilerPreset + @rolldown/plugin-babel in the web and desktop
vite configs, scoped to modules that can actually contain components
or hooks (JSX syntax or a react-ish import — the preset's default
filter babel-parsed every TS module). Both vitest configs run compiled
components, so rules-of-react violations fail in CI.
Also fixes the latent bug the compiler exposed: usePluginI18n kept a
stable translator identity over a mutating locale registry, so
memoized consumers (React.memo today, compiled components tomorrow)
served stale strings after a late bundle registration. The registry
version now keys the translator identity — correct with or without
the compiler.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
Self-hosted dashboards served over plain HTTP on a LAN have no
navigator.clipboard (insecure context), so every direct writeText call
silently failed. web/src/lib/clipboard.ts already ships the HTTP-safe
copyTextToClipboard fallback but only OAuthLoginModal used it; ChatPage
(OSC 52 + Ctrl/Cmd+Shift+C), ProfilesPage, SystemPage, and WebhooksPage
all bypassed it. Route them through the helper and add a source-level
regression test that rejects any new direct clipboard write outside
lib/clipboard.ts (clipboard reads are exempt: no legacy fallback exists).
Sabotage-verified: the guard test fails when a direct write is introduced.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Wire the continuity flag through every cron-creation surface, not just the
model tool:
- dashboard (web/): checkbox in the cron job editor; form state round-trips
the stored reserved 'self' entry into the toggle and strips it from the
context_from textarea; web_server dashboard validator skips 'self'
(create precedes the job's existence)
- Bot Mode Routines tab (hermes-bots plugin): Continuity checkbox in the
New Cronjob dialog, forwarded through cron.manage
- tui_gateway cron.manage RPC: optional continuity param on action=add
vitest cron-job suite 10/10 (4 new), tsc app project clean, py_compile clean.
ChatPage stays mounted (hidden) on every dashboard route so the PTY
survives tab switches. With the visualViewport listeners attached
unconditionally in the PTY effect, the NS-434 scroll pin
(window.scrollTo(0, 0)) fired whenever a soft keyboard opened on ANY
page — fighting iOS Safari's own scroll-into-view for focused inputs on
Settings, Sessions, etc.
Move listener attachment into an isActive-gated effect: attach on
chat-tab activation, detach on deactivation (clean lifecycle, no
if-check inside the hot handler). The handler reads through refs
populated by the PTY effect, so the two lifecycles stay independent.
Deactivation also clears any applied inset padding so a keyboard left
open during navigation can't strand stale bottom padding on the hidden
terminal wrapper.
Adds a component-level test asserting listeners attach only while
isActive and detach on deactivation.
On mobile the on-screen keyboard overlays the layout viewport instead of
resizing it (iOS Safari always; Android Chrome under its default
interactive-widget=resizes-visual). The dashboard shell is a fixed h-dvh
column, so the xterm host's bounding box never changed when the keyboard
opened: fit() computed identical (cols, rows), no RESIZE reached the PTY,
and the Ink input line — drawn at the bottom of the grid — stayed hidden
under the keyboard.
Fix, in three parts:
1. Keyboard-inset handling (new web/src/lib/keyboard-inset.ts).
computeKeyboardInset() measures the layout-viewport region obscured by
the keyboard via window.visualViewport
(innerHeight - vv.height - vv.offsetTop, with an 80px floor so
collapsing URL-bar chrome doesn't thrash the grid). ChatPage applies
it as bottom padding on the terminal wrapper, which shrinks the host →
the existing ResizeObserver/fit path recomputes rows and sends RESIZE →
Ink redraws the input line above the keyboard. Listens on both vv
resize and scroll (offsetTop changes arrive as scroll events on iOS).
2. interactive-widget=resizes-content in the viewport meta. Android
Chrome 108+ then resizes the layout viewport natively and the JS inset
computes ~0 (harmless no-op); iOS ignores the directive and takes the
JS path.
3. Scroll pinning. iOS auto-scrolls the page to reveal xterm's hidden
textarea on focus, which drags the fixed shell offscreen. While a
keyboard inset is active we pin window/scrollingElement scroll back to
0 and term.scrollToBottom() so the freshly-resized input line stays in
view.
Unit tests cover the inset math (thresholds, offsetTop, rotation races,
fractional geometry, non-finite guards). Grid-level behavior needs a real
device pass — DevTools emulation doesn't model keyboard insets.
The word-delete branches sent ^W / ESC d whenever the socket was OPEN,
bypassing the shouldBlockPtyInput gate that term.onData applies to
every other keystroke, so a reconnecting/closed session could still
receive shortcut bytes. Route both branches through a shared guarded
sender that applies the identical socket + PTY state check, and cover
the non-open PTY states in the lib tests.
The dashboard chat is an xterm.js terminal in a browser tab, so several
editing keys never reached the input:
- Bare Ctrl+V fell through to the TUI, whose server-side clipboard read
can't see the browser/OS clipboard, so it reported "No image found in
clipboard". Route Ctrl+V through the same navigator.clipboard path as
Ctrl+Shift+V (image-or-text). Fixes#24860.
- Ctrl+Backspace / Ctrl+Delete never sent a word-delete: xterm.js emits
bare DEL regardless of modifier. Send ^W (0x17) and Alt+d (ESC d) to the
PTY so readline / prompt_toolkit delete the previous / next word.
Ctrl+W itself stays unavailable in a browser tab: it is a reserved
shortcut (close tab) that preventDefault cannot suppress. Ctrl+Backspace
covers word-delete there; the Electron desktop app can bind Ctrl+W.
The right-hand chat panel (model picker + session list) is a fixed 240px
column on desktop with no way to hide it. Add a collapse button (X) in the
panel header and a floating 'panel' button over the terminal to reopen it,
mirroring the collapsible app sidebar. The choice is persisted in
localStorage (hermes-chat-panel-collapsed) so it survives reloads.
(cherry picked from commit 854f1325a5490bcc968723a82746c939118f6b67)
React 18's root-level event delegation intercepts keydown events with
keyCode 229 (the "composition in progress" signal sent by the browser
during non-Latin IME input) and synthesises an onCompositionStart event.
That synthetic path sets internal composing state that interferes with
xterm.js's own IME handling on its hidden textarea, causing the first
keystroke of each composition chunk to be silently dropped.
The fix adds a capture-phase keydown listener on the terminal host div
that stops propagation of keyCode-229 events before they reach React's
delegation layer. xterm.js relies on native compositionstart/
compositionend on its internal textarea — not on keydown — so blocking
the propagation is safe.
Fixes#52111
The salvaged kanban ConfirmDialog added confirmDoneMany/confirmArchiveMany/
confirmBlockedMany to types.ts and en.ts only; the strict locale type
requires every locale to carry them. English fallback pending translation.
Migrates 8 of 12 native dialog call sites in the kanban dashboard plugin
to the SDK's ConfirmDialog primitive (added in PR #50550):
- moveTask, moveSelected, applyBulk, deleteTask, deleteSelected,
archiveBoard, removeAttachment, doPatch
The 4 remaining carve-outs (window.prompt for completion summary,
window.alert for missing summary, cli_hint clipboard fallback) are
documented inline — the host's ConfirmDialog hardcodes onClick → unmount,
preventing the keep-open-across-validation behavior the completion-summary
form needs. Followup: upstream a `disabled` prop to ConfirmDialog and
rebuild the completion body using host Dialog components.
New architecture:
- useKanbanDialogs(t) — Promise-based dialog state machine at
KanbanPage scope. request({kind, ...}) returns {confirmed, summary?}.
- KanbanDialog component — renders ConfirmDialog from SDK for kind=confirm.
- performMoveTask(taskId, newStatus, count, summary) — extracted shared
dispatch path for single + bulk moves (optimistic UI + PATCH/POST +
error recovery).
- requestDialog prop threading — KanbanPage → BoardSwitcher,
TaskDrawer → TaskDetail → doPatch/AttachmentsSection. Every call
site has a defensive fallback to window.confirm if the prop is
missing (verified by test_dashboard_done_actions_prompt_for_completion_summary
counting the cancel guards + destructive:true markers in the bundle).
New host i18n keys (web/src/i18n/en.ts + types.ts):
- kanban.confirmDoneMany / confirmArchiveMany / confirmBlockedMany
- kanban.trash.confirmTitle / confirmManyTitle
Tests:
- Replaced bundle-string-only completion-summary test with behavioral
coverage: bundle cancel-guard count + destructive marker count, plus
backend tests that confirm cancel preserves old status and confirm
dispatches the expected PATCH/DELETE body.
- Removed the SDK_CONTRACT_VERSION snapshot test from
web/src/plugins/registry.test.ts (forbidden by AGENTS.md
"Don't write change-detector tests"; the two remaining tests in that
file already cover the new SDK surface behaviorally).
Closes#50547 (consumers of #50550).
Cross-vendor re-review: Gemini 3.5 Flash + GPT-OSS 120B (both SHOULD-FIX,
no remaining BLOCKERs after these fixes).
Additive expansion of window.__HERMES_PLUGIN_SDK__. Plugins can now render
host-styled dialogs, confirmations, and toasts instead of falling back to
window.alert/confirm/prompt.
New components: Dialog, DialogClose, DialogContent, DialogDescription,
DialogFooter, DialogHeader, DialogTitle, ConfirmDialog, Toast.
New hooks: useToast (replaces showToast/toast pair), useConfirmDelete
(single-id delete-confirm state machine).
SDK_CONTRACT_VERSION unchanged at 1.1.0 — additive surface per
sdk.d.ts:23-25 (no major bump required).
Consumer: kanban plugin's 'replace native dialogs' work, see issue #50547.
A reference prototype showing 4 design variants is committed at
docs/design/kanban-dialogs/index.html.
Adds web/src/plugins/registry.test.ts (3 vitest cases) that smoke-test the
new keys are wired and that the version constant is unchanged.
- Shared per-job trigger controller (apps/shared) coalesces duplicate
clicks for the same profile+job inside a mounted client while letting
unrelated jobs run independently; the backend durable claim remains
authoritative across windows/processes.
- Two-phase feedback everywhere: the action stays disabled/spinning
while the request is in flight and the terminal success/error is
reported once, after the HTTP response — no premature success toast
(Web), matching the Desktop info notification.
- Desktop keeps the 24h trigger timeout for the synchronous long
operation and fences stale profile/list responses and unmounted
surfaces; the sidebar trigger button shows a spinner while busy.
The resume follow-scroll fix subscribes term.onScroll and reads
term.buffer.active; the ChatPage test fake predates both and crashed
the suite with 'term.onScroll is not a function' (4 unhandled errors).
Follow-up to the salvaged #59713 commits: main's ChatPage gained the
PTY resume sanitizer + hydration machinery after the PR was authored, so
the write-callback follow-scroll is applied to the sanitized write path
(term.write(rendered, followScroll)) rather than the raw ev.data writes
the original diff targeted. Also adds the contributor email mapping.
Extract the resume-scroll decision into lib/pty-scroll (isViewportPinnedToBottom, shouldFollowPtyOutput) and add focused vitest coverage, per review on #59591.
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers
The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.
The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.
Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:
* /auth/native/authorize now accepts a supports_password provider:
register the pending broker authorization as usual, then 302 the
system browser to the interactive /login form with the opaque
broker_state in the gateway's PKCE cookie (the same server-controlled
channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
broker handle, a successful credential check completes the pending
authorization exactly like the /auth/callback native branch — mint
the one-time loopback code, return the loopback redirect (validated
loopback-only at authorize time) as `next`, clear the PKCE cookie,
and set NO session cookies. A lapsed broker is a clean 400 telling
the user to restart sign-in; a failed credential attempt leaves the
pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
session provider is registered (previously only for non-password
providers), so the desktop selects the system-browser strategy for
password-only gateways.
Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.
Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(dashboard-auth): bind native password completion to the authorize-time provider
Review follow-ups for #75808:
* /auth/password-login now enforces that body.provider matches the
provider recorded in the server-set PKCE cookie by
/auth/native/authorize before completing a pending native
authorization. /login renders a form for every session provider, so
without this a native flow started for provider A could be completed
with provider B's credentials, binding B's session into A's pending
entry. The mismatch is rejected BEFORE credential verification (no
session minted, no oracle) and preserves both the pending entry and
the cookie, so the user can still submit the correct provider's form.
Covered by a two-password-provider E2E regression test.
* Update the two docs spots that still said password-only providers do
not advertise native_pkce (website desktop-native-signin guide and the
auth_flows type comment in web/src/lib/api.ts).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: map contributor email for #75808 (buffpesos)
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Review edge case: OOM dismissal was boot-keyed but live critical/elevated
dismissals were keyed only by severity. Dismiss critical, gateway reboots,
next poll still critical with no observed "ok" in between — the new
incident stayed hidden, and masked the OOM notice too, since critical
takes precedence.
Every dismissal key now embeds boot_id, so a gateway restart invalidates
prior dismissals of any kind. Within a boot, semantics are unchanged:
severity-scoped masking, escalation re-opens, confirmed "ok" recovery
clears live dismissals, "unknown" clears nothing. Missing boot_id
(pre-NS-656 image) degrades to a shared per-severity bucket as before.
Addresses the human review findings on the memory-pressure feature:
* [P2] Dismissal hid later incidents of the same kind. The gateway now
publishes `boot_id` (the lifecycle sentinel's started_at — changes on
every gateway life) in the /api/status memory block, and the dashboard
keys OOM-restart dismissal on it: acknowledging one restart no longer
mutes the NEXT one (the OOM-loop case this banner exists for). Live
pressure dismissals now also reset once pressure is demonstrably back
to "ok" — "unknown" (stale heartbeat) is absence of evidence and
clears nothing. Dismissal storage moved to a JSON list; old bare-string
entries fail JSON.parse and degrade to a clean reset.
* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
heartbeat), not proof the OOM killer acted — banner copy now says
"restarted unexpectedly, most likely because it ran out of memory"
instead of stating OOM as fact.
* [P3] Mobile header clearance was applied per-banner (mt-14 on both
MemoryPressureBanner and ProfileScopeBanner) AND on the content
(pt-14), double/triple-stacking 56px gaps when banners were visible.
Replaced with a single h-14 spacer above the banner stack.
api.ts already declares MemoryStatus for the /api/memory providers
endpoint; the NS-656 pressure block reused the name, and TypeScript
declaration merging fused the two shapes — 'tsc -b' in the Docker image
build failed on every test fixture. Renamed to MemoryPressureStatus.
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
* fix(dashboard): retry stalled events feed reconnects
* fix(dashboard): bound the PTY ticket request before the socket exists
ChatPage's connect awaits a single-use ticket from `api.buildWsUrl()`
before `new WebSocket()`. That request produces no socket, so a
rejection or a hang emits no `close` event and never arms
PTY_CONNECTING_TIMEOUT_MS (set after the socket is constructed). The
tab stranded on "connecting" with `connectInFlightRef` stuck true,
which also suppresses the page-resume reconnect path.
Give the ticket phase its own deadline and route both failure modes
into the existing backoff. A `ticketSuperseded` flag invalidates a late
ticket result so a timed-out attempt cannot open a socket behind the
replacement it scheduled, and cleanup clears the timer on unmount.
`scheduleReconnect` now takes `number | null` so an attempt that died
before any socket existed omits the "(code N)" banner suffix instead of
inventing one.
Same bug class as the events-feed fix in the preceding commit, on the
main chat surface.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
* test(dashboard): cover the PTY ticket connect deadline
Mirrors the events-feed cases in ChatSidebar.test.tsx: a rejected ticket
retries, a stalled ticket times out and its late resolution cannot open
a superseded socket, and a settled ticket disarms the deadline so
PTY_CONNECTING_TIMEOUT_MS remains the only guard on a wedged handshake
(NS-591 regression).
Both failure cases fail against ChatPage.tsx without the preceding fix.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
---------
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
* fix(dashboard): add events-feed reconnect policy helpers
Extract the reconnect arithmetic and close-code classification for the
ChatSidebar /api/events socket into a pure module so both can be tested
without a fake WebSocket or a mounted component.
Two decisions live here rather than inline in the effect:
- `shouldRetryEventsClose` — 1000 (normal) and 4401/4403 (auth) are
terminal; everything else, including 1005/1006 from a killed gateway
or a dropped network, is retryable.
- `isEventsFeedMessage` — the sidebar's banner is shared with
`info.credential_warning` and the JSON-RPC sidecar, so a reconnect may
only clear a message the events feed wrote itself.
Co-authored-by: Ishan Parihar <ishan@supreme-god.dev>
Co-authored-by: Vyre <vyre@ishanparihar.com>
Co-authored-by: eric-senyao <178080753+eric-senyao@users.noreply.github.com>
* fix(dashboard): auto-reconnect the events WebSocket with backoff
The chat sidebar's /api/events subscriber surfaced a static "disconnected"
banner on a transient drop and never retried, so a gateway restart or a
network blip left the feed dead until the user reloaded the page. The feed
drives the live chat title (session.info) and dashboard.new_session_requested,
both of which silently stopped working.
Reconnect with exponential backoff (1s → 2s → 4s → … → 30s cap, 15 attempts
then a terminal banner). Specifically:
- One scheduling path. `close` always follows `error` for a failed socket,
so scheduling from both — as the superseded PRs did — queues two timers
and leaks the one that is no longer tracked for cleanup. `scheduleReconnect`
returns early when a retry is already pending.
- The auth ticket is re-minted per attempt via `buildWsUrl`; tickets are
single-use with a short TTL, so replaying the first URL would 4401 on the
second attempt.
- A superseded socket's late close cannot schedule a retry on top of its
replacement (`isCurrent` generation check).
- A successful open resets the backoff and clears only the events feed's own
banner, leaving a credential warning or sidecar error visible.
- The pending timer is cleared on unmount, not merely neutered by the
`unmounting` flag.
Also retires two strings the tools box left behind when it was removed in
47fccc073 (#51737): the banner no longer promises "tool calls may not
appear" and the button reads "reconnect events feed".
Co-authored-by: Ishan Parihar <ishan@supreme-god.dev>
Co-authored-by: Vyre <vyre@ishanparihar.com>
Co-authored-by: eric-senyao <178080753+eric-senyao@users.noreply.github.com>
* test(dashboard): cover the events-feed reconnect bug class
Fake-timer coverage for the behaviors the superseded PRs changed without
tests. Each case was mutation-checked — reverting the corresponding guard
in ChatSidebar.tsx makes exactly that test fail:
- transient close reconnects, and backoff grows 1s → 2s → 4s
- error + close on one socket schedules ONE retry, not two
- a successful open resets the backoff to 1s
- 4401/4403 and a normal 1000 close never retry
- the attempt cap stops the loop instead of retrying forever
- reconnect clears the feed's own banner but not a credential warning
- unmount clears the pending timer (asserted via `vi.getTimerCount()`,
since the `unmounting` flag alone hides a leaked timer)
Co-authored-by: Ishan Parihar <ishan@supreme-god.dev>
Co-authored-by: Vyre <vyre@ishanparihar.com>
Co-authored-by: eric-senyao <178080753+eric-senyao@users.noreply.github.com>
* fix(dashboard): stop the events feed overwriting a foreign banner
Review catch: `clearEventsBanner` guarded the shared banner but `surface`
did not, so the guard was only half applied. A sidecar error or
`credential_warning` already on screen when the feed dropped was replaced
by "events feed disconnected" — and lost for good, since `error` is that
message's only home and the sidecar does not re-emit.
`surface` now writes only over an empty banner or one of the feed's own
messages. Declining to write does not affect the retry itself; the
reconnect still runs on schedule, it just stays silent while a more
important message holds the banner.
Both directions are covered: a foreign banner survives a drop, and the
reconnect still fires while suppressed.
---------
Co-authored-by: Ishan Parihar <ishan@supreme-god.dev>
Co-authored-by: Vyre <vyre@ishanparihar.com>
Co-authored-by: eric-senyao <178080753+eric-senyao@users.noreply.github.com>
Sibling site missed by PR #54022 — /api/console WebSocket in
HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401
close path as the PTY and events WebSockets. Without this guard,
opening the console after a dashboard restart shows 'Console closed
(4401). auth: token_mismatch' with no recovery.
react-router v7 exports MemoryRouter from 'react-router', not
'react-router-dom'. The test was written when the repo still imported
from 'react-router-dom' (4000+ commits ago).