Commit Graph

46022 Commits

Author SHA1 Message Date
Mohamad Kanso
e21291663f fix(desktop): aggregate registered gateway sessions 2026-09-28 20:48:58 -05:00
John Paul Soliva
c129d9897b fix(desktop): Maintenance tails a second run of the same op
The Maintenance panel's only log poller was an effect keyed on the
action name. The backend spawns each op under one fixed name ('doctor',
'security-audit', 'backup', 'curator-run'), so running the same op a
second time set the name state to the value it already held. React
skipped the effect and no poll started. launch() had already cleared the
status, which hid the log panel and re-enabled every op button while the
new run was still going.

Measured with the real panel against a fake backend that keeps the
fixed-name contract: after a second Run doctor, getActionStatus had been
called once in total (the first run's poll), nothing was rendered and
the button stayed enabled. The status response for the second run was
never fetched, so its rail task was never upserted either.

Key the tail on a fresh object per launch instead of the bare name, so
every launch restarts the poll and cancels the previous one. With the
fix the second run is polled, its output and the running indicator
show, the button is disabled, and the rail task holds the new pid.
2026-09-28 20:41:54 -05:00
chelsealong
8b228c37d5 fix(desktop): stop trusting composer previewUrl as a full-res upload cache
readImageForRemoteAttach reused attachment.previewUrl as if it always held
the full-resolution bytes for remote image uploads, but attachImagePath
drops previewUrl the moment a thumbnail exists — the only large data URL
kept afterward is the bounded (<=512px) thumbnailUrl. The cache lookup was
therefore always a miss in normal usage and the fallback disk read masked
it, but the contract was false and any future change that populated
previewUrl with the retained thumbnail would silently upload a downscaled
image with no error (#93324).

Removes the dead cache-reuse path and always reads the image fresh from
disk for remote/cross-filesystem uploads, matching the producer's actual
guarantee.
2026-09-28 20:41:44 -05:00
Adolanium
7c15d6994d fix(desktop): comment mode redacts the whole picked tree and keeps nested groups together
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-28 20:37:46 -05:00
beardthelion
399bf7938f fix(desktop): resolve the serve subcommand positionally when a profile is named "serve"
dashboardFallbackArgs rewrote args.indexOf('serve'), the first token equal
to 'serve'. For a profile literally named 'serve' that is the --profile
value, so the fallback corrupted the profile and left the real serve
subcommand in place for a runtime that cannot parse it. The remote
ownership verifiers had the same collision: a remote dashboard spawned
under a 'serve' profile carried two 'serve' tokens in its cmdline and was
judged foreign, so stale cleanup refused to reap it.

Both now locate the subcommand positionally, skipping the value consumed
by -m/--profile/-p, and the remote verifiers count only non-value 'serve'
tokens.
2026-09-28 20:37:18 -05:00
Xipong
43d2f8493e fix(desktop): bound Preview console string retention 2026-09-28 20:30:16 -05:00
liuhao1024
805ddfba5c fix(desktop): tolerate a throwing disposer so a failed plugin load cannot wedge the registry
A disposer with its own bug aborted the bare forEach in both
unloadRuntimePlugin() and failRegistration(), stranding the
loaded.delete() behind it. The registration stayed held, so every
later reload re-ran the same broken disposers and died before the
fresh register() — file edits looked inert until an app restart
(#126338).

Release the record first and run each disposer in its own try/catch;
failRegistration() now reuses unloadRuntimePlugin() so the rollback
path is tolerant the same way.
2026-09-28 20:29:49 -05:00
finn763
e801147cc2 fix(desktop): surface dead client wake capture instead of staying deaf forever
Client wake capture had no health monitoring: after a platform capture
failure (macOS PLAS IPC delegate error -> StopSourceOnError) getUserMedia
stays resolved but the PCM is dead (ended track, halted callbacks, or
endless zeros). The ear kept showing listening while the gateway detector
could never fire, and wake.feed refusals plus feed RPC errors were
swallowed (no onError wired). Add track-ended/stall/silence watchdogs and
consecutive feed-refusal escalation; wire onError to an honest off state
with the reason and lease release. Closes #119089
2026-09-28 20:29:19 -05:00
rodricksz4h5
fd0b98f625 fix(desktop): stop background polls from out-dialing the reconnect ladder
Both request paths into a non-active profile - gatewayForProfile (a plain
local profile) and requestGatewayForAgent (a registry route) - dialed
whenever the socket was down, straight past scheduleReconnect's full-jitter
ladder and stable-open rule. openSecondary only coalesces concurrent dials,
so a background poller against a backend that accepts and then drops the
socket redialed once per poll: the reconnect loop and renderer CPU spin in
#121865. On the local path every attempt was two dials, because
sharedPrimaryRoute probes main before the secondary dials.

Keep a per-scope dial-failure record (last failure, streak). Inside a window
that starts at 300ms and doubles to the ladder's 15s cap, background callers
fail fast and leave redialing to the ladder; the throttle sits ahead of the
sharedPrimaryRoute probe. Foreground callers always dial and a first dial is
never blocked. A socket that opens and dies before RECONNECT_STABLE_OPEN_MS
counts as a failure and a bare open does not clear the record, so an
accept-then-close backend keeps escalating.

The record lives in module state keyed by scope, like reauthFailures: a
background request lease disposes its entry when the request fails, so
per-entry history reset on every poll. It clears on a served RPC or a socket
that lived, and on a foreground re-arm, connection removal or full teardown.
Auth rejections are left to reauthFailures.
2026-09-28 20:18:01 -05:00
Yingliang Zhang
b8b38e55dd fix(state): handoff rows keep the durable parent timestamp for lineage dedupe
Rotation re-inserts the compressor's handoff tail as fresh child rows. Carried
tail dicts arrive without a timestamp, so _insert_message_rows stamped them
with publish-time now(); the parent's durable originals kept their own
timestamps. The two copies then differ in _display_dedupe_key
(role, content, timestamp, tool fields) and _dedupe_display_generations shows
both -> duplicated rows in the ancestor-lineage display read (#59661), once
per segment boundary in a long chain.

publish_compression_child now adopts the durable parent row's timestamp onto
carried rows before insert (best-effort, never overwrites a timestamp the dict
already carries). archive_and_compact gets the same carry for in-place double
compaction, and _insert_message_rows back-stamps the persisted timestamp so
later rotations inherit it. The existing dedupe key and every read path are
unchanged; the collapse happens because identities now agree.
2026-09-28 20:17:37 -05:00
Yingliang Zhang
cfd5dc3876 fix(api): scope ancestor reads to display surfaces; guard after_id; stamp read-path fixture
Quad-review follow-up (4 legs ACCEPT-WITH-NITS, P0s independently reproduced):

- _conversation_history_for_session reverts to tip-only: it feeds
  conversation_history for session chat / chat stream / v1 runs (model-fed
  restore must not regrow what compaction summarized away); the
  include_ancestors=True expansion stays on the DISPLAY reads
  (_handle_session_messages + web router). Restores the delegation delivery
  contract test (probe pinned the model history to the tip segment).
- get_messages now raises ValueError for after_id + include_ancestors
  instead of silently ignoring the cursor on a merged-lineage read.
- test_session_db_read_path_split fixture stamps the parent
  end_reason='compression' the way a real rotation (publish_compression_child)
  does; the old unverified parent link is a shape no production writer
  creates and the verified walker correctly rejects.

Verification: test_session_db_read_path_split 9 passed;
test_api_delegation_delivery_contract 2 passed; test_hermes_state.py 270
passed / 1 pre-existing FTS5 base failure (fails identically on 6e07eb4838);
tests/hermes_state/ 310 passed 3 skipped; gateway session-api +
compaction-projection + read-path 44 passed; resume-display + conversation-root
+ tui resume 32 passed; ruff + py_compile + git diff --check clean.
2026-09-28 20:17:37 -05:00
Yingliang Zhang
e5867d9bae fix(state): verified compression lineage for ancestor reads + fork marker
- _resume_lineage_ids walks verified compression links only (each parent ended
  'compression', child not an explicit fork), so API forks, reset children, and
  subagents stay single-session — fixing the [assistant, user, assistant]
  duplicated-answer history ehz0ah reproduced on API forks (#59661 review).
- _handle_fork_session writes _branched_from like the CLI/TUI branch paths.
- _handle_session_messages passes include_ancestors=True (REST session API
  half of the original fix, lost in a later squash).
2026-09-28 20:17:37 -05:00
Yingliang Zhang
cf63a03e39 fix(rest): include compression-ancestor messages in /api/sessions/{id}/messages
Port of #59661 onto post-#102117 main. The refactor moved get_messages
into hermes_state_messages.py (SessionMessagesMixin) and made the REST
endpoint resolve the resume tip first, so the port adds the
include_ancestors kwarg to the new home — reusing _resume_lineage_ids
(same branch-session exemption as get_messages_as_conversation) and the
IN-clause pattern already used by _fetch_conversation_rows — and passes
include_ancestors=True at the REST messages endpoint and the gateway
API server's history loader.

After a compression rotation the transcript spans parent sessions; the
REST endpoint previously returned only the child continuation, hiding
the pre-compaction transcript while session.resume showed it (#51058).

Tests: 3 new tests in tests/test_hermes_state.py (lineage expansion,
explicit-branch exemption, merged-set paging); 34/34 lineage/branch/resume
suite green.
2026-09-28 20:17:37 -05:00
unsupportedpastels
31c5d57ec6 docs(whatsapp): the bridge runs in its own session, not the gateway's process group
The #127047 comment said a supervisor's stop signal reaches the bridge
through the shared process group. The bridge is spawned with
start_new_session=True; what reaches it is a container-wide broadcast
such as s6-overlay's stage-3 SIGTERM.
2026-09-28 21:06:08 -04:00
unsupportedpastels
99b24916ce fix(photon): don't read a signal-raced sidecar exit as a fatal crash
Same race as the WhatsApp bridge (#127047): a container/supervisor stop
signals the whole process tree, so the Photon sidecar (its own session)
can die before the gateway's stop flow reaches disconnect(), while
_inbound_running is still True. _supervise_sidecar() then reported
SIDECAR_CRASHED and queued a reconnect during shutdown.

Consult the runner's _stop_requested_by_signal, which the signal handler
sets before any stop work runs. A sidecar exit while the gateway keeps
running is unchanged: still a retryable fatal error.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-28 21:06:08 -04:00
liuhao1024
f5a3cd6bd8 fix(whatsapp): don't read a signal-raced bridge exit as a fatal crash
A supervisor's stop signal (docker stop, s6-supervise, systemd) reaches the
bridge child in the same process group before the gateway's stop flow reaches
WhatsAppAdapter.disconnect(), so _shutting_down is still False when the poll
loop observes the -15 exit. _check_managed_bridge_exit() then reports a
retryable fatal adapter error, the stranded-platform check exits the gateway
with failure, and the supervisor restarts it into a crash loop (#127047).

Also consult the runner's _stop_requested_by_signal, which the signal handler
flips before any stop work runs. A bridge SIGTERM/SIGINT while the gateway
keeps running is unchanged: still a retryable fatal error so the reconnect
watcher revives it.

(cherry picked from commit 252ae311f06e402f9a8ef1482c87aaefe8d9ac39)
2026-09-28 21:06:08 -04:00
Austin Pickett
835501f92f fix(desktop): keep a streamed reply over an empty committed tool shell
A turn whose store rows end in a tool round with content='' (no prose row to
fold into it) hydrates as a text-less assistant shell. preserveLocalPendingTurnMessages
paired the settled live bubble with that shell and dropped it, because
localPendingSupersedes only accepts live-tail authoritative rows, deleting the
answer the user just watched stream.

Keep a settled, non-interim local final with answer text over an error-free
committed shell with no answer text. Resume reconciliation keeps the stricter
localPendingSupersedes rule; a committed prose row still folds into the shell and
wins, so no duplicate.

Based on #123065 by @kvnloo.

Fixes #123047
2026-09-28 21:05:52 -04:00
Austin Pickett
e050902e7c fix(desktop): keep the NVIDIA proprietary driver on XWayland by default
The Wayland launcher (#121377) passes --ozone-platform=wayland on every
local Wayland session. On the NVIDIA proprietary driver (615.x) the GPU
process dies under Wayland ozone with the bundled Chromium, while x11
launches. Default to x11 when /proc/driver/nvidia/version exists and the
user set no platform or x11/wayland hint; explicit choices still win.

Fixes #126013
2026-09-28 21:05:03 -04:00
Austin Pickett
e33f1a0faf fix(desktop): refuse window-owned pane requests at once when no window shows the session
preview.read / preview.act / terminal.read / window.read / tour are answered
only by the Desktop window that shows the requesting session (#114981). When
no attached window showed it, every window stayed silent and the agent waited
out the full 30-45s bridge deadline, then got the generic "No preview tab is
open, or the read timed out."

A window not hosting the session now declines with JSON-RPC error 4404
instead of staying silent. The gateway counts a decline as that client's vote,
not as the answer: the request stays open for the owner window and settles
only once every attached answering client declined, with a distinct refusal
("No Hermes Desktop window is showing this chat ... Ask the user to open this
chat in the Desktop app, then retry."). A bystander window therefore still
cannot beat the owner (#113348).

The backend advertises the counting via client.capabilities
(declines_not_shown); the shared channel only sends a decline to a backend
that did, so a new app against an older backend keeps today's silence.

Fixes #119333
2026-09-28 21:04:47 -04:00
Austin Pickett
0f50687327 fix(update): report a symlinked other-install profile gateway as external, not unaccounted
A profiles/<name> symlinked to a separate Hermes install is inventoried by the
update plan, left alone by the restart phase, and already classified
'external' by the post-restart fleet matrix. Reconciliation had no such
outcome, so the gateway ended 'unaccounted' and every hermes update exited 1.

Feed the fleet's verified external gateway pids into match_runtime_outcomes:
an untouched gateway in that set is 'external', reported informationally and
excluded from the exit-1 tripwire. No fleet evidence keeps it unaccounted.

Fixes #120240
2026-09-28 21:04:13 -04:00
Austin Pickett
bbdf0a34fc fix(desktop): keep clarify/connection cards open for /bg side tasks
Generalize the /btw exemption to every registry action that runs beside
the live turn (/btw, /bg, /background) via isSideTaskSlashCommand, keep
the ordinary-message behavior when attachments ride along, and fix the
gateway mock cast so desktop typecheck passes.

Fixes #126188

Co-authored-by: Charan Rathore <charan-rathore@users.noreply.github.com>
2026-09-28 21:03:58 -04:00
Charan Rathore
06cc14bca3 fix(desktop): preserve connection request for /btw 2026-09-28 21:03:58 -04:00
Charan Rathore
3a2a626821 fix(desktop): preserve clarify when asking /btw 2026-09-28 21:03:58 -04:00
Austin Pickett
f0410eb4f0 fix(desktop): drop voice-conversation barge captures that echo the spoken reply
Over speakers, the reply bleeds into the mic, trips the renderer's
playback-phase barge monitor, and the transcript was submitted as a user
turn (with the "user interrupted" note). Port the CLI's is_tts_echo rule
(difflib ratio >= 0.6 + transcript-sized sliding window) as a pure TS
helper and apply it in submitCapturedUtterance: a playback-phase capture
matching the reply being spoken is treated as silence — no submit, the
interruption latch is cleared, and the loop resumes listening.

Covers the TTS-echo half of #126708; the client-direct STT hallucination
filter and the Desktop voice.barge_in pref are untouched.
2026-09-28 21:03:40 -04:00
Austin Pickett
3881d85dc7 fix(desktop): a redirect inside a regenerated turn keeps the partial reply on screen
A mid-turn correction sealed the live bubble but left the regenerate's
pendingBranchGroup armed, so the post-correction reply joined the branch
group. The runtime repository parents every group member to the group's
user row, which made that reply a sibling of the sealed partial: the
partial and the correction dropped off the rendered branch, and
applyBranchVisibility then marked the partial hidden in the store.

appendMidTurnUserMessage (shared by the main composer and session tiles)
now ends the branch group at the correction.

Covers the regenerate half of #119015; a plain-turn redirect already kept
the partial on main (steer-arrival-order suite). Related to #119686.
2026-09-28 21:03:34 -04:00
Austin Pickett
6397a84a91 fix(sessions): un-hide an auto-archived chat when it is resumed or compressed
The idle sweep archives a whole compression lineage. When the chat later
resumed and compressed, the new tip was inserted with archived=0 under the
archived root, and the sidebar admits a lineage by its root's flag, so the
live chat stayed hidden.

Record sweep provenance in a new sessions.auto_archived column. Publishing a
compression child or reopening a session under a sweep-only archive
un-hides the lineage; a deliberate archive (sidebar, API, CLI) clears the
provenance and keeps the chat hidden. Compression children now inherit the
parent's archive state so a manually archived lineage stays uniform.

Fixes #117713
2026-09-28 21:03:09 -04:00
Austin Pickett
0c266b5150 fix(desktop): skip macOS SSH remotes with a live serve in update-all
The POSIX managed-update drain only signals the Desktop-owned serve through
a pidfd so the signal is bound to the verified process. Darwin has no
equivalent, so terminateOwnedDashboardForUpdate deliberately refuses there
rather than accept a PID-reuse window. The update lifecycle still cycled
every forward first, hit the refusal, and reported the row as a failed
update ("Update fan-out failed") on every attempt.

Detect a live Desktop-owned Darwin serve right after scope capture, before
any forward or serve is touched, and return a structured refusal with a
skip reason. "Update all instances" now reports that connection as skipped
with a clear per-remote message; macOS remotes with no live Desktop serve
need no drain and still update, and the other rows are unaffected.

Fixes #124617
2026-09-28 21:02:45 -04:00
Austin Pickett
99bc0d1842 fix(desktop): bound managed SSH durable recovery when a scope never restores
Durable recovery cleared its journal only when every scope restored, and the
update gate fences dials, edits and removal while a journal record exists. One
scope that never restores therefore wedged the SSH connection forever.

Once the remote install marker is positively clear, recovery now ends when the
correlated remote exit code is 0, or after MAX_MANAGED_SSH_RECOVERY_ATTEMPTS
failed attempts (persisted as `attempts` on the record). The journal record is
cleared, the gate stops fencing, and the unrestored profiles are logged; the
next ordinary dial starts them again. The in-session update path clears the
journal the same way after a proven-successful update.

Fixes #107827
2026-09-28 21:02:45 -04:00
Austin Pickett
d619b67ec2 fix(local-runtime): spill only the recurrent FFN blocks a hybrid needs
A spilled hybrid got `-ot blk\.\d+\.ffn_.*\.weight=CPU`, which moves every
block's FFN — full-attention blocks included — regardless of how many bytes
actually had to leave the GPU. The GGUF reader now records FFN weight bytes
per block, and spill_overrides picks the fewest recurrent (n_head_kv==0)
blocks whose FFNs cover the decision's spill_bytes, emitting them as an
explicit `blk\.(i|j|...)\.ffn_.*\.weight=CPU` selector. Unknown sizes fall
back to every recurrent block, never a full-attention one.

Based on #120778 by @kvnloo (recurrent-index selector).

Covers the placement half of #113329; the overhead/logits calibration
(RUNTIME_OVERHEAD_BYTES, ub_logits_bytes) is untouched pending measured data.
Refs #113329
2026-09-28 21:02:39 -04:00
Austin Pickett
f6df13bac6 fix(update): keep never-pushed and unmerged branches pinned in source check
An empty ls-remote advertisement cannot tell "merged and deleted upstream"
from "never pushed". The source check now re-pins a missing branch to main
only when the local branch was once published (remote-tracking ref or a
configured upstream) and git cherry finds no commits missing from main.
Otherwise it keeps the branch and reports branch-local-only.

Diagnosis from #105053 by @0xalydev and #105055 by @kokhlo, which targeted
the Electron healer since moved into hermes_cli/source_check.py.

Fixes #105042
2026-09-28 21:02:35 -04:00
Austin Pickett
90b0b21537 fix(update): record Desktop's correlation_id in the update receipt
Desktop's managed SSH update exports HERMES_UPDATE_CORRELATION_ID and
accepts only a receipt whose correlation_id equals it. The backend never
read the variable, so every successful remote update was reported as
failed and restored. Persist it as correlation_id when the receipt is
created, and keep it across the interpreter handoff.

Idea from @LordMelkor on the issue.

Fixes #101516
2026-09-28 21:02:29 -04:00
Austin Pickett
449fae030a perf(desktop): share thread-message position lookups across transcript rows
Every mounted UserMessage and assistant row re-ran a full scan of
thread.messages on each streamed chunk (runtimeUserOrdinal,
interAgentSender), so per-chunk cost grew with rows x transcript length.
Build one id->index / id->user-ordinal map per messages array and share
it across rows. At 400 turns this cuts per-chunk time from ~42 ms to
~30 ms in jsdom; 20 turns unchanged.

Covers the per-row transcript-scan part of #126486. Prior art: #69151.
2026-09-28 20:23:15 -04:00
Austin Pickett
05d832d685 fix(desktop): only the last holder exits a shared SSH ControlMaster
Every SshConnection for one scope/host/identity hashes to the same
ControlPath, and ControlMaster=auto attaches later connections to the
master already listening there. A stale bootstrap attempt (timed out,
superseded, rolled back) that closed late ran `ssh -O exit` on that
socket and tore down the master its successor had attached to, killing
the live backend's forward after "Remote Hermes backend is ready".

Connections now claim their ControlPath from the moment open() starts
dialing and release it on close or failed open; close() runs -O exit
(and the socket-unlink fallback) only when no other connection still
holds the socket. No-mux (Windows) is unchanged.

Covers the stale-attempt teardown half of #97264; the hardcoded 45s
ready timeout is untouched (tracked in #94642 / #94665).

Refs #97264
2026-09-28 20:21:57 -04:00
Austin Pickett
81f481b2db fix(agent): carry a custom provider's extra_body into TUI/Desktop agents and aux calls
TUI/Desktop `_make_agent` (plus `hermes -z` and ACP) built AIAgent without the
resolver's `request_overrides`, so a custom entry's `extra_body` only reached
the agent through the init-time re-derivation, which matches `custom:*` alone.
Auxiliary requests never read the entry's `extra_body` at all. A proxy that
requires a body field (`user`) therefore 400'd in the app and on smart
approval while `hermes chat` worked.

- Pass `runtime["request_overrides"]` at every remaining AIAgent build site
  (ACP only when no explicit base_url points at another endpoint).
- `_build_call_kwargs` layers the destination entry's `extra_body` under the
  task/caller body, using the agent's own matcher, so primary, retry and
  fallback requests each get their own endpoint's body and nothing leaks.

Covers holes 1 and 3 of #103738; hole 2 (matching bare/unprefixed names) is
#124290 and is untouched here.
2026-09-28 19:49:00 -04:00
Austin Pickett
bf84561b9d fix(tui_gateway): keep a container backend's terminal.cwd as the session workspace
Under a container terminal backend (docker/modal/...), terminal.cwd names a path
inside the sandbox (/workspace). _completion_cwd host-validated it with
os.path.isdir, failed, and fell back to the gateway's own cwd ($HOME for the
desktop). Every host file then looked "inside the workspace", so file.attach
never staged it into the bind-mounted attachments/ dir and the agent was handed
a host path that does not exist in the container.

- _completion_cwd keeps an absolute container cwd unvalidated when the bound
  backend is non-local (matching _terminal_task_cwd), including a named
  profile's own container terminal.cwd and a config-only (unbridged) backend.
- _session_is_local_backend reads the effective backend (env or config), so a
  container cwd is never "healed" to its nearest host ancestor (/).
- map_cache_path_to_container also matches through a symlinked HERMES_HOME:
  @file: expansion passes resolved paths, the mount roots keep the configured
  spelling, so staged attachments leaked their host path.

Covers the desktop attachment path of #103147. Typed @file: refs to binary
files outside the mounted dirs still render the host path.
2026-09-28 19:41:57 -04:00
Austin Pickett
b30bbe2f53 fix(desktop,i18n): regenerate _keys.desktop.json for the two keys main added
commandCenter.toggleBrowser (dbc5ec8f9b) and zones.zoneMenuLabel (ad2d4822e1)
landed in en.ts without the generated key index, so the i18n-keys test in
JS & TS checks is red on main and on every desktop PR.
2026-09-28 19:37:47 -04:00
Austin Pickett
5d40f8e95d docs(website): move Terms/Privacy links from catalog footers into the hero
Legal links at the very bottom of the Skills Hub and Plugin Catalog were
easy to miss. Show 'Terms • Privacy Policy' in the hero instead, linking
to the same portal.nousresearch.com pages, and drop the now-unused
legalNote footer markup and CSS.
2026-09-28 19:31:44 -04:00
xxxigm
29e3228a42 test(desktop): pin a Bot Mode side chat as a listed session
Drives `openNewSessionTile` through the exact shape "New chat with this bot"
and the Bot Mode tab-strip "+" use — a bots workspace scope with the bot's
own route — and asserts `session.create` carries no `hidden` param.

A contract, not a snapshot: it says a hand-started conversation in a bot's
profile is an ordinary listed session, which is what makes it findable,
resumable, renameable and deletable from the Sessions sidebar and `/resume`.
Red on the unfixed tree (`hidden: true`), green with the blanket flag gone.
2026-09-28 18:25:40 -05:00
xxxigm
a109601657 fix(desktop): keep Bot Mode side chats in the Sessions sidebar
`openNewSessionTile` stamped `hidden: true` on every session created while
the workspace was in Bot Mode. That path only ever mints side chats — "New
chat with this bot" and the Bot Mode tab-strip "+" / ⌘T — so each one was
born unlisted and stranded: absent from the Sessions sidebar, skipped by
`/resume`, and unreachable once its tab closed, because the bot row opens the
canonical chat and "Open recent session" reads `last_session`, which never
reports a hidden row. Users had no way to find, rename, or delete a bot
conversation they started by hand.

Only Bot Mode's PLUMBING sessions are meant to be hidden, and each already
mints its own row with the flag set: the canonical Bot Chat in
`hermes-bots/canonical-chat.ts` and group member sessions in
`hermes-bots/group-turns.ts`. The hide sweep agrees — its title allow-list is
deliberately exact so "a user's real conversation inside a bot profile ...
is never touched" (`hermes-bots/session-sweep.ts`) — and so does
`apps/desktop/src/AGENTS.md`: side chats "stay visible in the sidebar".

Drop the blanket flag. The sidebar concern it was reaching for is already
handled where it belongs: `mergeSessionPage` and the refresh keep-list in
`store/session.ts` drop `hidden` rows, so a canonical Bot Chat cannot be
resurrected into the list. `listTileSessionRow` keeps its Bot Mode guard, so
a side chat surfaces on the next refresh once its first turn persists,
exactly like a Sessions-mode ⌘T tab.
2026-09-28 18:25:40 -05:00
Austin Pickett
7154128fe1 fix(tui-gateway): answer gateway.ping while a WS handler blocks dispatch
The WS read loop awaited each inline dispatch() before reading the next
frame, so a handler stuck behind a long compaction (or any lock/GIL/DB
stall) left gateway.ping unread. The desktop's 45s heartbeat deadline
then tore down a busy but healthy backend and killed the turn.

Dispatch now runs on a per-connection task that keeps the serial arrival
order; the reader answers pings directly. Teardown still waits for the
in-flight handler and frames read before the disconnect.

Fixes #108325
2026-09-28 18:58:29 -04:00
kshitijk4poor
1fd412e8a0 test(sessions): drop the SCHEMA_VERSION import the trigram test no longer reads 2026-09-29 04:15:21 +05:30
Eva
19d54c0611 fix(gateway): replay keeps message identity on plain user and assistant rows
_build_gateway_agent_history rebuilds a session's own transcript for its
next turn. Rows that pass through whole (tool-calling assistants, tool
results) kept message_uid and the merge witness, but _build_replay_entry
rebuilt plain user and assistant rows from an allowlist that left both out.
A context engine therefore saw the store's uids on some roles and none on
others in the same replayed conversation.

The replay entry now copies message_uid and the merge witness on every
role. It leaves out the tool-call uid maps: a plain row whose unanswered
tool_calls repair pruned can still hold a map, and replaying it would name
calls the row no longer carries.
2026-09-29 04:15:21 +05:30
kshitijk4poor
0cabdf3f64 fix(sessions): an assistant fold ignores a non-dict uid map on the surviving turn too
The absorbed turn's map already had the dict guard; a plugin-built survivor carrying a malformed
map still raised in the per-occurrence expansion. Pre-existing (the old merge raised the same way).
2026-09-29 04:15:21 +05:30
kshitijk4poor
536a2611c9 refactor(sessions): record_absorbed_message dedupes with dict.fromkeys
No behaviour change.
2026-09-29 04:15:21 +05:30
kshitijk4poor
dd66c132f2 fix(sessions): keep tool-call uids paired through rewrites, folds of repeated ids and chunked copies
Found by the pre-merge correctness and hermes-specific review of this stack:
- A fold after a row whose one response repeated a provider id (one shared uid) appended the
  absorbed turn's uid in the wrong slot. The shared uid now fills each of the row's own
  occurrences before the fold appends (per_occurrence_tool_call_uids). A single response keeps one
  shared uid for a repeated id: every result carries it, so no call looks unanswered.
- A digest-less re-flush of a restored fold (replay heal, adopt after a digest mismatch) let the
  stored pre-fold map overwrite the longer live list, so the later occurrence's results paired
  with nothing. A live list that starts with every stored occurrence is kept (stored uid or list).
- A dict that lost its uid map (a clone) rewrote tool_call_uids to NULL. The row-addressed
  rewrite now fills stored uids for the calls the dict still names; a call it dropped keeps none.
- append_messages_batch(chunk_rows=...) reset the pairing index per chunk, so a result in the chunk
  after its call was stored without tool_call_uid (branch/seed copies). The whole batch is paired
  before it is chunked.
- _merge_assistant_into keeps its dict guard on the absorbed turn's map (a plugin-built dict with a
  non-dict map no longer raises).
2026-09-29 04:15:21 +05:30
kshitijk4poor
aac4d5812b chore(sessions): label the v31 schema-history event with a commit on this branch
20d67ffb5e was a pre-rebase SHA that exists nowhere; eed37d63ce is the commit that adds the
columns here. As with 1c6683e8e0 / d9e64e9165, the label follows the landed SHA once merged.
2026-09-29 04:15:21 +05:30
kshitijk4poor
2be9fa5f9d refactor(sessions): simplify-code pass over the identity helpers
- without_persistence_fields() replaces the two hand-written strips (context-engine selection,
  token-estimate fingerprint).
- mint_uid() is the one uid minter; uid_list() the one uid-list normalizer (message_metadata's
  _absorbed_uids and the codec's _uid_list share it).
- _live_or_column() replaces four copies of 'live key first, column name second'; the absorbed
  list writer takes the dict, not a (live_key, column) pair its one caller always passed the same.
- _restore_row_identity() is the one stored-identity-onto-live-dict step, shared by restore and
  the row-addressed rewrite (transcript_repair had a hand copy). Restore skips the decoders for
  NULL columns (the common row).
- index_tool_call_uids() returns the variants it named, so the flush walk stops computing them
  twice; tool_call_uid_from_history drops guards its one caller can't trip.
- rewind returns the replacement dict (the insert already stamps _row_id and message_uid)
  instead of re-querying last_insert_rowid.
- _remember_absorbed_row's folded is required, so every call site states the witness rule.
- The backfill reads/deletes state_meta through _meta_row/_delete_meta; chat_completions strips
  MESSAGE_UID by constant.
No behaviour change.
2026-09-29 04:15:21 +05:30
kshitijk4poor
4808a2bff1 refactor(sessions): move the message-identity column codec to hermes_state_identity
hermes_state_messages.py crossed 2000 lines with the v31 identity columns. The codec between
the live identity keys and the four columns (_uid_list/_uid_map and their JSON writers, the
restore decoder) moves to its own sibling; transcript repair imports _uid_map from there.
_json_or moves to hermes_state_common so both modules share it (same 'hermes_state' logger).
No behaviour change.
2026-09-29 04:15:21 +05:30
kshitijk4poor
5d8b17f865 fix(sessions): an assistant fold that discards the later text records no merge witness
_merge_assistant_into never joins multimodal (list) content, so a non-empty later turn
beside a list (or a list beside non-empty text) is dropped, yet the fold still recorded its
uid in _absorbed_message_uids, a false claim that its text lives on. The merge now reports
whether the later text was kept and the witness follows it; the retired row id is still
recorded.

Reported by @andrexibiza on #126307.
2026-09-29 04:15:21 +05:30
kshitijk4poor
d9808867c7 fix(sessions): an assistant fold keeps every occurrence uid of a reused provider id
Folding two assistant turns that reuse a provider id (llama.cpp-style constant ids) leaves
both calls in tool_calls, but the {id: uid} union kept only the later uid, so the first
occurrence lost its identity. The id now maps to a list of uids, one per occurrence in
tool_calls order (merge_tool_call_uids); index_tool_call_uids pairs each call with its own
occurrence, so a following result still pairs with the nearest call. Provider ids are never
rewritten. Persisted and restored as-is (the JSON map accepts list values).

Reported by @andrexibiza on #126307.
2026-09-29 04:15:21 +05:30