Commit Graph

11694 Commits

Author SHA1 Message Date
Nicolas Formenton
c75d835555 feat(desktop): mark a session as unread/read with a persisted watermark 2026-08-15 01:21:40 -07:00
Moisés Valero
fef9c537d7 fix(cli): convert Alt key shortcuts to sequence tuple for prompt_toolkit (#74169) 2026-08-15 01:05:39 -07:00
John Lussier
b3447c2129 test: align submit bindings with multiline default 2026-08-15 01:05:39 -07:00
John Lussier
2ae7884ffa fix: make CLI multiline shortcuts work by default 2026-08-15 01:05:39 -07:00
doncazper
6a375a24a5 fix(gateway): timestamp shared stderr output 2026-08-15 01:04:19 -07:00
doncazper
1db9273584 fix(gateway): timestamp launchd error log lines 2026-08-15 01:04:19 -07:00
coe0718
7f7aefe5cb fix: restore complete message timestamp coverage 2026-08-15 01:04:19 -07:00
Tuck
03636ab33f test: add message_metadata unit tests 2026-08-15 01:04:19 -07:00
Teknium
f2a30fa400 test: derive lost-and-found synthetic width from the live schema
The mapper test pinned the sessions column count (54, then 55, then 56
within one week as git_metadata_generation and the hidden flag landed).
Every ordinary column addition broke it. Derive max_fields from
PRAGMA table_info at runtime with a >= floor so the test keeps asserting
the rebuild contract without change-detecting the schema width.
2026-08-15 00:55:20 -07:00
Paul BlackSwan
d5a865882a fix(desktop): save remote gateway files natively 2026-08-15 00:37:00 -07:00
Teknium
46fe44fa8d test(cli): stub portable-MCP lookup in completer read test; bound resolve_toolset memo
- The readonly-loader completer test now stubs
  get_portable_mcp_server_names_nowait — real plugin discovery runs
  load_config() during one-time process init, which is not the
  per-keystroke read the test guards against.
- Cap _resolve_toolset_memo at 256 entries: generation-keyed entries
  from stale generations are never hit again, so clear on overflow to
  keep long sessions bounded.
2026-08-15 00:36:03 -07:00
jackoconner55
2f54ad4023 perf(cli): skip launcher-side plugin discovery for TUI handoff
(cherry picked from commit e2e0edd6b8ec10d02e867ff11dd1f9b4961a1ba7)
2026-08-15 00:36:03 -07:00
spfcraze
b9f7525a1b test(agent): pin measured-work regression for display-flag config cache
(cherry picked from commit 57a7044c20965150e74b0ecadb2550f0ed81ee8e)
2026-08-15 00:36:03 -07:00
spfcraze
e25cafc83e perf(agent): cache per-turn display-flag config reads
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.

The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.

(cherry picked from commit f2d0e00d71ba35fa3a53cd41076261e3c5feb45d)
2026-08-15 00:36:03 -07:00
spfcraze
7b45d1d049 perf(toolsets): memoise resolve_toolset keyed on registry generation
resolve_toolset() recursively walks toolset includes and, with
include_registry=True, merges registry-registered tools on every call —
each external call re-runs the includes walk and takes a fresh registry
snapshot under the registry lock. It is called dozens of times per
_get_platform_tools() (every /tools completion keystroke, per picker
render) and at 19 call sites across the CLI/gateway.

Memoise the external-entry result keyed on (name, include_registry,
registry id, registry generation). tools.registry exposes a monotonic
_generation counter bumped on every register/deregister/alias/MCP
refresh (its docstring explicitly invites generation-keyed memoisation),
so a cache entry is valid until the registry changes. External callers
never pass visited, so the memo engages exactly at the public entry
and the internal cycle-detection recursion is untouched.

Measured: _get_platform_tools drops 165us -> 59us per call (3x) with the
xAI credential fix simulated; /tools completion ~2ms -> ~77us/keystroke
combined. Regression tests: repeat resolution is a memo hit (get_toolset
called once), a generation bump forces a fresh resolve, and the memoised
result is identical to a fresh resolution.

(cherry picked from commit 3d36ecb273f6652d00556a7b9d846c08047801aa)
2026-08-15 00:36:03 -07:00
spfcraze
47400fe2af test: make memo pins pre-fix-safe (raising=False resets)
(cherry picked from commit 4822daed5d9238348c50bfbdf8c3c4795adc4986)
2026-08-15 00:36:03 -07:00
spfcraze
4bd746c6e9 perf(cli): memoise default-hermes-root resolution and global auth-store read
get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).

_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.

Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).

Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).

(cherry picked from commit be348f32e5bd7479c26fabb652de549fb9c8a1e1)
2026-08-15 00:36:03 -07:00
spfcraze
473490a2e9 perf(cli): stop per-keystroke config re-reads in slash completers
The /tools and /personality completers run on every keystroke while the
user types those commands (complete_while_typing), and both re-read +
re-parse the full config on each keypress:

- _tools_completions called load_config() — the defensive deepcopy
  (~340us/call on cache hit) even though it only reads toolset enable
  state + MCP server names. Switched to load_config_readonly() (the
  perf(agent) #74322 pattern; this per-keystroke site was missed).
- _personality_completions called load_cli_config() — a full YAML parse
  + deep merge of the built-in defaults (~110us) — on every keystroke.
  Memoised keyed on the config file path+mtime (same pattern as load_env
  / _nous_auth_status_cache), so the parse runs once per config state.

Measured: /tools 357us -> 18us per keystroke; /personality parse drops
from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse).

Regression tests: _tools_completions uses the readonly loader (deepcopy
loader never called); personality memo parses once across repeated
completions and re-parses once after a config mtime change.

(cherry picked from commit 2b3f897171f93dc6b099848ec5ab3763b7294870)
2026-08-15 00:36:03 -07:00
Adolanium
ee1731b13c perf(cli): re-export decomposed command modules lazily, ~60ms off every CLI start
The main.py decomposition re-exported the sessions/update/dashboard command
surface with eager from-imports, so every hermes invocation (including
hermes --version) paid for update_cmd's dependency chain (jwt, click,
cryptography). Resolve the re-exports through the existing PEP 562 module
__getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on
first actual use. Internal call sites go through a _self() helper because
bare-name lookups do not trigger __getattr__; _self() imports sys locally
since update tests patch hermes_cli.main.sys. The
_warn_stale_dashboard_processes back-compat alias moves into the lazy
surface, and the sessions argparse dispatch defers sessions_cmd to call
time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a
real module attribute, which shadows __getattr__.

Measured (Windows 11, Python 3.11, median of 7 warm runs):
import hermes_cli.main 253ms -> 196ms (-22%).

(cherry picked from commit cad1083b71635f98815b698d69a18f7f58e15517)
2026-08-15 00:36:03 -07:00
DannyFengTianYu
07b9090259 fix(desktop): preserve complete history when branching
Read the durable display transcript when creating a branch instead of copying the compacted model projection. Hydrate the Desktop branch boundary from persisted history, avoid stale whole-chat counts, and seed the new tile from the backend snapshot. Add regression coverage for compacted histories, visible-message counts, selected prefixes, and hydration races.

(cherry picked from commit c3d2d759ae104395adde215b2e293ccf8e895684)
2026-08-15 00:35:40 -07:00
DannyFengTianYu
b066f2b373 fix(history): isolate branch transcripts from parent updates
(cherry picked from commit 57c51cc401e867e0b315e051ba138d0b8f7a5f27)
2026-08-15 00:35:40 -07:00
jdgg777
6da30f72a2 fix(tui_gateway): fall back to session id when session_key is NULL in truncation persist
CLI-origin sessions have no session_key; the Desktop history-truncation
path called replace_messages(session["session_key"], ...) with None,
whose reinsert violated the messages.session_id FK -> "FOREIGN KEY
constraint failed" -> "Restore failed" on resume. Key the persist off
the durable session id instead.

Extracted from PR #81904 (the scope=compacted API half was superseded by
include_compacted, #86595). The PR's companion change defaulting
session_key to the session id at insert time is deliberately NOT taken:
main treats a non-NULL session_key as "this is a gateway session"
(list_gateway_sessions, orphan gateway-session repair), so the default
would misclassify every CLI session.

(extracted from PR #81904, commit cef9b9b27d)
2026-08-15 00:35:16 -07:00
lepetitprince716-prog
7e439dbb1b perf: parallelize provider model-list fetches in model picker
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.

Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
  that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
  concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
  with threading.Lock to prevent concurrent write races

Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)

Closes #80413

(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
2026-08-15 00:34:29 -07:00
joaomarcos
2162d583b1 perf(run-agent): reuse the Anthropic request-local client instead of rebuilding it per call
_create_request_anthropic_client() built a fresh anthropic.Anthropic
client (and httpx pool) on every single LLM call, and
_close_request_anthropic_client() always fully closed it right after
- unlike the OpenAI-wire path, which caches and reuses one warm
client across sequential calls via a single-slot cache keyed on the
effective client kwargs.

Add the same single-slot cache to the Anthropic-wire path: keyed on
credentials, base URL/Bedrock region, per-model timeout, and the
1M-beta flag; in_use guards concurrent calls from sharing one pool's
close/abort lifecycle; poisoned marks a cross-thread-aborted slot so
the owner-thread close discards it; reuse only on request_complete /
stream_request_complete (the same _REQUEST_CLIENT_REUSE_REASONS the
OpenAI path already uses). Wires a teardown hook into
release_clients()/close() mirroring _close_cached_request_openai_client.

Fixes #HPA-02

(cherry picked from commit 37f90df15593e6ded0390f827f8e5604bf0acc86)
2026-08-15 00:34:29 -07:00
blunkjamie-dev
4415f917b4 test(session-search): lock positional parameter prefix
(cherry picked from commit 73592200c69a4f0b6d7c290ce45832847df608e2)
2026-08-15 00:34:29 -07:00
blunkjamie-dev
163d7af310 fix(session-search): forward detail through agent paths
(cherry picked from commit 5f6de984f170ef470c7fbbd7662484bfaffc821d)
2026-08-15 00:34:29 -07:00
blunkjamie-dev
6e1bdc0a18 perf(session-search): adapt discovery result hydration
(cherry picked from commit 60a3530444f65c2cdcd4e5b983e4aa380ed651c7)
2026-08-15 00:34:29 -07:00
Adolanium
ee9ec6164c perf(state): stop selecting full message content in session search
Every search route in _search_messages_impl (FTS, CJK bigram, trigram,
LIKE fallback, rebuild-gap supplement) selected m.content, then the
result tail popped it unread. On DBs with multi-MB tool rows, each
search read and materialized up to `limit` full rows only to discard
them. Snippets come from snippet()/substr() in SQL and the context
window is re-fetched by id, so no code path ever read the column.

Drop the column from all six SELECT lists. Returned dicts are
unchanged: content was never part of the public result (the pop ran
before return), and tests/test_hermes_state.py already documents that
contract.

(cherry picked from commit d0c3af167e7dd4eb18e1bea29ba107a94911ea24)
2026-08-15 00:34:29 -07:00
briandevans
4c24629bc9 fix(console): skip the checkpoints prune confirmation the console already took
Hermes Console registers `checkpoints prune`, `clear` and `clear-legacy` as
mutating, so it takes a console-level confirmation before dispatching any of
them. `_apply_confirmed_defaults` then exists to keep the CLI layer from
asking a second time — its docstring says so — but it only force-defaults
`clear` and `clear-legacy`. `prune` was left out, even though `cmd_prune`
gates its orphan preview on the identical `not args.force` shape.

`_capture_output` redirects stdout and stderr but never stdin, so the
unskipped `_confirm()` call hits `input()` with no terminal behind it:
`EOFError` propagates into `_confirm`, which returns False, and `cmd_prune`
prints "Aborted." and returns 1. The console turns that non-zero exit into a
ConsoleCommandError, so `checkpoints prune` fails outright for any user who
has at least one orphan checkpoint project — after that user already
confirmed. When the server does happen to inherit a foreground terminal, the
same call instead blocks a console worker thread and eats the operator's
keystrokes.

Forcing the flag is the documented behavior here rather than a weakening of
the recent orphan-allowlist hardening. `orphan_allowlist` binds a deletion to
the identities shown in the preview, guarding the window where a workdir
disappears while the command waits on `input()`. Under the console there is
no preview and no wait, which is exactly the `--force` case the comment on
`cmd_prune` describes as "no restriction".
2026-08-15 00:33:32 -07:00
David Metcalfe
08f32a6335 fix(kanban): replace native browser dialogs with in-app ConfirmDialog
Migrates 8 of 12 native dialog call sites in the kanban dashboard plugin
to the SDK's ConfirmDialog primitive (added in PR #50550):
  - moveTask, moveSelected, applyBulk, deleteTask, deleteSelected,
    archiveBoard, removeAttachment, doPatch

The 4 remaining carve-outs (window.prompt for completion summary,
window.alert for missing summary, cli_hint clipboard fallback) are
documented inline — the host's ConfirmDialog hardcodes onClick → unmount,
preventing the keep-open-across-validation behavior the completion-summary
form needs. Followup: upstream a `disabled` prop to ConfirmDialog and
rebuild the completion body using host Dialog components.

New architecture:
  - useKanbanDialogs(t) — Promise-based dialog state machine at
    KanbanPage scope. request({kind, ...}) returns {confirmed, summary?}.
  - KanbanDialog component — renders ConfirmDialog from SDK for kind=confirm.
  - performMoveTask(taskId, newStatus, count, summary) — extracted shared
    dispatch path for single + bulk moves (optimistic UI + PATCH/POST +
    error recovery).
  - requestDialog prop threading — KanbanPage → BoardSwitcher,
    TaskDrawer → TaskDetail → doPatch/AttachmentsSection. Every call
    site has a defensive fallback to window.confirm if the prop is
    missing (verified by test_dashboard_done_actions_prompt_for_completion_summary
    counting the cancel guards + destructive:true markers in the bundle).

New host i18n keys (web/src/i18n/en.ts + types.ts):
  - kanban.confirmDoneMany / confirmArchiveMany / confirmBlockedMany
  - kanban.trash.confirmTitle / confirmManyTitle

Tests:
  - Replaced bundle-string-only completion-summary test with behavioral
    coverage: bundle cancel-guard count + destructive marker count, plus
    backend tests that confirm cancel preserves old status and confirm
    dispatches the expected PATCH/DELETE body.
  - Removed the SDK_CONTRACT_VERSION snapshot test from
    web/src/plugins/registry.test.ts (forbidden by AGENTS.md
    "Don't write change-detector tests"; the two remaining tests in that
    file already cover the new SDK surface behaviorally).

Closes #50547 (consumers of #50550).

Cross-vendor re-review: Gemini 3.5 Flash + GPT-OSS 120B (both SHOULD-FIX,
no remaining BLOCKERs after these fixes).
2026-08-15 00:33:32 -07:00
Teknium
d16326bb25 test: align lost-and-found schema pins with git_metadata_generation column
The salvaged #76716 adds git_metadata_generation to sessions (54 -> 55
columns). Update the synthetic-rebuild test's pinned widths and row
builders to the new current layout.
2026-08-15 00:33:11 -07:00
Teknium
cab6eb78f9 test(projects): widen lane-id derivation regression coverage
Cover the kanban ::kanban id, the -wt- suffix raw-path lane, and
Windows separator/trailing-slash spellings collapsing to one lane key.
2026-08-15 00:33:11 -07:00
Teknium
f378a8fb3b fix(projects): dedup project_create by primary_path (#75820)
Creating a project whose resolved primary path already belongs to a
non-archived project now raises a clear ValueError naming the existing
project (create_project) — duplicated projects each seeded an identical
copy of the repo subtree, multiplying the duplicate-lane bug per copy.
The agent-facing project_create tool is idempotent instead: it re-activates
the existing project rather than erroring. allow_duplicate_path=True keeps
deliberate duplicates possible. Also updates the legacy non-git lane-id
expectation to the branch-style id introduced for #53329.
2026-08-15 00:33:11 -07:00
embwl0x
e89532d97e fix(desktop): order async session git metadata 2026-08-15 00:33:11 -07:00
briandevans
f7de2ca416 fix(desktop): order project-tree lanes by recency in the overview, not alphabetically
_build_repos emptied each lane's sessions array for the overview (hydrate=False)
payload BEFORE _sort_lanes ran. _lane_sort_key derives a lane's activity from
max(_session_time(s) for s in group["sessions"]), so with the rows already gone
every non-trunk lane scored activity=0.0 and the sort key
(is_trunk, is_kanban, -activity, label) collapsed to alphabetical-by-label. The
documented intent — branches and linked worktrees sort by most-recent activity,
then label — was silently defeated on the projects.tree RPC that feeds the
desktop sidebar overview, while the drill-in path (hydrate=True) kept the rows
and sorted correctly. Any repo with two or more non-trunk lanes showed a
different order in the overview than when opened.

Move the session-clearing to after _sort_lanes/_disambiguate_labels so the sort
reads real recency. Lane counts are still captured before clearing, so
sessionCount and the slim overview payload are unchanged — only the order is
fixed, and the overview now matches the drill-in.
2026-08-15 00:33:11 -07:00
Tranquil-Flow
0d07fe63f9 fix(projects): use _branch_lane_id for non-git folders to prevent duplicate lanes (#53329)
_place_by_heuristic used the raw path as the lane key for non-git
project folders, while the desktop overlay independently computed
::branch::main for the same session (since git_branch was null).
The ID mismatch caused duplicate lanes — one from the backend with
the folder name, one from the overlay labeled 'main'.

Use _branch_lane_id(path, DEFAULT_BRANCH_LABEL) so the backend's
lane key matches the overlay's expected ::branch::main scheme,
eliminating the duplicate lane.
2026-08-15 00:33:11 -07:00
Teknium
94ce8396e8 fix(sessions): release active-session leases against their acquisition registry
A gateway active-session lease is acquired against the root HERMES_HOME,
but release_active_session()/transfer_active_session() re-resolved the
registry path from the *current* HERMES_HOME. Under native multiplex a
routed turn runs agent cleanup inside _profile_runtime_scope, so the
release looked under the named profile while the root entry stayed
alive — after max_concurrent_sessions routed turns every new session was
rejected with 'Hermes is at the active session limit' (#85431).

Pin state/lock paths on the lease at acquisition time and prefer them on
release and transfer. Fixes #85431.
2026-08-15 00:33:01 -07:00
rainbowgits
aba4934274 fix(agent): omit unsupported metadata on Relay scope.pop
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
2026-08-15 00:33:01 -07:00
Adolanium
b9672ea24e fix(pets): remove the non-PNG base draft after hardening
generate_base_drafts hardens every base draft to a transparent PNG with
_harden_transparency. When the provider returned a non-PNG file (webp,
jpg, or gif), the hardened PNG is saved under a new path and the original
draft is left in cache/images. Nothing prunes that directory outside the
gateway housekeeping loop, so a CLI or desktop draft round leaks one
original per non-PNG draft.

Remove the original after a successful hardening when the output path
differs from the input. PNG inputs (including mixed-case suffixes like
.PNG) are hardened in place so a case-insensitive filesystem cannot
treat with_suffix(".png") as a different file and unlink the output.
2026-08-15 00:33:01 -07:00
Adolanium
7de5a65906 fix(pets): delete row strips after extracting their frames
Each hatch generates one row strip per state into cache/images and
extracts the animation frames from it, but never removes the strip. The
only cleanup for that directory runs in the gateway housekeeping loop,
which a CLI, desktop, or cron hatch never starts, so the strips
accumulate for good.

Drop the strip after every attempt once its frames are decoded into
memory, including failed or retried attempts, so a hatch no longer grows
the image cache without bound.
2026-08-15 00:33:01 -07:00
andy
c99a45b28e fix(browser): reap leaked agent-browser daemons whose owner is still alive
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:

1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
   loop started, so a leak appearing after boot could never be recovered.

2. `owner_alive is True` skipped unconditionally. In-memory session tracking
   is lost on any exception path between spawn and registration, but the
   owner PID stays up — so such a daemon was skipped forever.

The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.

Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.

Changes:

- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
  the cleanup loop. Cycle 0 preserves the existing startup reap.

- When the owner is alive but the session is untracked, fall back to idle
  age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
  `max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.

- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
  socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
  there, making it a last-activity marker that survives hermes restarts and
  does not depend on in-memory bookkeeping surviving an exception path. It
  scans directory entries rather than reading the directory mtime alone:
  command names repeat, and rewriting an existing `_stdout_click` updates
  that file's mtime but not the directory's, so a dir-mtime-only check would
  report a busy session as idle and reap it.

Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.

Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:33:01 -07:00
Reksely
3d73821e9d fix(agent): stop thread output descriptor leaks 2026-08-15 00:33:01 -07:00
Teknium
edd73daaf4 test(kanban): regression for idle-board WS disconnect detection (#77833) 2026-08-15 00:32:53 -07:00
Teknium
67a1c1ed1a fix(dashboard): use a fixed sidebar cache TTL (no HERMES_* env var for non-secret config) 2026-08-15 00:32:53 -07:00
Christopher
5bceb3e84b fix(dashboard): add idle back-off to PTY pump loop (#42627) 2026-08-15 00:32:53 -07:00
Lucas Oliveira
f0cfe5a56f perf(dashboard): bound multi-profile sidebar polling 2026-08-15 00:32:53 -07:00
Teknium
4f29374662 fix(gateway): send-once spritesheet semantics for pet.info (#54730)
pet.info accepts knownRevision; when it matches the active sheet's
revision the multi-MB spritesheetBase64 is elided and
spritesheetUnchanged=true is returned. The desktop floating pet passes
the revision it already holds and keeps its cached bytes, so backstop
refreshes no longer resend ~3.2MB frames over the WS (write-loop stalls,
disconnect storms). Legacy callers omitting knownRevision get the full
payload unchanged.
2026-08-15 00:32:44 -07:00
thatssoheil
8052d5dd24 test(pets): lock the quoted-false behavior on the CLI surfaces
Review follow-up (final round PASS with a repeated suggestion): the pets
CLI regression test only exercised real bools, leaving the exact bug this
fix shipped untested. Add a quoted-'false' test driving _has_active_pet
(now False) and toggle_pet_display (now takes the ENABLE branch,
distinguished by the 'no pets installed' error instead of err=None from
the old wrong-way disable). Verified RED on the pre-fix pets.py and GREEN
on the fix.
2026-08-15 00:32:44 -07:00
thatssoheil
9e8828999d fix(petdex): quoted 'false' now disables display.pet.enabled everywhere
Three bare bool() reads of display.pet.enabled (deep-merged config, so a
hand-edited quoted YAML value lands as the string 'false'): the pet.cells
gate, the pet.gallery enabled echo, and the shared pet-state helper.
bool('false') is True, so a quoted value kept the mascot enabled against
the operator's explicit intent.

All three now go through utils.is_truthy_value (default False, matching
DEFAULT_CONFIG). Regression test drives pet.gallery with a quoted 'false'
config and asserts enabled=False; verified RED on the old code.
2026-08-15 00:32:44 -07:00
Teknium
fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00