_dedupe_display_generations chose the right representative row per logical
message but sorted the survivors by that representative's id. A protected-tail
copy written into a newer compaction generation has a higher id than messages
emitted after the original, so include_compacted reads came back as C, A, B.
Anchor the sort on the logical message's first-ever row id instead.
Salvaged from #93869 (the tui_gateway half of that PR is superseded by #100504
and #104137); the code moved from hermes_state.py to hermes_state_messages.py
since, so the change is re-applied to its new home with the PR's regression
test verbatim.
Whether dangerous commands run unasked is state worth seeing at a glance,
so the approval pill (yolo lightning) leaves STATUSBAR_HIDDEN_BY_DEFAULT.
Existing stores were seeded with it hidden, so the hidden-set key moves to
`hermes.desktop.statusbarHidden.v2`, seeded from v1 minus `approval-mode`:
other customizations survive, the zap appears once on update, and hiding
it again persists under the new key.
The whole-bar preference lived at `hermes.desktop.statusbarVisible`. For a
stretch (d399c164 → 120e465c) the atom's fallback was `false`, so any
install that launched in that window persisted a hidden bar the user never
chose, and flipping the fallback back to `true` only helped fresh stores.
Move the preference to `hermes.desktop.statusbarVisible.v2` and do not seed
it from v1: every existing install comes back to "on" once on update, and a
hide made afterwards persists under the new key. Fresh installs are on by
default as before.
spawnPoolBackend() is not on every dial path: a primary route (startHermes),
a registry remote scope (connectRegistryBackend), a reused primary SSH
backend, or a guard rejection all settle the claim without requesting a
slot, so a foreground mark set for that dial stayed in pendingForegroundSpawns
and would have upgraded the next background hydration spawn of the same key.
applySpawnPriority() now returns the cleanup; both IPC handlers run it in a
finally around the claim. The mark is also taken right before the slot
request instead of at function entry, so the remote branch never consumes it.
Every user open first probes its route (sharedPrimaryRoute /
isAttachedSharedRemote) with getConnection / getConnectionFor, and only then
dials the secondary. With #102496 only the second dial carried
priority: 'foreground', so main started (or joined) the spawn as a background
slot wait on the probe and the click still waited out the probe's 20 s
RECONNECT_ATTEMPT_TIMEOUT_MS before promotion kicked in. Thread the priority
into both probes; the activation doors (ensureGatewayForProfile /
ensureGatewayForAgent) pass 'foreground' explicitly.
Also drop the renderer-side isBackgroundSlotWaitTimeout + the try/catch whose
two branches both rethrew: Electron rebuilds IPC rejections as a plain Error,
so name/silent/priority never reached the renderer and the helper was dead.
Test: gateway-spawn-priority.test.ts asserts every dial of a foreground open
carries the tag and an untagged open never does (red on the #102496 head).
Follow-ups to the #102496 salvage in the Electron main process:
- pendingForegroundSpawns leaked: promoteInFlightLocalSpawn marked the key
even when the pool entry already existed (the common click path), and
nothing consumed it. After that backend was reaped, the next dial for the
key - normally 10 s roster hydration - spawned as foreground and sat in the
reserved slot. Mark only when no entry exists yet; spawnPoolBackend consumes
the mark before any early return (remote route included) so it never
outlives the dial.
- The "waiting for a free local slot" log fired for a foreground request that
was granted the reserved slot immediately (condition was queuedCount > 0
after request()). The request now reports `queued`; log only then.
- One promotePoolEntry() and one logPoolSpawnFailure() replace three copies of
the promote snippet and two copies of the background/foreground log branch;
spawnPoolBackend reads entry.spawnPriority instead of a second opts channel;
isBackgroundSlotWaitTimeout is an instanceof check (same process as the
class, no duck typing).
LocalBackendSpawnCoordinator is FIFO with maxBackends=3, so launch
hydration queues ~27 ensureBackend calls and a user click times out
waiting for a slot. Reserve a foreground slot, drain foreground first,
and fail background slot-wait quietly.
Fixes#102281.
It is lifecycle logic (detached sentinel, reap cancel) and its siblings
_reattach_refusal / _cancel_ws_orphan_reap are already there; server.py is the
facade and gets no new definitions. Also trims the _claim_or_reuse_live comment
to point at where the reap cancel now happens.
_resume_live_unpersisted hand-rolled the same transport + viewers + reap-cancel
sequence _rebind_live_transport now owns. Route it through the helper; the
stdio case (no current transport) keeps cancelling the reap as before.
The salvaged fix correctly put the activate guard + transport rebind under the
process-wide resume lock, but it dragged _live_session_payload in with them. The
Desktop passes omit_messages=true (cheap), the Ink TUI does not: every TUI session
switch then read the full persisted history from the profile DB while holding the
lock that serializes every resume, disconnect and reap Timer.
Extract _rebind_live_transport from _live_session_payload; activate does guard +
rebind under the lock (the part that must be atomic with grace expiry) and builds
the payload after releasing it.
The 4007/4009 guard the salvaged fix added to prompt.submit, session.activate,
_resume_live_unpersisted and _resume_reuse_live_locked was the same six lines
four times. One helper in session_lifecycle.py (next to _cancel_ws_orphan_reap,
whose contract it mirrors) so the next reattach path cannot drift from the others.
Behaviour unchanged; the 19 race tests cover all four call sites.
Re-scope #98106 onto the activity-based orphan policy from #100504. Keep timer ownership across callbacks and continuations, serialize reconnect paths against interrupt claims, and avoid recursive eager-resume locking. Leave cleanup polling armed when concurrent cold reuse is rejected.
- probe_only is consumed inside __init__ only; no instance attribute.
- _sync_manager is assigned only on the probe branch, after the connection is
up, so a normal SSHEnvironment whose constructor fails still behaves exactly
as before (its __del__ cleanup does not reach the shared socket teardown).
- _remote_home has no reader on the probe path; not assigned.
- prompt_builder passes probe_only=True unconditionally: which backends honor
it is the builder table's decision, not a caller-side env_type check.
- Tests trimmed to the probe_only contract: the cleanup-raises case was green
on main (pre-existing try/except), the socket-name length and double-cleanup
assertions covered pre-existing behaviour.
The prompt-time backend probe built a normal SSHEnvironment just to run a
one-line `uname`. That constructor detects the remote home, creates the
~/.hermes tree, force-uploads every sync file and snapshots a login session;
when the throwaway object was later garbage-collected, __del__ -> cleanup()
ran sync_back() and `ssh -O exit` against the ControlMaster socket the
agent's real environment shares (keyed by user@host:port).
Add an internal probe_only construction path for SSH: an isolated,
same-length ControlMaster socket (keyed by the instance's session id), no
remote dir setup, no FileSyncManager, no session snapshot. The probe now
tears its own connection down explicitly on success, non-zero exit and
exception, without replacing the probe result when cleanup fails. Normal
SSH callers and non-SSH backends are unchanged.
Salvaged from #77933 onto the facade/sibling layout (the probe body moved to
_run_backend_probe, _create_environment to tools/terminal_tool_backends.py).
get_all_toolsets() copied static TOOLSETS entries verbatim, so /toolsets and the
dashboard showed only the built-in tools for a colliding name even though
get_toolset() now unions the MCP alias in. Route those names through
get_toolset() and document the alias/collision behaviour.
If an MCP server registers itself under the same name as a built-in
toolset (e.g. an MCP "homeassistant" running alongside the built-in
`homeassistant` toolset), `get_toolset()` returned the static definition
early and the MCP tools were silently shadowed — they got registered
into `mcp-homeassistant` but never surfaced when the agent looked up
`homeassistant`.
Detect the collision via the registry alias (`alias_target` starting
with `mcp-`) and merge the two tool lists. The description is annotated
so the source is visible in `/toolsets info`.
The default backend needs no Reddit account, login, cookie or API key; the
optional upgrade is a free 'script' app registration using the app-only
client_credentials grant, never a user login. Said in the skill's
Prerequisites (with a comparison table), Procedure, Pitfalls, the runtime
doctor/thread notes, and a delimited .env.example block.
Two zero-install research skills plus routing guidance, ported as ideas
(not code) from the Agent Reach skill's per-platform backend routing.
reddit-reading (skills/social-media): subreddit listings, site/subreddit
search, threads with comments, user pages. Live-verified from a datacentre
IP: www .json, api.reddit.com, old.reddit, r.jina.ai and the browser tool
all return 403 / an empty shell / a humanity check; the Atom .rss endpoints
are the only anonymous path and are throttled to ~1 request/min/IP. The
script waits out the x-ratelimit-reset window once and retries, and
switches to the OAuth API (scores, nested comments, ~100 req/min) when
REDDIT_CLIENT_ID/SECRET are present. `doctor` reports the active backend.
rss-feeds (skills/research): RSS 2.0 / RSS 1.0 / Atom / JSON Feed parsing
with UTC-normalised dates, --since/--limit, and feed discovery behind a
page URL (<link rel=alternate>, then well-known paths). Stdlib only; the
optional blogwatcher skill remains the stateful many-feed reader and now
points at this one for one-off reads.
grounded-citations gains a "Multi-Platform Sweeps" section routing
"what are people saying about X" tasks across web, Reddit, feeds, video,
code and X with per-platform attribution and coverage-gap reporting;
competitor-news-monitor references the new sources.
Tests: two invariant tests per skill (format normalisation + discovery;
anonymous 429 handling + OAuth routing/flattening), no network. The
reddit test caught a real bug: the feed footer's "/u/author" leaked into
post bodies.
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.
redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.
Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
DockerEnvironment.cleanup() runs docker stop + docker rm -f on a daemon thread and
keeps the handle on the env. The idle reaper pops the env out of _active_environments
BEFORE calling cleanup(), and the atexit drain only iterated that registry — so a
detached env's worker was unreachable and died with the interpreter, leaving a
stopped (or, if the exit came fast enough, still running) labeled container behind
while the log said the environment was cleaned.
Every teardown worker is now also recorded in a module-level set in docker.py;
_atexit_cleanup joins that set after the registry pass via
DockerEnvironment.wait_for_all_teardowns (re-snapshotting each pass, since the
reaper can start a worker while the drain runs). Finished workers are dropped from
the set on each drain so it cannot grow across a long gateway life.
Mechanism from #86344 by @PRATHAMESH75; this is the slim redo — one shared set
plus one static drain hooked into the existing terminal_tool atexit, instead of a
second atexit registration in docker.py.
On hosts where Docker ships as a snap (Ubuntu cloud images / Azure VMs), the
snap's AppArmor confinement turns two sandbox hardening flags into a dead
container at start: `--init` fails with "exec /sbin/docker-init: operation not
permitted" and `--security-opt no-new-privileges` then fails every exec the
same way ("exec /usr/bin/sleep: operation not permitted"). This is snapd
LP#1908448 — not probeable from the client, and docker_extra_args cannot remove
flags we add.
`terminal.docker_snap_compat: true` drops exactly those two flags; cap-drop ALL,
the tmpfs hardening, PID limits and the privdrop caps are unchanged, and a
warning is logged at container start. Bridged everywhere the other docker_*
keys are (CLI env map, gateway env map, `hermes config set` sync, terminal_tool
env read, the shared container_config shaper, DEFAULT_CONFIG).
Two stacked Windows-only failures in FileSyncManager.sync_back:
1. The staging tar was a NamedTemporaryFile held open while the backend's
bulk_download_fn reopened the same path for writing — PermissionError on
Windows (exclusive handle). mkstemp + close + unlink in finally instead.
2. Staged member keys came from os.path.relpath and mapping parents from
Path(remote).parent — both stringify with backslashes on Windows, so no
staged file ever matched its POSIX remote key and every edit was skipped
as "no host mapping". Keys are now built with as_posix(); parents with
posixpath.dirname; the inferred host path is joined with the host's Path.
(cherry picked from commit e8cab109db, re-applied on the current file)
serve --isolated is detached on purpose (setsid/nohup, PPID 1) so it survives the SSH
channel closing, and every teardown path lived on the client. A laptop that sleeps
mid-session (dark wake reconnects the tunnel, spawns a backend, sleeps again) therefore
left a new backend behind every cycle, each one an extra writer on state.db — the
multi-writer source behind the WAL corruption incidents.
Only for backends started with --ssh-session-token-file: an ASGI wrapper counts accepted
WebSocket sessions on every dashboard route without touching the handlers; a watchdog
requests a graceful uvicorn exit (WAL checkpoint, exit 0) once no client has been
connected for dashboard.ssh_isolated_idle_grace_s (default 15 min) AND no agent turn is
running. An unreadable turn state fails closed (the backend stays up). Loopback normally
disables the WS ping, but across a tunnel the local socket stays healthy while the far
end sleeps, so these backends keep a slow ping (60s / 10 min) that a GIL-holding turn
cannot trip.
Design, client-count/turn-probe/fail-closed shape and the tunnel-ping rationale from
#101678 by @StanleyStetson; this is the slim redo on the decomposed web server (no
exclusive home lock / handover protocol — the newcomer never needs to evict an idle
incumbent once idle incumbents exit on their own).
_probe_regular_file returned 'missing' for ANY non-zero exit, so a docker container that
was still starting, paused, or otherwise unable to exec made read_file / read_file_raw /
read_file_bytes report a false 'File not found' — which the model then trusted for the
rest of the session. The probe now echoes a sentinel for a genuinely missing path and a
non-zero exit without either sentinel is reported as an environment failure with a retry
hint; search_files does the same when its existence probe returns neither marker.
Salvaged from #44753 by @dredozubov.
'Skipping unsafe MEDIA directive path' fired for every non-existent file too, sending
operators down a security-policy rabbit hole for what was a hallucinated or untranslated
path. The warning now names the reason: 'not found on this host' vs 'denied by the
delivery policy'.
docker info needs the /info endpoint, which socket proxies such as tecnativa block by
default, so doctor reported 'docker daemon not running' against a fully working
DOCKER_HOST. docker version hits /version and is what the backend itself probes with.
_find_git_root walked cwd's parents with Path.exists(); on hosts where a parent is
mode 700 for another user that raised PermissionError out of prompt construction.
Treat an unreadable ancestor as 'no .git here'.
_HOST_CWD_PREFIXES only listed C:\ and C:/, so a D:\ or E:/ working directory on a
Windows host reached docker run -w and the container started in a directory that does
not exist. The two prefix scans now go through one _is_host_cwd predicate that matches
any drive letter with either slash.
The bare deepseek/deepseek-v4-flash slug is the pre-snapshot release; both
aggregators now carry deepseek-v4-flash-0731 (and Nous exposes the rolling
~deepseek/deepseek-v4-flash-latest alias). Keeping both rows in the curated
picker just duplicates the flash tier.
- OPENROUTER_MODELS (derives the nous list): drop deepseek/deepseek-v4-flash
- model-catalog.json regenerated
Deliberately KEPT: alibaba-token-plan / opencode-go / commandcode / deepseek
direct plugin fallback_models + default_aux_model (bare id is the wire slug
those providers serve), DEFAULT_CONTEXT_LENGTHS / reasoning floor / pricing
snapshot entries (manually-typed id still behaves), and the
deepseek-chat/deepseek-reasoner -> deepseek-v4-flash alias normalization.
The hermes_tools stub module a remote kernel imports is generated from sandbox_tools once,
at spawn, but the registry key was (owner, env_type, task_env_id) only: a later
execute_code call with a different tool set (skill loaded, toolset toggled) reused the
kernel and got stale stubs. The tool set is now part of the key; a different set gets
its own kernel and the over-cap eviction keeps the newest.
Salvaged from #97265 by @Liuzikaii.