`hermes logs --since` and `--level` passed every line that had no leading
timestamp. A traceback's frames are written without one, so an old error
printed its frames without the header that was filtered out, and
`--component` dropped the frames of a matching record. `hermes logs` now
reads each unstamped line as part of the record above it: the line gets
that record's verdict for every filter, in the tail read and in `-f`.
Lines before the first stamp in the read window have an unknown time and
level, so they are dropped when `--since` or `--level` is set.
mcp-stderr.log had no parseable stamp at all: the banner started with
`=====` and server output was copied raw, so `hermes logs mcp --since`
printed the whole file. The stderr tee already reads each server's
stderr in a thread, so it now writes one line at a time, each prefixed
with the asctime-shaped local stamp the Python logs use. The banner
starts with the same stamp. The stamp comes from new public
`timestamp()`/`stamp_line()` in hermes_cli/stderr_timestamp.py, which
stays stdlib-only. The desktop MCP log view accepts both banner shapes.
A test now requires a real writer's sample line for every LOG_FILES
entry to parse with `_parse_line_timestamp`. The docs no longer say the
`timezone` key changes log timestamps; log lines use the machine's
local time.
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.
Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
rename_profile removed the launchd/systemd unit only when the gateway
was running, and never touched the s6 slot. A unit installed under the
old name but stopped stayed behind: it runs --profile <old> with
HERMES_HOME pinned to the moved directory, so the next login or
container boot crash-looped it for a profile that no longer exists,
and deleting the renamed profile never found it.
The old unit is now removed whether or not the gateway runs, the s6
slot moves to the new name, and the command prints how to reinstall
the service under the new name.
(cherry picked from commit 76002ea29ab8f8871f2f962ef4535408f29c590f)
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(desktop): bound NVIDIA EGL fallback to confirmed-broken driver series
The 580+ check was open-ended: every future NVIDIA series (615.x, #123203)
was forced onto CPU SwiftShader rendering even though only 580.x has
confirmed EGL probe reports (#40077: 580.159.03, 580.173.02) and 570.x is
the recommended downgrade. Replace the >= floor with a closed set of
known-broken series majors so healthy newer drivers keep GPU rendering.
HERMES_DESKTOP_NVIDIA_SWIFTSHADER still overrides both ways.
* docs(desktop): document the NVIDIA SwiftShader override and scope its force-on claim
- add HERMES_DESKTOP_NVIDIA_SWIFTSHADER to environment-variables.md (both
directions, with the recovery-hatch framing the closed-set design leans on)
- correct the module header: force-on does not apply where an earlier gate
already returned (remote display, WSLg, DISABLE_GPU=0)
Review follow-ups on PR #123213.
* test(desktop): pin NVIDIA EGL detection to set membership across a driver sweep
The closed-set test only checked that 580 is a member, so it still passed
if a neighbouring series was added to the set or the check drifted back
to a floor. Sweep majors 500-700 and require enable === set.has(major);
this fails against a '>= 580' floor. Also clears the prettier arrow-paren
diff that js-autofix would have reverted on merge.
---------
Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.
- cua-driver is a default PM package, so the installers, a bare
`hermes pm install` and `hermes update` carry it on every target it builds
for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
and `hermes update` both run, provisions the Browser Use CLI for the default
and explicit Browser Use backends. `--skip-browser` declines it along with
agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
`--without cua-driver`).
- The update message stops calling every default "browser tools".
CI run 36189416163 failed after "git: unpacking": executing the cached
fetch-<sha> PortableGit PE in place left it handle-held (Defender
on-execute scan / the stub's RunProgram child chain) past pm's ~2 s
_remove_entry retry, so download cleanup raised WinError 32. Review
(teknium1, 5 threads) adds the rest:
- Git.unpack now copies the artifact into a .sfx-* dir beside the
staging tree and executes the copy; pm's .staging-* teardown
(ignore_errors) owns that path, so any hold lands on a disposable
path and the cache dir only ever holds read handles
- refuse off-Windows with the cross-host trade-off stated instead of a
raw PermissionError; documented in package-management.md and the PR
- the extractor is a GUI-subsystem stub that is silent under -y: error
messages now carry the exit code and the usual causes (disk full,
path length, antivirus) instead of promising captured output; the
docstring no longer claims "no GUI" and records that the stub shows
an Extracting window and runs the vendor post-install
- install.ps1: WaitForExit(600000) + Kill() mirrors pm's timeout=600,
Fail reports the exit code and the silence
- tests: the OS-refused-exec assertion and the not-in-text change
detectors are replaced by subprocess-argv invariants (scratch copy
location, exit-code message, off-Windows guard before any execution)
Related to #122512
(cherry picked from commit adaf76a286bebe30c89aee4174df27f1945b60c9)
A `terminal(background=true, heartbeat=N)` tick queued a notification every N seconds
whether or not the process had printed anything, and every queued event costs the owning
session a full model turn. On Desktop and the TUI that turn painted the wake as a user
bubble ("[Background process ... heartbeat #9 ... (no new output since the last
heartbeat)]") followed by the model's "Still running normally." — over and over, for a
process whose row on the status stack already said it was running — and while the wake
held the session's turn, the user's own prompt sat queued behind it.
- `ProcessRegistry._emit_heartbeat` skips a tick with no new output. The sequence counts
delivered beats only; the "(no new output)" placeholder in the formatter is gone.
- TUI/Desktop type heartbeat rows `display_kind: hidden` (the kind both clients and the
transcript preview already honour); the CLI paints a one-line receipt and persists the
row hidden, so reopening the session in Desktop shows only the agent's reply.
- Desktop hydration drops heartbeat rows persisted by older backends the same way.
- `display.background_process_notifications: off` is honored by the TUI/Desktop poller and
the CLI drain, not just the messaging gateway. `off` mutes process-driven wakes only:
a finished `delegate_task(background=true)` still lands.
Supersedes #123123 (cherry-picked; scoped so `off` keeps subagent results) and #119202
(cherry-picked; `heartbeat: 0` is schema-valid so models that materialize every field
stop tripping the foreground guard).
latent-spaces/brag (MIT) turns the project you just built into a short
launch video with music, motion and share copy. It ships two skills:
/brag, the Hyperframes workflow with a bundled music and SFX library,
and /brag-slim, a single SKILL.md where the model builds the whole video
with local tools. /brag hands off to its bundled copy of brag-slim on
Claude Opus 5.5.
Both follow the impeccable/archify pattern: catalog stubs whose
metadata.hermes.upstream pointer makes
`hermes skills install official/creative/<name>` pull the live tree
through OptionalSkillSource._fetch_from_upstream. Nothing is vendored.
The brag stub documents that its Hyperframes path loads HeyGen's
hyperframes-* domain skills by name, and that the skills guard blocks
four of the five today (hyperframes-creative scores dangerous).
brag-slim has no such dependency.
Docs: two generated pages, two catalog rows, two sidebar lines.
Credit: Shunit Haviv Hakimi (shunithaviv), upstream author.
HERMES_DESKTOP_IGNORE_EXISTING=1 only wrapped the PATH probe, so a usable
active install or system Python still started a local serve. Skipping those
rungs on every resolve would make the post-bootstrap re-resolve return
bootstrap-needed and start the installer again. The first pass now falls
through to connect/onboarding; the re-resolve keeps the runtime just installed.
Fixes#117682
The previous commit stops auto-routing a non-sk-or- OPENAI_API_KEY to
OpenRouter. The env-var reference still described OPENAI_API_KEY only as a
custom-endpoint key and the non-interactive setup hint listed it as an
OpenRouter alternative. Say what the key now selects so users with an
OpenRouter key in OPENAI_API_KEY know to move it to OPENROUTER_API_KEY.
Co-authored-by: notwitcheer <notwitcheer@users.noreply.github.com>
They are not "never auto-repaired": resuming such a session re-pins the
full tool surface (restore_agent_tool_prefix), after which a scan sees
skill_manage without the Skill Safety guidance and clears the prompt.
The repair-prompts detector cleared two legitimate prompts:
- a pin with skills_list/skill_view but no skill_manage and zero skills
installed: build_skills_system_prompt returns '' and SKILLS_GUIDANCE is
only emitted with skill_manage, so the healthy prompt has neither marker;
- an exact memory-only pin, which is also a user toolsets=[memory] config;
the healthy rebuild was re-flagged on every run (not idempotent).
Key the decision on the missing '## Skill Safety' guidance, which is
unconditional when skill_manage is in the pin, and require skill_manage in
the pin. Memory-only rows are now reported as unverifiable; the explicit
SESSION_ID override still clears them and their pin.
Docs: describe the tightened rule and note that a running gateway keeps
cached prompts in memory, so it must be restarted after --apply.
Merging an owned category per skill means a skill the author removes from
the distribution is no longer deleted on update, matching top-level
skills/. Say so in the English and zh-Hans reference so authors are not
surprised that retired skills linger.
A non-elevated Windows ARM64 install threw "run setup-hermes.ps1 in an
Administrator PowerShell" whenever Visual Studio ARM64 C++/Clang were
missing. Interactive runs now launch the signed VS installer through a
UAC prompt; CI, ssh and scheduled runs keep the explicit instruction.
Browser tools find agent-browser only in PM's store or on PATH and their
readiness check never installs it, so a fresh install silently had no
browser_* tools. Package.default marks an optional package that a bare
`pm install` also carries; a failed download of it warns instead of
failing the install. `pm install --without NAME` records the opt-out in
declined-packages.json beside PM's install state, and naming the package
explicitly clears it. agent-browser and chromium gain a Termux gap: Termux
owns its browser stack and there is no bionic Chromium.
Contributors had to run `python -m pm.build_env --source . --lock-only`, then
re-source activate, then call sync_venv for opt-in extras, while `hermes pm
lock` only pinned tool artifacts in pm/lock.json. Now `hermes pm lock` with no
arguments is the one step: it runs the same PM check_project_lock /
lock_project operations as build_env's --check-lock / --lock-only (same
exclude-newer quarantine), writes nothing when the lock is current, changes
no environment, and prints the activate command to run next. Activation
syncs [all] plus recorded extras only, and --test-extras replaces the
default, so a newly added extra outside [all] gets
`source ./activate --test-extras all,NAME` (`-TestExtras` on PowerShell).
`hermes pm lock --bump NAME VERSION` keeps writing pm/lock.json; NAME and
VERSION are now only accepted together. build_env --lock-only is unchanged
for scripts. Contributor docs, AGENTS.md, CONTRIBUTING.es.md, the pyproject
comment, and the uv.lock CI remediation text point at `hermes pm lock`.
The image refused every lazy install (HERMES_DISABLE_LAZY_INSTALLS=1), so
edge-tts and the other opt-in SDKs could never be installed at runtime. PM
never writes the sealed /opt/hermes/.venv: it builds a generation under
$HERMES_HOME/installs and commits it in facts.json there, which already
survives container recreates and image updates. Drop the refusal.
Surviving updates means a new image boots under a selection resolved against
the previous image's lock. refresh_dependencies() re-resolves the recorded
extras and plugins against the current inputs; if that fails (offline), it
deselects the generation so the image's own environment boots, keeping the
extras recorded for the next boot or install. stage2 runs it as hermes before
any service starts, then collects generations nothing selects any more, since
nothing else collects them automatically and each one is a full venv.
HERMES_DATA_DIR_SUFFIX, HERMES_REPO_URL, HERMES_UPDATE_STATUS_FILE and
HERMES_UPDATE_UI_ACTIVE are set by Hermes' own bundles, installers and
update shim, not by users, but only the suffix was documented and it
read as a user knob. Add an "Internal bridge variables" table that
says who sets each one and why it cannot live in config, and move the
suffix row there.
`hermes import` only checked the central directory (`is_zipfile`,
`namelist`), so an archive with one member whose deflate stream or CRC is
rotten passed validation and blew up mid-restore with a zlib.error
traceback -- after config.yaml and everything before the bad member had
already been replaced, with later members never written (#121258).
Add a pre-flight pass in run_import that streams every member through
1 MiB reads (zipfile verifies the CRC at EOF) and collects every
BadZipFile / zlib.error / EOFError. If any member is damaged the command
prints a capped list and returns 1 with the home untouched. Stdlib
`ZipFile.testzip()` is deliberately not used: it lets zlib.error escape
and names at most the first bad member.
Widen the per-member catch from the previous commit with EOFError so a
member that rots between the two passes still becomes a "skipped" warning
+ `Import incomplete` / exit 1 instead of a traceback. Rework that
commit's test to drive the per-member path (the archive now never reaches
it with a corrupt member), and let `_break_member` serve the pre-flight
read before failing the restore's own read.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
`_import_members` catches the PermissionError/OSError of each member it
cannot publish and records it in `errors`. That includes the state.db the
live-safe restore refuses because a gateway or dashboard still holds it
(#100960, #110179). `run_import` printed those under "Warnings (N files
skipped)", then printed "Done. Your Hermes configuration has been
restored." and returned None. `cmd_import` did not forward a return value,
so the shell status was 0. A script chained on `hermes import --force`
carried on over a partial restore, and the dashboard's import action
showed the green "done" badge. For the refused state.db case, the user's
sessions were never restored.
`run_import` now returns 1 when `errors` is non-empty. It words both the
summary header and the final line as "Import incomplete", and `cmd_import`
forwards the return code, the same contract `hermes backup` has for an
incomplete archive. Runtime files the import deliberately keeps
(gateway.pid, SQLite sidecars) and the older-backup session warning stay
warnings. The gateway revive still runs, so the files that did land come
up as before.
Measured on main with a fresh HERMES_HOME through `main()`: a 3-file
backup with one read-only target directory, and a refused state.db (live
DB left at 3 sessions / 12 messages). Both used to exit 0 after "Done ...
restored". Both now exit 1 after "Import incomplete: 1 file(s) were not
restored". A clean import still exits 0 and prints "Done".
(cherry picked from commit a0f808ec9f868128f4174e4816e095ac17f18d28)
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Follow-up to the salvaged mirror commit.
- Opt-in (`mirror_to_session: true`, or `hermes webhook subscribe --mirror-to-session`),
default off like cron's `mirror_delivery`. The mirrored text lands with user authority in
the target chat, and on `deliver_only` routes it is the raw rendered payload, so a route
author has to ask for it. Only a real boolean `true` opts in (a YAML string "false" no
longer does).
- The mirror runs inside the routed profile's scope, so a `/p/<profile>/` delivery writes
into THAT profile's state.db. A Telegram DM chat_id is the user's id on every bot; an
unscoped mirror landed in the default profile's DM with the same person (live-probed).
- Tests trimmed to two invariants against a real state.db: an opted-in delivery lands in the
routed profile's chat session and not the default profile's; a route without the opt-in
(including `mirror_to_session: "false"`) never touches the target transcript.
- Docs: webhooks guide section "Replying to a delivery" (default, profile scope, trust note),
CLI reference row, bundled hermes-agent skill reference.
cryptography ships no win_arm64 wheel, so every Windows ARM64 venv sync
compiles it from the sdist and needs MSVC, Clang, Rust and static OpenSSL.
Only setup-hermes.ps1 (and so activate.ps1) prepared that environment,
between a `pm install --tools-only` and the real sync. install.ps1,
`hermes update` and repair ran the same sync without it and failed in
openssl-sys.
PM owns the sync, so PM prepares it. pm/native_build.py holds the adapter
(moved from scripts/build/windows_deps.py) plus source_build_environment(),
which prepares only on win32-arm64 when the synced project carries the
provider script. A payload has prebuilt dependencies and needs no compiler.
VenvPackage.apply and build_environment pass the result to uv children
only. It carries the bridged pip index settings, which managed_environment
applies only to the ambient environment. The state root stays the store
parent, so existing vcpkg/OpenSSL builds are reused.
setup-hermes.ps1 collapses to one `pm install`: the tools-only split existed
only for this preparation, and pm install already puts its tools on PATH
before the venv sync (pm/cli.py activate check).
Not yet verified live on Windows ARM64.
Follow-up to the salvaged wrapper: the argument layer now lives in a topical
sibling (agent/legacy_cli.py) instead of growing the run_agent facade, and
`python run_agent.py` (the installer's PATH launcher for hermes-agent) routes
through the same parser instead of fire, so `--version` and a bare invocation
no longer run the demo turn there either.
- --help/-h/--version use argparse's own exit path; options carry help text
- the runner is injected (`run=`) so `python run_agent.py` does not import
run_agent a second time
- tests trimmed to two invariants that resolve the console-script target
from pyproject exactly like pip's wrapper (red on main, green here)
- docs: hermes-agent section in the CLI reference
OpenRouter relays OpenAI's "this user has been blocked for a previous
policy violation" as HTTP 200 + an SSE error event. The SDK raises a
status-less APIError that no classifier rule matched, so it landed in
the retryable `unknown` bucket: every retry was re-sent (3 streamed
requests by default), a configured fallback only engaged after the
ladder, and the user was told the provider "looks temporarily
unavailable. Wait a minute and send /retry".
- error_classifier: match the ban phrase status-agnostically in
_provider_special_cases -> provider_policy_blocked (non-retryable,
fallback, no credential rotation: every key on a banned account is
banned; a 403 variant is not a bad key).
- turn_failure_copy: provider_policy_blocked chat copy now covers an
account block as well as data/privacy settings; add its cause gloss
so cron and subagent notices explain it instead of printing raw text.
- cron: provider_policy_blocked action (retrying won't help; pin another
model).
- docs: FAQ entry for the reply; auto-recovery ladder exclusion list.
source ./activate and . .\activate.ps1 always run setup, and setup always
runs `python -m pm.cli install`. An explicit install re-hashes every
published tool entry so it can repair a corrupted one. On this machine
that hash reads about 550 MiB and takes 6.0 of the 6.6 seconds, even
when nothing changed.
Shell activation now passes --trust-recorded and trusts the digest the
install recorded, the same check startup already uses. A missing tool
is still installed and a stale venv is still rebuilt. `hermes pm
install` and `hermes update` keep the byte check, and the flag refuses
names, --extra, and --target so it cannot narrow an install someone
asked for by name.
Measured on this machine, already up to date: the activation path drops
from 6.6 s to 0.9 s. A bare `python -m pm.cli install` stays at 7.2 s.
The remaining 0.9 s is process startup and imports, not hashing.
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:
- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
the credential pool, not the environment. Chat and image_gen/openrouter find
it through resolve_runtime_provider. video_gen reported OpenRouter
unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
routed profile's video jobs were submitted, polled and downloaded with the
launch profile's key and billed to that account. A profile whose key lived
only in its own .env could not use the backend at all.
The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.
OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
--clone carried memory.provider (e.g. hindsight) in config.yaml but not the
provider's own config under the profile home (hindsight/config.json,
mem0.json, ...), so the clone booted with the provider selected and silently
unavailable. Copy the ACTIVE provider's <provider>/ dir and/or <provider>.json
by the same convention the dashboard memory-provider routers read, guard the
name against path traversal, tighten copies to 0600 (they can hold an API
key), and say so in the CLI notice. Convention-based on purpose: hindsight is
a catalog plugin now, so a hook the plugin must implement could not fix the
reported case, and no provider module is imported during profile create.
Supersedes #43107 (credit @bionicbutterfly13 for the direction).
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.
Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.
The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
tools-reference.md listed manage_catalog under the connections section, which the
reference-docs contract resolves against the connections toolset. The agent-runtime
post-hook ownership test enumerates every tool in AGENT_RUNTIME_POST_HOOK_TOOL_NAMES;
manage_catalog now runs through both executor paths there.
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog:
- search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with
whether each is already installed in the default profile. Read-only.
- install: opens the same connection operation manage_connections opens, with rows of kind
plugin / skill. Nothing installs until the user approves a row. An approved row installs
into `default` (or the profile the Advanced modal named) through dashboard_install_plugin /
the hub's headless install, so the catalog pin, kill list, security scan and live
activation (#119644) are the host's. The row settles with the live MCP tool names and the
plugin's skill.
The model sends catalog ids and an action only; every other key is refused before anything
runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own
text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the
result is the `hermes plugins install` / `hermes skills install` pointer.
- contract: plugin/skill targets follow the MCP transitions.
- run.apply_answer / reissue route a card answer to the module that owns the operation.
- tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never
deferred behind tool_search.
- docs: tools reference, toolsets reference, plugin catalog page.
Linear NS-964.