CI run 36189416163 failed after "git: unpacking": executing the cached
fetch-<sha> PortableGit PE in place left it handle-held (Defender
on-execute scan / the stub's RunProgram child chain) past pm's ~2 s
_remove_entry retry, so download cleanup raised WinError 32. Review
(teknium1, 5 threads) adds the rest:
- Git.unpack now copies the artifact into a .sfx-* dir beside the
staging tree and executes the copy; pm's .staging-* teardown
(ignore_errors) owns that path, so any hold lands on a disposable
path and the cache dir only ever holds read handles
- refuse off-Windows with the cross-host trade-off stated instead of a
raw PermissionError; documented in package-management.md and the PR
- the extractor is a GUI-subsystem stub that is silent under -y: error
messages now carry the exit code and the usual causes (disk full,
path length, antivirus) instead of promising captured output; the
docstring no longer claims "no GUI" and records that the stub shows
an Extracting window and runs the vendor post-install
- install.ps1: WaitForExit(600000) + Kill() mirrors pm's timeout=600,
Fail reports the exit code and the silence
- tests: the OS-refused-exec assertion and the not-in-text change
detectors are replaced by subprocess-argv invariants (scratch copy
location, exit-code message, off-Windows guard before any execution)
Related to #122512
(cherry picked from commit adaf76a286bebe30c89aee4174df27f1945b60c9)
A `terminal(background=true, heartbeat=N)` tick queued a notification every N seconds
whether or not the process had printed anything, and every queued event costs the owning
session a full model turn. On Desktop and the TUI that turn painted the wake as a user
bubble ("[Background process ... heartbeat #9 ... (no new output since the last
heartbeat)]") followed by the model's "Still running normally." — over and over, for a
process whose row on the status stack already said it was running — and while the wake
held the session's turn, the user's own prompt sat queued behind it.
- `ProcessRegistry._emit_heartbeat` skips a tick with no new output. The sequence counts
delivered beats only; the "(no new output)" placeholder in the formatter is gone.
- TUI/Desktop type heartbeat rows `display_kind: hidden` (the kind both clients and the
transcript preview already honour); the CLI paints a one-line receipt and persists the
row hidden, so reopening the session in Desktop shows only the agent's reply.
- Desktop hydration drops heartbeat rows persisted by older backends the same way.
- `display.background_process_notifications: off` is honored by the TUI/Desktop poller and
the CLI drain, not just the messaging gateway. `off` mutes process-driven wakes only:
a finished `delegate_task(background=true)` still lands.
Supersedes #123123 (cherry-picked; scoped so `off` keeps subagent results) and #119202
(cherry-picked; `heartbeat: 0` is schema-valid so models that materialize every field
stop tripping the foreground guard).
latent-spaces/brag (MIT) turns the project you just built into a short
launch video with music, motion and share copy. It ships two skills:
/brag, the Hyperframes workflow with a bundled music and SFX library,
and /brag-slim, a single SKILL.md where the model builds the whole video
with local tools. /brag hands off to its bundled copy of brag-slim on
Claude Opus 5.5.
Both follow the impeccable/archify pattern: catalog stubs whose
metadata.hermes.upstream pointer makes
`hermes skills install official/creative/<name>` pull the live tree
through OptionalSkillSource._fetch_from_upstream. Nothing is vendored.
The brag stub documents that its Hyperframes path loads HeyGen's
hyperframes-* domain skills by name, and that the skills guard blocks
four of the five today (hyperframes-creative scores dangerous).
brag-slim has no such dependency.
Docs: two generated pages, two catalog rows, two sidebar lines.
Credit: Shunit Haviv Hakimi (shunithaviv), upstream author.
HERMES_DESKTOP_IGNORE_EXISTING=1 only wrapped the PATH probe, so a usable
active install or system Python still started a local serve. Skipping those
rungs on every resolve would make the post-bootstrap re-resolve return
bootstrap-needed and start the installer again. The first pass now falls
through to connect/onboarding; the re-resolve keeps the runtime just installed.
Fixes#117682
The previous commit stops auto-routing a non-sk-or- OPENAI_API_KEY to
OpenRouter. The env-var reference still described OPENAI_API_KEY only as a
custom-endpoint key and the non-interactive setup hint listed it as an
OpenRouter alternative. Say what the key now selects so users with an
OpenRouter key in OPENAI_API_KEY know to move it to OPENROUTER_API_KEY.
Co-authored-by: notwitcheer <notwitcheer@users.noreply.github.com>
They are not "never auto-repaired": resuming such a session re-pins the
full tool surface (restore_agent_tool_prefix), after which a scan sees
skill_manage without the Skill Safety guidance and clears the prompt.
The repair-prompts detector cleared two legitimate prompts:
- a pin with skills_list/skill_view but no skill_manage and zero skills
installed: build_skills_system_prompt returns '' and SKILLS_GUIDANCE is
only emitted with skill_manage, so the healthy prompt has neither marker;
- an exact memory-only pin, which is also a user toolsets=[memory] config;
the healthy rebuild was re-flagged on every run (not idempotent).
Key the decision on the missing '## Skill Safety' guidance, which is
unconditional when skill_manage is in the pin, and require skill_manage in
the pin. Memory-only rows are now reported as unverifiable; the explicit
SESSION_ID override still clears them and their pin.
Docs: describe the tightened rule and note that a running gateway keeps
cached prompts in memory, so it must be restarted after --apply.
Merging an owned category per skill means a skill the author removes from
the distribution is no longer deleted on update, matching top-level
skills/. Say so in the English and zh-Hans reference so authors are not
surprised that retired skills linger.
A non-elevated Windows ARM64 install threw "run setup-hermes.ps1 in an
Administrator PowerShell" whenever Visual Studio ARM64 C++/Clang were
missing. Interactive runs now launch the signed VS installer through a
UAC prompt; CI, ssh and scheduled runs keep the explicit instruction.
Browser tools find agent-browser only in PM's store or on PATH and their
readiness check never installs it, so a fresh install silently had no
browser_* tools. Package.default marks an optional package that a bare
`pm install` also carries; a failed download of it warns instead of
failing the install. `pm install --without NAME` records the opt-out in
declined-packages.json beside PM's install state, and naming the package
explicitly clears it. agent-browser and chromium gain a Termux gap: Termux
owns its browser stack and there is no bionic Chromium.
Contributors had to run `python -m pm.build_env --source . --lock-only`, then
re-source activate, then call sync_venv for opt-in extras, while `hermes pm
lock` only pinned tool artifacts in pm/lock.json. Now `hermes pm lock` with no
arguments is the one step: it runs the same PM check_project_lock /
lock_project operations as build_env's --check-lock / --lock-only (same
exclude-newer quarantine), writes nothing when the lock is current, changes
no environment, and prints the activate command to run next. Activation
syncs [all] plus recorded extras only, and --test-extras replaces the
default, so a newly added extra outside [all] gets
`source ./activate --test-extras all,NAME` (`-TestExtras` on PowerShell).
`hermes pm lock --bump NAME VERSION` keeps writing pm/lock.json; NAME and
VERSION are now only accepted together. build_env --lock-only is unchanged
for scripts. Contributor docs, AGENTS.md, CONTRIBUTING.es.md, the pyproject
comment, and the uv.lock CI remediation text point at `hermes pm lock`.
The image refused every lazy install (HERMES_DISABLE_LAZY_INSTALLS=1), so
edge-tts and the other opt-in SDKs could never be installed at runtime. PM
never writes the sealed /opt/hermes/.venv: it builds a generation under
$HERMES_HOME/installs and commits it in facts.json there, which already
survives container recreates and image updates. Drop the refusal.
Surviving updates means a new image boots under a selection resolved against
the previous image's lock. refresh_dependencies() re-resolves the recorded
extras and plugins against the current inputs; if that fails (offline), it
deselects the generation so the image's own environment boots, keeping the
extras recorded for the next boot or install. stage2 runs it as hermes before
any service starts, then collects generations nothing selects any more, since
nothing else collects them automatically and each one is a full venv.
HERMES_DATA_DIR_SUFFIX, HERMES_REPO_URL, HERMES_UPDATE_STATUS_FILE and
HERMES_UPDATE_UI_ACTIVE are set by Hermes' own bundles, installers and
update shim, not by users, but only the suffix was documented and it
read as a user knob. Add an "Internal bridge variables" table that
says who sets each one and why it cannot live in config, and move the
suffix row there.
`hermes import` only checked the central directory (`is_zipfile`,
`namelist`), so an archive with one member whose deflate stream or CRC is
rotten passed validation and blew up mid-restore with a zlib.error
traceback -- after config.yaml and everything before the bad member had
already been replaced, with later members never written (#121258).
Add a pre-flight pass in run_import that streams every member through
1 MiB reads (zipfile verifies the CRC at EOF) and collects every
BadZipFile / zlib.error / EOFError. If any member is damaged the command
prints a capped list and returns 1 with the home untouched. Stdlib
`ZipFile.testzip()` is deliberately not used: it lets zlib.error escape
and names at most the first bad member.
Widen the per-member catch from the previous commit with EOFError so a
member that rots between the two passes still becomes a "skipped" warning
+ `Import incomplete` / exit 1 instead of a traceback. Rework that
commit's test to drive the per-member path (the archive now never reaches
it with a corrupt member), and let `_break_member` serve the pre-flight
read before failing the restore's own read.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
`_import_members` catches the PermissionError/OSError of each member it
cannot publish and records it in `errors`. That includes the state.db the
live-safe restore refuses because a gateway or dashboard still holds it
(#100960, #110179). `run_import` printed those under "Warnings (N files
skipped)", then printed "Done. Your Hermes configuration has been
restored." and returned None. `cmd_import` did not forward a return value,
so the shell status was 0. A script chained on `hermes import --force`
carried on over a partial restore, and the dashboard's import action
showed the green "done" badge. For the refused state.db case, the user's
sessions were never restored.
`run_import` now returns 1 when `errors` is non-empty. It words both the
summary header and the final line as "Import incomplete", and `cmd_import`
forwards the return code, the same contract `hermes backup` has for an
incomplete archive. Runtime files the import deliberately keeps
(gateway.pid, SQLite sidecars) and the older-backup session warning stay
warnings. The gateway revive still runs, so the files that did land come
up as before.
Measured on main with a fresh HERMES_HOME through `main()`: a 3-file
backup with one read-only target directory, and a refused state.db (live
DB left at 3 sessions / 12 messages). Both used to exit 0 after "Done ...
restored". Both now exit 1 after "Import incomplete: 1 file(s) were not
restored". A clean import still exits 0 and prints "Done".
(cherry picked from commit a0f808ec9f868128f4174e4816e095ac17f18d28)
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Follow-up to the salvaged mirror commit.
- Opt-in (`mirror_to_session: true`, or `hermes webhook subscribe --mirror-to-session`),
default off like cron's `mirror_delivery`. The mirrored text lands with user authority in
the target chat, and on `deliver_only` routes it is the raw rendered payload, so a route
author has to ask for it. Only a real boolean `true` opts in (a YAML string "false" no
longer does).
- The mirror runs inside the routed profile's scope, so a `/p/<profile>/` delivery writes
into THAT profile's state.db. A Telegram DM chat_id is the user's id on every bot; an
unscoped mirror landed in the default profile's DM with the same person (live-probed).
- Tests trimmed to two invariants against a real state.db: an opted-in delivery lands in the
routed profile's chat session and not the default profile's; a route without the opt-in
(including `mirror_to_session: "false"`) never touches the target transcript.
- Docs: webhooks guide section "Replying to a delivery" (default, profile scope, trust note),
CLI reference row, bundled hermes-agent skill reference.
cryptography ships no win_arm64 wheel, so every Windows ARM64 venv sync
compiles it from the sdist and needs MSVC, Clang, Rust and static OpenSSL.
Only setup-hermes.ps1 (and so activate.ps1) prepared that environment,
between a `pm install --tools-only` and the real sync. install.ps1,
`hermes update` and repair ran the same sync without it and failed in
openssl-sys.
PM owns the sync, so PM prepares it. pm/native_build.py holds the adapter
(moved from scripts/build/windows_deps.py) plus source_build_environment(),
which prepares only on win32-arm64 when the synced project carries the
provider script. A payload has prebuilt dependencies and needs no compiler.
VenvPackage.apply and build_environment pass the result to uv children
only. It carries the bridged pip index settings, which managed_environment
applies only to the ambient environment. The state root stays the store
parent, so existing vcpkg/OpenSSL builds are reused.
setup-hermes.ps1 collapses to one `pm install`: the tools-only split existed
only for this preparation, and pm install already puts its tools on PATH
before the venv sync (pm/cli.py activate check).
Not yet verified live on Windows ARM64.
Follow-up to the salvaged wrapper: the argument layer now lives in a topical
sibling (agent/legacy_cli.py) instead of growing the run_agent facade, and
`python run_agent.py` (the installer's PATH launcher for hermes-agent) routes
through the same parser instead of fire, so `--version` and a bare invocation
no longer run the demo turn there either.
- --help/-h/--version use argparse's own exit path; options carry help text
- the runner is injected (`run=`) so `python run_agent.py` does not import
run_agent a second time
- tests trimmed to two invariants that resolve the console-script target
from pyproject exactly like pip's wrapper (red on main, green here)
- docs: hermes-agent section in the CLI reference
OpenRouter relays OpenAI's "this user has been blocked for a previous
policy violation" as HTTP 200 + an SSE error event. The SDK raises a
status-less APIError that no classifier rule matched, so it landed in
the retryable `unknown` bucket: every retry was re-sent (3 streamed
requests by default), a configured fallback only engaged after the
ladder, and the user was told the provider "looks temporarily
unavailable. Wait a minute and send /retry".
- error_classifier: match the ban phrase status-agnostically in
_provider_special_cases -> provider_policy_blocked (non-retryable,
fallback, no credential rotation: every key on a banned account is
banned; a 403 variant is not a bad key).
- turn_failure_copy: provider_policy_blocked chat copy now covers an
account block as well as data/privacy settings; add its cause gloss
so cron and subagent notices explain it instead of printing raw text.
- cron: provider_policy_blocked action (retrying won't help; pin another
model).
- docs: FAQ entry for the reply; auto-recovery ladder exclusion list.
source ./activate and . .\activate.ps1 always run setup, and setup always
runs `python -m pm.cli install`. An explicit install re-hashes every
published tool entry so it can repair a corrupted one. On this machine
that hash reads about 550 MiB and takes 6.0 of the 6.6 seconds, even
when nothing changed.
Shell activation now passes --trust-recorded and trusts the digest the
install recorded, the same check startup already uses. A missing tool
is still installed and a stale venv is still rebuilt. `hermes pm
install` and `hermes update` keep the byte check, and the flag refuses
names, --extra, and --target so it cannot narrow an install someone
asked for by name.
Measured on this machine, already up to date: the activation path drops
from 6.6 s to 0.9 s. A bare `python -m pm.cli install` stays at 7.2 s.
The remaining 0.9 s is process startup and imports, not hashing.
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:
- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
the credential pool, not the environment. Chat and image_gen/openrouter find
it through resolve_runtime_provider. video_gen reported OpenRouter
unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
routed profile's video jobs were submitted, polled and downloaded with the
launch profile's key and billed to that account. A profile whose key lived
only in its own .env could not use the backend at all.
The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.
OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
--clone carried memory.provider (e.g. hindsight) in config.yaml but not the
provider's own config under the profile home (hindsight/config.json,
mem0.json, ...), so the clone booted with the provider selected and silently
unavailable. Copy the ACTIVE provider's <provider>/ dir and/or <provider>.json
by the same convention the dashboard memory-provider routers read, guard the
name against path traversal, tighten copies to 0600 (they can hold an API
key), and say so in the CLI notice. Convention-based on purpose: hindsight is
a catalog plugin now, so a hook the plugin must implement could not fix the
reported case, and no provider module is imported during profile create.
Supersedes #43107 (credit @bionicbutterfly13 for the direction).
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.
Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.
The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
tools-reference.md listed manage_catalog under the connections section, which the
reference-docs contract resolves against the connections toolset. The agent-runtime
post-hook ownership test enumerates every tool in AGENT_RUNTIME_POST_HOOK_TOOL_NAMES;
manage_catalog now runs through both executor paths there.
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog:
- search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with
whether each is already installed in the default profile. Read-only.
- install: opens the same connection operation manage_connections opens, with rows of kind
plugin / skill. Nothing installs until the user approves a row. An approved row installs
into `default` (or the profile the Advanced modal named) through dashboard_install_plugin /
the hub's headless install, so the catalog pin, kill list, security scan and live
activation (#119644) are the host's. The row settles with the live MCP tool names and the
plugin's skill.
The model sends catalog ids and an action only; every other key is refused before anything
runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own
text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the
result is the `hermes plugins install` / `hermes skills install` pointer.
- contract: plugin/skill targets follow the MCP transitions.
- run.apply_answer / reissue route a card answer to the module that owns the operation.
- tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never
deferred behind tool_search.
- docs: tools reference, toolsets reference, plugin catalog page.
Linear NS-964.
The setup toolset (empty until NS-964 registers request_catalog_install)
is reserved for the profile whose backend-written profile.yaml carries
role: setup. Two points enforce it:
- Grant: tui_gateway/server.py::_load_enabled_toolsets folds the in-scope
profile's role toolsets into all three return paths (configured CLI
toolsets, coding posture, HERMES_TUI_TOOLSETS pin), next to the
client-surface set. The pin keeps it too: an operator pin picks
configurable toolsets and must not strip the profile's own.
- Deny: model_tools._select_tool_names strips toolsets reserved for any
other role from every selection, including None/"all", a saved
platform_toolsets list, profiles.configure and the env pin. That is the
one point every surface's selection passes, so a default session cannot
get the tool by any route.
The role is read with read_profile_meta on get_hermes_home(), which under a
session's home override is that session's profile dir (no directory scan).
setup stays out of _HERMES_CORE_TOOLS, CONFIGURABLE_TOOLSETS and the
platform-native recovery loop (it has no tools, so the loop skips it).
Linear NS-963.
Windows PowerShell 5.1 returns every match from Get-Command. .Source on
that array joins the paths with a space, and the call operator then treats
the joined string as one program name. Git for Windows ships git.exe in
cmd\ and bin\, so setup died with CommandNotFoundException before pm
install ran.
setup-hermes.ps1 now installs the tool closure first (`pm install
--tools-only`), then prepares the ARM64 compiler environment, then syncs
the venv. The sync inherits that compiler environment. A bare `pm install`
and the update takeover path publish tools and put them on PATH before
uv sync. A missing tool stops the sync. A missing venv does not.
Verified: scripts/run_tests.sh on test_install_default_closure.py,
test_install_extra.py, and test_windows_build_deps.py — 13 passed.
Activation defined PATH and PYTHONPATH, but `hermes` still resolved to a
global command or an MSIX alias. The shell now defines `hermes` as a
function for this worktree. It runs this checkout's CLI only while the
shell is inside the tree, and it refuses outside it. The prompt gains a
branch prefix and drops it outside the tree. `deactivate` removes the
function and the prefix.
Since the first-strike escalation, a closed transport (socket_closed /
client_closed) forces the reconnect on the first unhealthy sample and the
failure threshold only applies to soft signals (ack staleness, latency,
event silence). The user guide (website/docs/user-guide/messaging/discord.md:89)
and the env-var reference (website/docs/reference/environment-variables.md:780)
still described the threshold as gating every unhealthy sample, so an
operator reading `1/2` followed by a forced reconnect would think the
knob was ignored. One sentence each, citing #118487.