update_stage rebuilt the home from HERMES_HOME with a literal ~/.hermes
fallback while update_lock resolves the same marker through
get_process_hermes_home(). Where the platform default differs (sudo invoker,
data-dir suffix) the two disagreed and the old-shim UI fallback found no
marker, leaving the hand-off window frozen. hermes_constants is stdlib-only,
so the lazy import keeps this module usable under -I -S -B.
prepare_launch treated "dependencies current" as "update finished", but the
sync commits the generation before the product builds and the maintenance
run. A crash between the two left a current install that never built anything
and never would. The tail is now owed by a pending marker written before the
sync and cleared after the tail succeeds, so a repeat launch finishes it.
That tail imports the application, whose entry point is this same function:
inside the launching process's update-lock claim (holder pid is our ancestor)
prepare_launch is a no-op, otherwise the marker recursed forever. The claim
also keeps `hermes update` from racing the launch-time completion.
Completion output goes to stderr: it runs in front of whatever the user
typed, which may be emitting machine-readable stdout. Metadata queries
(--version, -V, --help) skip completion entirely; they read no dependencies.
A completion that fails offline no longer exits 1 from hermes_bootstrap: the
previous generation is still selected (a failed sync commits nothing), so
warn, point at `hermes update`, and launch.
Every update route now finishes through update_completion._complete_selected,
which restarted the whole fleet unconditionally -- including the "Already up to
date" route that main sent through the pending-restart catch-up. Net effect:
each cron tick and each profile's `hermes update` drained and re-killed the one
multiplexed gateway.
Port the catch-up path's two live guards into the completion tail:
- host_restart_already_completed(checkout sha): a sibling profile attaches to
the restart this host already stamped (#95294); the restart phase now stamps
it via mark_host_restart_completed.
- every planned runtime AND every live fleet row current at the checkout sha
(#117051, d6b0d37ece). Both are required: the live matrix lists gateways
only, so a planned serve still on pre-update code keeps the restart.
The already-current route also arms the host obligation with the checkout sha;
an SHA-less arm replaced the standing record and wiped the restarted proof.
The config comment and both docs pages named six of the eight sources in
MACHINE_PACED_SOURCES; "tool" and "batch" were missing, so an operator
reading the docs could not predict the tier those sessions get.
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.
Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".
Every tool schema rides every API call; the first draft read like docs (111 est. tokens).
Kept the two facts the model needs — when to use it and what it delivers — the rest lives in
tools-reference.md.
The execute_code sandbox refuses background/notify modifiers on terminal(); heartbeat implies
notify and rides the same delivery path, so it joins the blocked set (and the stub-drift test's
mirror of it).
`terminal(background=true, heartbeat=N)` emits a `heartbeat` event every N seconds
(floor 60) carrying only the output produced since the previous one, plus the usual
completion notice. The agent stays current on a long bounded job (merge train, full
test suite, deploy) and reacts to a failure within N seconds instead of at exit.
Why: `notify=true` fires once at exit, so multi-hour merge trains ran silently — the
parent's prompt cache went cold and every "what's up" cost 5–10 rediscovery calls;
30 days of batch sessions show 4,930 hand status polls (14% of all tool calls).
`notify=[pattern]` cannot serve this: its lifetime cap (8) exists precisely to stop
periodic output from flooding the agent.
Mechanism: one daemon timer thread for every heartbeat session (reader threads block
on the pipe and cannot keep time); delta output via a total-ingested counter over the
rolling buffer; heartbeats only where the completion notice can be delivered (same
async-support/subagent gates), never after exit. `heartbeat_seconds` is checkpointed.
Rendered on CLI, gateway and TUI through the existing process-notification paths.
The merge landed main's `minimize-to-tray` and `hud-modifier-types` modules next to
ours, leaving two `perfectionist/sort-imports` errors — the only blocking failures in
the `JS & TS checks` job (`apps/desktop :: check:lint`):
electron/main.ts:321 ./minimize-to-tray before ./native-auth-decisions
src/global.d.ts:6 ../electron/hud-modifier-types before ../electron/machine-profile
Resolved with the workspace's own `npm run fix` (eslint --fix + prettier), which also
pads the `padding-line-between-statements` warnings it can. Lint is 0 errors after.
The runner's per-run temp roots live under a fixed `/var/tmp/hermes-pytest`, chosen for
real reasons (disk-backed, not hidden, short enough for AF_UNIX sun_path). But a fixed
literal in a world-writable sticky dir belongs to whoever creates it first: a root-owned
root — a container or system-service run — makes every later `makedirs`/`mkdtemp` there
fail with EPERM for every other user on the host, with no way back that does not need
root. That is exactly what happened on luna: /var/tmp/hermes-pytest is root:root 755, so
every local suite run by the login user died at `runner crashed: PermissionError(13)`
before collecting a single test.
Key the name by uid (with the same non-/var/tmp fallback), so no run can be blocked by
another user's leftovers. Invariant test: two uids never share a scratch root, and the
root is created under the expected parent.
Conflicts, all inside the local-runtime / local-models subsystem:
- hermes_cli/local_runtime/binaries.py, hermes_cli/web_routers/local_models.py:
main's own transfer machinery (urllib ranged download, shutil.move publish) is
retired here — pm/downloader owns every fetch. Took ours.
- tests/hermes_cli/test_local_runtime_downloads.py: deleted here (its subject,
binaries._download, is gone); kept deleted.
- tests/hermes_cli/test_local_models_routes.py: ours (the pause/resume suite) plus
main's held-finished-file contract, re-seamed onto the pm downloader.
main's fix 515412f415 ("wait out a held finished download instead of copying it and
failing") is re-homed so it fits the PM model: the Windows rename that fails while an
antivirus or indexing scan still holds a just-closed multi-gigabyte file is now handled
once, in pm/downloader.replace_when_released, which is the single publish path for PM
archives and model files alike. A cleanup that cannot remove its leftover can no longer
mask the error that left it.
Tests: helper retries a transient refusal and lands the file; gives up with the
plain-language message chained to the OS error; a download publishes through two refused
renames with nothing left behind; a stuck hold reports the rename error, not the failed
cleanup. Route-level: the job lands the file and reports done.
At 601px the select columns were squeezed to 89/103/113px while the selects
are floored at 120px, so each control overran its neighbour (row scrollWidth
603 vs 569 clientWidth). Hold the same floor on the column, letting the
search field absorb the shrink instead.
The plugins panel had the same viewport-drift the drawer had: leaving the
mobile band left `filtersOpen` true, so `aria-expanded` stayed "true" on the
hidden toggle. Reset it from the same media query.
The new compact controls hid `.sidebarToggle` at every width — the 1024px
block's `.layout > .sidebarToggle` rule outranks the later 900px
`display: flex` — so its button, `activeCatBadge` and three CSS blocks were
unreachable, and the drawer's `onKeyDown` was shadowed by the focus trap's
capture-phase listener. Delete them.
Also drop declarations that cannot paint under `display: contents`
(`padding`/`border-bottom`/`backdrop-filter`/`background !important` on
`.controlsBar` at <=600px), the redundant `.compactSort` hide (its parent is
already hidden), and the empty effect left over on the plugins page.
`.filterPanel.filterPanelOpen` keeps the open panel working regardless of
source order instead of relying on the 600px block coming last.
Tapping Filters scrolled 391px even when the panel was already on screen,
because the handler called `scrollIntoView({ block: "start" })` on every
open. Reveal it from an effect instead, with `block: "nearest"`, and keep the
click handler a plain state toggle (the cancelled rAF no longer rides along
on the handler).
Opening the filters drawer at <=600px and then resizing to a wider viewport
left `sidebarOpen` true: the hidden toggle kept `aria-expanded="true"`, and
coming back to mobile silently reopened the drawer and re-locked body scroll.
Reset it in the matchMedia sync, and name the query constant so it stays in
step with the `max-width: 600px` CSS blocks it mirrors.
`?embed=picker` hides the navbar and pins the controls bar to the frame top,
but that override still names `.controlsBar` — which is `display: contents`
at <=600px, so the sticky element (`.controlsTopRow`) kept its 60px navbar
offset and left a dead 60px gap with cards sliding under it. Extend the
override to the row, and add the same rule to the skills page, which had no
picker offset at all.
Between 601px and ~870px the compact selects were squeezed to as little as
13px of text area (SOURCE needed 64px), so the control displayed no
readable value while the pills it replaced were hidden — the band had no
way to see the active filter.
Floor the selects at 7.5rem, let the search field shrink to 8rem instead of
12rem, and drop the uppercase labels below 820px where they leave no room.
The new sticky control bar hardcoded the dark surface
(`rgba(7, 7, 13, 0.94)` plus gold borders and a gold-on-gold count badge),
but the site respects `prefers-color-scheme`, so light-mode users got dark
text on a near-black bar: the Filters label measured 1.00 contrast (1.5 by
pixel sample) and the skills search input dropped from 3.34 to 1.12.
Use the theme tokens instead — `--ifm-navbar-background-color` and
`--ifm-hr-border-color` resolve to the previous dark values exactly — and
scope the gold gem/border accents to the dark theme, with
`--ifm-color-emphasis-*` for light mode.
One test pins the contract (if/then/else + oneOf/not fragments survive
normalization and validate with the intended semantics, using the NBX
set_effects shape from #107141), one pins the guard (root and declared
nested objects still get the #4651 dangling-required repair).
_repair_schema didn't recurse into if/then/else, so a bare conditional
wrapper (e.g. an allOf branch expressing "when goal is set, mode must be
one of ...") fell through _fill_missing_type's default and was stamped
with type: "string". That corrupts the schema: the wrapper's actual
instance is an object, so Moonshot (and any strict validator) now
rejects every real argument against it.
if/then/else are added to the recursed node keys, and _fill_missing_type
leaves a bare conditional node untyped instead of defaulting to "string"
since it constrains the parent instance rather than describing its own
shape.
_repair_object_shape() treated every dict carrying `required` as an object
declaration. A constraint fragment — {"required": ["chain"]} inside an
allOf/oneOf/anyOf/if branch, with no properties and no type — is not one.
Repairing it synthesised `properties: {}` and then pruned every name out of
`required`, so sibling branches collapsed into identical always-true schemas
and the enclosing oneOf had two matches: the tool appeared in the catalog and
every call failed validation.
Only the parameters-schema root still receives the dangling-`required` repair,
which is the single node handed to providers as the argument object.
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.
Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.