Commit Graph

3037 Commits

Author SHA1 Message Date
Teknium
7c799a6565 feat(vercel): fresh sandboxes use a managed image (universal:latest); runtime presets deprecated, migration 49 (#126741)
* feat(vercel): start fresh sandboxes from a managed image instead of the deprecated runtime

Vercel deprecated Sandbox runtimes (node24/node22/python3.13) in Aug 2026 in favour of
images, and rejects runtime+image together and runtime with a snapshot source. New
terminal.vercel_image (default vercel/sandbox/universal:latest, Node 24 + Python 3.14)
picks the image for fresh sandboxes; a pinned terminal.vercel_runtime still works, wins
over the image and logs a deprecation warning; snapshot restores send neither.

Setup wizard prompts for the image, dashboard exposes both keys, status/config show the
effective choice, TERMINAL_VERCEL_IMAGE bridges config to the tool like its siblings.

* feat(config): migration 49 drops the seeded node24 Vercel runtime pin

Every pre-49 config.yaml carries terminal.vercel_runtime: node24 (the template default) and
the setup wizard mirrored it into .env as TERMINAL_VERCEL_RUNTIME. Both are the default
copied, not a choice, so the migration drops them and fresh sandboxes follow vercel_image;
node22 / python3.13 pins are the user's and survive. Persisted sandboxes are unaffected:
a snapshot restore never sends a runtime or an image.
2026-09-28 12:33:01 -07:00
Teknium
3f871425af fix(modal): persistent sandbox snapshots no longer expire after 30 days (modal 1.5.5, ttl=None) (#126740)
* fix(modal): keep persistent-sandbox snapshots past the SDK's 30-day TTL

modal>=1.5 gives Sandbox.snapshot_filesystem() a default ttl of 30 days, so an idle
persistent Modal sandbox silently lost its filesystem and restarted from the base image.
Pass ttl=None (retain until deleted) and bump the modal extra from 1.3.4 (no ttl
parameter; legacy RPC) to 1.5.5 so the kwarg exists on every install.

* fix(modal): drop the dead modal.Mount credential-mount block

modal.Mount left the public API in modal 1.0, so _modal.Mount.from_local_file raised
AttributeError into the surrounding except on every sandbox start and the block never
mounted anything. The FileSyncManager created right after already uploads the same
credential, skills and cache files (iter_sync_files), so delete the duplicate; the test
fake stops exporting a Mount the real SDK does not have.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 12:31:03 -07:00
teknium1
a2cdaa7388 fix(review): run TTS warm/release command hooks under the caller's secret scope (#120481)
The warm_command/release_command hook thread was a bare threading.Thread, so under
the multiplexer resolve_passthrough_value() saw no secret scope, raised
UnscopedSecretError, and the best-effort hook swallowed it at debug level: every
command provider with a non-empty env_passthrough silently never warmed or released.
Bind the thread with ctx_bound (copy_context().run), the same seam the keep-warm
timer already uses. Regression test: the hook thread resolves the bound profile's
passthrough value.

Docs (review minor): the multiplex isolation table now names per-profile slash-command
gating and the fail-closed empty-admin policy for a served profile with no cached config.
2026-09-28 12:23:25 -07:00
teknium1
2f0b102ce3 fix(review): /api/profiles/active fails closed on current; document the hub-action secret scope
Review finding 2: `_or_default` in `get_active_profile_endpoint` answered
"default" for `current` when `get_active_profile_name()` raised — exactly the
value that lets the SPA's `shouldAdoptActiveProfile` retarget a dashboard to
the sticky active profile. A dashboard that cannot name its own home must not
read as the machine dashboard, so the failure fallback for `current` is now
"custom" (the same adoption-refusing answer the function itself gives an
unresolvable home). The empty case keeps `or "default"`, and `active` keeps
"default" on failure: "custom" there would itself trigger adoption of a
non-existent profile when `current == "default"`.

Review finding 1 (docs half): hub actions targeting `default` from the
machine dashboard now take the same scrubbed, HERMES_HOME-pinned environment
as every named-profile action. Kept on purpose (one env contract per hub
target); the user-visible rule — children run with the target profile's own
.env and secret sources, not the dashboard process environment — is now in
web-dashboard.md. The PR body carries the explicit behaviour-change note.
2026-09-28 12:18:58 -07:00
John Paul Soliva
9ad7156324 fix(dashboard): cron delivery targets, blueprints, recommended default, voice config, official skills and debug share answer an unknown ?profile= with 404
Each route's catch-all or fallback swallowed the profile scope's 404: the
GETs answered 200 with default or empty data (local-only targets, static
deliver options, an empty model, relay voice mode, no installed marks) and
POST /api/ops/debug-share answered 500. HTTPException now passes through
ahead of each catch-all, as in the rest of this change; for the hub skills
catalog the pass-through sits in _installed_hub_identifiers, so the search
and sources routes that share it keep their existing 404.

The unknown-profile test gains a row per GET route; the endpoint list in
web-dashboard.md names them.
2026-09-28 12:18:58 -07:00
John Paul Soliva
c2e4772b13 docs(web-dashboard): list learning graph and plugins hub as profile-scoped 2026-09-28 12:18:58 -07:00
alt-glitch
fc4dbe32df fix(logs): hermes logs --since/--level handle unstamped lines and MCP output
`hermes logs --since` and `--level` passed every line that had no leading
timestamp. A traceback's frames are written without one, so an old error
printed its frames without the header that was filtered out, and
`--component` dropped the frames of a matching record. `hermes logs` now
reads each unstamped line as part of the record above it: the line gets
that record's verdict for every filter, in the tail read and in `-f`.
Lines before the first stamp in the read window have an unknown time and
level, so they are dropped when `--since` or `--level` is set.

mcp-stderr.log had no parseable stamp at all: the banner started with
`=====` and server output was copied raw, so `hermes logs mcp --since`
printed the whole file. The stderr tee already reads each server's
stderr in a thread, so it now writes one line at a time, each prefixed
with the asctime-shaped local stamp the Python logs use. The banner
starts with the same stamp. The stamp comes from new public
`timestamp()`/`stamp_line()` in hermes_cli/stderr_timestamp.py, which
stays stdlib-only. The desktop MCP log view accepts both banner shapes.

A test now requires a real writer's sample line for every LOG_FILES
entry to parse with `_parse_line_timestamp`. The docs no longer say the
`timezone` key changes log timestamps; log lines use the machine's
local time.
2026-09-28 19:02:20 +05:30
John Paul Soliva
ef3fa2a6f9 fix(profiles): rename removes the old name's gateway service even when stopped (#124401, salvage #124402)
rename_profile removed the launchd/systemd unit only when the gateway
was running, and never touched the s6 slot. A unit installed under the
old name but stopped stayed behind: it runs --profile <old> with
HERMES_HOME pinned to the moved directory, so the next login or
container boot crash-looped it for a profile that no longer exists,
and deleting the renamed profile never found it.

The old unit is now removed whether or not the gateway runs, the s6
slot moves to the new name, and the command prints how to reinstall
the service under the new name.

(cherry picked from commit 76002ea29ab8f8871f2f962ef4535408f29c590f)
2026-09-28 04:04:23 -07:00
John Paul Soliva
8e9d1b90c9 fix(docker): boot a gateway.standalone profile's own s6 slot at container start (#120524, salvage #120525)
The boot reconciler (`reconcile_profile_gateways`) registered every named
slot down and folded any slot's `running` intent into the root slot. The
root multiplexer never serves a `gateway.standalone` profile
(`profiles_to_serve` excludes it), so after every container restart that
profile had no gateway while the boot log claimed the root served it.

A standalone profile's slot now boots from its own intent and is kept out
of the fold; every other named slot stays registered down. The reconcile
comment and the fold notice name the exception, and the standalone docs say
its listener needs its own port beside a host gateway that enables the same
one.

`gateway.parked` is deliberately NOT consulted here: parked is orthogonal to
standalone (a host-served profile the operator took offline), the in-process
migration path (`gateway_migrate`) already leaves standalone-by-config
profiles alone, and the root multiplexer skips parked profiles itself.

(cherry picked from commit 1b28dee46d96e5509bd00b567c4658f5e47600d1)
2026-09-28 04:04:23 -07:00
teknium1
37daf85b2a fix(review): PR acceptance never falls through to the launch user's gh login
Review findings on the assignee-login gate (#122689):

- MAJOR: served_profile_child_env(inherit_credentials=True) overlays only the
  assignee's own GH_TOKEN/GH_CONFIG_DIR but HOME/XDG_CONFIG_HOME stay the
  launch process's, so a profile with no login of its own fell through to
  ~/.config/gh/hosts.yml - the ambient login the gate promises never to use.
  For a routed assignee home with neither token nor config dir, pin
  GH_CONFIG_DIR to <profile_home>/gh; gh then exits 4 (authentication
  required), which _api classifies as auth naming the profile.
- MINOR (a): an ASSIGNED card whose profile cannot be resolved was fail-open
  (env=None -> completing process's full ambient login). Resolution now
  happens inside collect_acceptance and raises _GateAuthError naming the
  profile; only genuinely unassigned cards keep the ambient path.
- MINOR (c): Codex refresh 200/non-JSON body no longer sets relogin_required,
  so the new exit-78 startup gate cannot turn an edge misfire into a sticky
  terminal block (invalid_json_relogin=False, as xAI already does).
- MINOR (b)+(d) docs: GH_TOKEN must live in the profile's .env/secret source
  (a shell/systemd export is scrubbed); skipped_nonspawnable {assignee} row
  in the worker telemetry table; startup relogin-required failure exits 78.
2026-09-28 03:37:09 -07:00
Yuan Li
fdac8a1896 fix(kanban): PR acceptance runs gh as the assignee profile's login
collect_acceptance's gh api subprocess inherited the calling process's
environment, so on a multi-profile host the gate read the contract repo as
the ambient (launch/default) gh login: a private repo in another org
returned 404 and the broad except classified it infra with 'check gh
authentication and retry', blocking green cards and holding them via the
blocker_auth respawn guard (#122689).

- Thread the assignee's profile home into every _api call and build the
  child env with served_profile_child_env(inherit_credentials=True): the
  profile's own GH_TOKEN/GH_CONFIG_DIR overlay, launch credential residue
  scrubbed (GH_CONFIG_DIR is a path, not a credential, so no scrub list
  saw it — drop it from the base for routed targets).
- Classify a refused read (gh HTTP 401/403/404, or GraphQL resolving the
  repository to null) as 'auth' naming the repository, distinct from
  retryable infra; persist only the status code + endpoint, never stderr.
- Unassigned/uninstalled cards and single-profile hosts keep the ambient
  env, unchanged.

Fixes #122689
2026-09-28 03:37:09 -07:00
Teknium
399d956903 feat(bot_desktop): Bot Screen, computer_use and the browser run inside the terminal backend (#121169)
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends

The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.

docker/sandbox-desktop.Dockerfile is that base plus:
  - the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
    zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
  - the exact package set the Hermes -desktop image installs (TigerVNC,
    Xfce components, dbus, xauth, fonts)
  - Playwright's headed Chromium (same build as the -desktop image)
  - cua-driver 0.28.2 from its pinned release tarball

No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.

docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.

* feat(docker): sandbox desktop base on python3.13-nodejs26

Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.

* feat(docker): bake agent-browser into hermes-sandbox:desktop

The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.

* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend

A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.

Now the desktop lives where the terminal lives:

- tools/environments/streams.py: one primitive per spawn-per-call backend, the
  local argv prefix that runs its remainder inside the sandbox with stdio open
  (`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
  vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
  the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
  prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
  gateway). `auto` follows the backend; a sandbox that cannot host a screen
  REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
  lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
  there; the host driver's runtime contract is irrelevant then; check_fn is
  true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
  with the daemon, socket dir and profile inside the sandbox; screenshots are
  fetched back so MEDIA: paths keep working; recycle closes the sandbox
  daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
  in a sandbox, a unix socket otherwise.

Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.

* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only

DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).

* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read

Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.

* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order

Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host

sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.

* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox

Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).

* chore: retrigger CI (zero-job dispatch failure)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore(config): template stamps v47 and shows the new default sandbox image

The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.

* feat(sandbox): the default image change is a decision, not a surprise

A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.

Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.

Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.

Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.

* fix(config): both plain defaults that preceded the desktop sandbox image are template copies

main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.

* ci(docker): build the sandbox image on release/dispatch, not every main push

Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.

* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta

The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.

* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)

* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)

* fix(bot_desktop): placement is the authority; sandbox screen survives restarts

Review findings on the sandbox-hosted Bot Screen, each reproduced live first.

Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.

Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.

SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.

CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.

pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.

Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.

Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.

Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.

* docs(bot-screen): no literal tmp path in the profile-location note

* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)

* fix(bot_desktop): adopting a screen the sandbox kept records the marker

Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.

Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).

* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)

* docs(bot-screen): what the Apptainer path inherits from the image and what it does not

* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 03:34:07 -07:00
kshitijk4poor
0d5aae5e24 fix(gateway): warn when config still declares a retired idle/daily session_reset
Time-triggered conversation rotation was removed in 1d5d059410 and core
has read nothing under session_reset since. Nothing told users whose
config still asked for it: their gateway conversations silently stopped
resetting, and a household that never types /new pays for an
ever-growing context.

Gateway startup (for every served profile) and `hermes doctor` now
report a session_reset block whose mode is idle, daily or both, and
point at the hermes-session-reset-policy catalog plugin, which reads
the same top-level block unchanged. The notice stays quiet while that
plugin is enabled. The config key is reported, never rewritten, since
the plugin consumes it. The gateway: form the old loader also accepted
is detected too, with a hint to move it to the top level.
2026-09-28 15:54:01 +05:30
Siddharth Balyan
98278833e2 Compaction follows compression.threshold again: no default 256K token cap (#117915, #125235) (#126064)
* fix(compression): compaction follows the ratio again; no default token cap

compression.threshold_tokens defaulted to 256000 since #115986, so every
window above ~341K compacted at 256K regardless of compression.threshold
or model_thresholds: a 1M-window model compacted at 25% of its window,
and threshold: 0.8 still compacted at 256K. No single token count suits
windows from 64K to 1M+, so the cap goes back to opt-in (null) and the
ratio decides, as it did before #115986.

The template seeder and `hermes doctor --fix` copied the 256000 default
into config.yaml, where it reads as a user choice and would keep capping
those installs. Config v47 drops threshold_tokens only when it equals
256000; any other explicit cap and an explicit null are preserved.

The rest of #115986 stays: the model-switch warning quoting the real
trigger and the shared _derive_trigger are correct with any default.
Tests that encoded 256K now derive expectations from DEFAULT_CONFIG.

* docs(compression): threshold_tokens is an optional cap, default null

User guide, developer guide and delegation page described the 256K
default cap; they now describe the ratio trigger as the default and
threshold_tokens as an opt-in cost ceiling.
2026-09-28 07:29:40 +00:00
teknium1
1d287d5375 fix(browser): ship the Browser Use CLI engine in every install, Desktop included
The default browser_exec tool ran the `browser-use` CLI from a PM side
environment (browser-use==0.13.10 in <home>/environments/browser-use),
provisioned by the installers and `hermes update`. Sealed Desktop payloads
skip that step, so the Desktop app never had it and silently fell back to
the built-in tools; the side env was also per-profile and 225 MB.

The CLI's execution path is only `browser_harness.run.main()`; the
browser-use agent framework (anthropic/openai/google-api pins, 93 MB of
googleapiclient) is never imported. browser-harness itself is 2.6 MB of
pure Python whose pins (Pillow 12.3.0, websockets 15.0.1) already match
Hermes's own, so it becomes a core dependency and runs on sys.executable:

- pyproject/uv.lock: browser-harness==0.1.13 (+ cdp-use, fetch-use).
- _find_cli() returns [sys.executable, -m, browser_harness.run]; the child
  env points PYTHONPATH at the harness site dir (the Desktop store
  interpreter boots without a venv and the harness daemon re-runs
  sys.executable), replacing whatever the agent inherited.
- The side-env provisioning (install_cli, the update/installer step) goes.
2026-09-27 23:53:40 -07:00
teknium1
27062c3474 fix: install cua-driver and the Browser Use CLI by default again
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.

- cua-driver is a default PM package, so the installers, a bare
  `hermes pm install` and `hermes update` carry it on every target it builds
  for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
  and `hermes update` both run, provisions the Browser Use CLI for the default
  and explicit Browser Use backends. `--skip-browser` declines it along with
  agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
  `--without cua-driver`).
- The update message stops calling every default "browser tools".
2026-09-27 19:18:58 -07:00
Brooklyn Nicholson
9cd114b499 feat(web): inline_images=false on GET /api/sessions/{id}/messages
The REST pages carry message content verbatim so the desktop's
extractEmbeddedImages can pull data URIs out of the text; a client reading
over a network had no way to ask for less (26.33 MiB per page read on the
measured conversation). inline_images=false (default true) routes content
through the same _coerce_message_text(image_urls=False) projection
session.resume uses, rendering [image] in place of the data URI so both
history surfaces agree by construction. Documented in the api-server
reference.

Fixes https://github.com/NousResearch/hermes-agent/issues/116511
2026-09-27 19:06:26 -05:00
Brooklyn Nicholson
511a6b1c16 fix(desktop): ⌘1…⌘9 switch the tab under the pointer again, and profiles when there is none
#92569 made profile.switch.N unconditional and shipped view.tabSlot.N
unbound, which threw away the hover → focused → workspace tab dispatch
from #74447: holding ⌘ still painted tab-number hints, but the chord
switched profiles. The reporter's real complaint was that the tab
dispatch was hardcoded inside the profile handler, so rebinding the
chord could not separate the two.

Give the keybind registry a `passthrough` action kind: the combo index
keeps every action bound to a chord in registration order, the
dispatcher runs the first and, when a passthrough handler returns
`false`, hands the chord to the next. view.tabSlot.N defaults to ⌘N
ahead of profile.switch.N and declines when no zone is a real tab
strip or the strip has no Nth tab, so ⌘N is "tab N" over a strip and
"profile N" anywhere else. Either action can be rebound on its own,
the panel does not flag the layered pair as a conflict, and the held-⌘
hints key off the tab action they describe.
2026-09-27 18:00:11 -05:00
M1racleShih
6f7cc7e74c feat(cron): allow Python scripts to use an external interpreter
Add an optional per-job `interpreter` field so a cron Python `script` /
`monitor_script` can run under a user-managed venv instead of Hermes' own
Python, letting scripts import packages the Hermes runtime does not carry
(#8714). Nothing is installed, frozen, or restored automatically.

- cron/jobs.py: persist + normalize the field (absent => record unchanged;
  empty string clears it on update).
- cron/scheduler_script.py: _resolve_cron_interpreter() validates the path
  at run time (absolute/~ required, regular file, executable on POSIX);
  _script_argv runs [interpreter, script] and skips the managed-store
  bootstrap/PYTHONPATH overlays, which exist for Hermes' own venv.
  Threaded through _run_job_script, the claim-heartbeat wrapper, the
  pre-run prompt path and monitor scripts.
- hermes_cli: --interpreter on `cron create` / `cron edit`; shown in
  details and `cron list`.
- tools/cronjob_tools.py: programmatic/CLI lane only, like model and
  reasoning_effort — absent from the model-facing schema.

Shell scripts (.sh/.bash) still always run under bash. Revives #8741.

Ported onto current main from #70500 (the scheduler moved to
cron/scheduler_script.py and the CLI/tool became table-driven since the
PR's base).

Co-authored-by: MestreY0d4-Uninter <241404605+MestreY0d4-Uninter@users.noreply.github.com>
2026-09-28 02:32:05 +05:30
Hermes Agent
c5380053b5 fix(install): keep the ffmpeg lock label, drop the network liveness test, document -SkipSetup
Review fix-ups on top of the re-pin:

- pm/lock.json: keep "version": "9.0.1". pm keys the store entry on
  `ffmpeg-<version>-<target>` and reinstalls on the artifact sha alone
  (pm/install.py::_identity / _entry_current), so the sha change already
  re-fetches the four BtbN targets. Bumping the label would also rename the
  macOS (martin-riedl, still 9.0.1) and Termux entries and re-stage unchanged
  bytes on every existing install for no reason.
- tests/pm/test_ffmpeg_pin_liveness.py: removed. It HEADs live GitHub URLs
  from the unit lane, and BtbN prunes dated autobuild tags after ~14 days
  (autobuild-2026-09-10-15-31 is already gone), so the test turns red for
  every PR on ~Oct 11 by construction. Upstream rot is what
  archive-inputs.yml (sha256 mirror on merge) and #122433 (re-pin on fetch
  failure) are for.
- website/docs/user-guide/windows-native.md + install.ps1 header: say that
  -SkipSetup is accepted as a deprecated alias for -NonInteractive instead of
  claiming it is rejected.

(cherry picked from commit def90331c2; ffmpeg lock/liveness hunks dropped, superseded by #125468)
2026-09-28 00:58:16 +05:30
Brooklyn Nicholson
aa25f9e85f fix(desktop): Kanban board switcher outside the full-page layout
A Kanban board opened in a split route tile rendered no board switcher:
the board contributed it to WORKSPACE_PAGE_HEADER_AREA unconditionally,
and only the workspace pane paints that area. The tile's contribution
also leaked into another page's header and shared its id with the full
page's, so closing the tile removed the page's switcher.

Add WorkspacePageHeaderControl (exported via the plugin SDK). The
workspace pane's render provides a private host context; inside it the
control projects into the page header, anywhere else it renders inline.
The board mounts BoardSwitcher once, through it, in its own header row.

Fixes #123597

Originally authored by Justin Haynes (@jhaynes).
2026-09-27 13:18:46 -05:00
shali10
40523600b0 fix(sessions): refuse to delete a session row a live turn still owns (#123583)
Refactor entry-side deletion refusal to execute in-transaction via
`_write_guards_reject(conn, sid)` (#123583), per maintainer review:

- Underlying `delete_session` and `delete_sessions` now accept an opt-in
  kwarg `exclude_active_write_guards=True` running inside `_do` write
  transaction, eliminating the race condition where a turn acquires the lease
  between check and delete.
- Raises `SessionActiveWriteGuardError` when refusing single delete, leaving
  the row untouched; `delete_sessions` atomically skips active rows.
- Checks both active turn leases and compression locks via the existing
  reclaim-aware `_write_guards_reject` helper.
- Covers all user-facing delete sinks:
  * Web `DELETE /api/sessions/{id}` -> 409 Conflict
  * Web `POST /api/sessions/bulk-delete` -> skips active rows
  * Web / CLI `prune` -> passes `exclude_active_write_guards=True` so lineage
    parents of active conversations are not pruned
  * API Server `DELETE /api/sessions/{id}` -> 409 session_active_turn
  * CLI `hermes sessions delete` & `export --delete-after-verified` -> exits 1
  * CLI browse picker -> refuses active delete
  * TUI Gateway `session.delete` -> 4023 error
- Conforms to rubric with 2 targeted invariant tests in
  `tests/hermes_state/test_delete_session_write_guards.py`.
- Updates user guide and web dashboard docs for 409 / exit 1.

(cherry picked from commit 2c037a7a79dc211b49bacc72e3140951ccf900cf)
2026-09-27 20:49:04 +05:30
kshitijk4poor
1ec84a2dae fix(simplex): document contactId-only allowlist and warn on name entries
After #44729 SIMPLEX_ALLOWED_USERS matches only the numeric contactId, but
the docs still told operators display names work, and existing name
entries would silently stop matching. Update the docs and log a one-time
warning at first connect listing non-numeric entries that are now ignored.
2026-09-27 20:47:41 +05:30
kshitijk4poor
f6ce8bb23b fix(context): assemble compaction head/tail from the pruned copy (#61932)
The salvaged lossless-history change rebuilt the carried head/tail from
canonical history, which undid _pressure_demote_tail's tool-result
shrinking and re-broke #61932 (an all-oversized tail could no longer
compress). Pruning no longer rewrites tool_calls, so the pruned copy's
arguments are already byte-identical to canonical history: assemble the
head and tail from the pruned copy, keeping tool-result demotions and
exact tool-call arguments at once. Docs updated to match.
2026-09-27 18:38:58 +05:30
JoaoMarcos44
a7baa5f5eb fix(context): keep compaction history lossless
(cherry picked from commit d51c8f4f5096badfd0beddd78646617643f6028f)
2026-09-27 18:38:58 +05:30
brooklyn!
9e7239acfd fix(desktop): pass wayland ozone on native Wayland sessions
Native Linux Wayland stayed on XWayland because the relaunch only ran for
WSLg. Append --ozone-platform=wayland when the user did not already choose
a platform. An explicit x11 hint and desktop.electron_flags still win.
2026-09-27 06:26:35 -05:00
Brooklyn Nicholson
8cb4fdc925 fix(process): heartbeats wake the agent only on new output, and never as a user bubble
A `terminal(background=true, heartbeat=N)` tick queued a notification every N seconds
whether or not the process had printed anything, and every queued event costs the owning
session a full model turn. On Desktop and the TUI that turn painted the wake as a user
bubble ("[Background process ... heartbeat #9 ... (no new output since the last
heartbeat)]") followed by the model's "Still running normally." — over and over, for a
process whose row on the status stack already said it was running — and while the wake
held the session's turn, the user's own prompt sat queued behind it.

- `ProcessRegistry._emit_heartbeat` skips a tick with no new output. The sequence counts
  delivered beats only; the "(no new output)" placeholder in the formatter is gone.
- TUI/Desktop type heartbeat rows `display_kind: hidden` (the kind both clients and the
  transcript preview already honour); the CLI paints a one-line receipt and persists the
  row hidden, so reopening the session in Desktop shows only the agent's reply.
- Desktop hydration drops heartbeat rows persisted by older backends the same way.
- `display.background_process_notifications: off` is honored by the TUI/Desktop poller and
  the CLI drain, not just the messaging gateway. `off` mutes process-driven wakes only:
  a finished `delegate_task(background=true)` still lands.

Supersedes #123123 (cherry-picked; scoped so `off` keeps subagent results) and #119202
(cherry-picked; `heartbeat: 0` is schema-valid so models that materialize every field
stop tripping the foreground guard).
2026-09-26 22:19:53 -05:00
Brooklyn Nicholson
33f45ca30b fix(gateway): keep one Windows gateway autostart mechanism
A successful Scheduled Task install returned without removing an
existing Startup-folder Hermes_Gateway.vbs or legacy .cmd, and the
fallback path wrote a Startup entry even while a task was still
registered. Both fire at logon, so the gateway launched twice.

install() now removes Startup entries after the task registers, the
fallback is skipped while a task exists, and reconcile_autostart_launchers()
converges an existing install to one mechanism.

Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
2026-09-26 21:44:50 -05:00
Brooklyn Nicholson
a7c080ca66 feat(skills): brag and brag-slim join the optional-skills catalog as upstream stubs
latent-spaces/brag (MIT) turns the project you just built into a short
launch video with music, motion and share copy. It ships two skills:
/brag, the Hyperframes workflow with a bundled music and SFX library,
and /brag-slim, a single SKILL.md where the model builds the whole video
with local tools. /brag hands off to its bundled copy of brag-slim on
Claude Opus 5.5.

Both follow the impeccable/archify pattern: catalog stubs whose
metadata.hermes.upstream pointer makes
`hermes skills install official/creative/<name>` pull the live tree
through OptionalSkillSource._fetch_from_upstream. Nothing is vendored.

The brag stub documents that its Hyperframes path loads HeyGen's
hyperframes-* domain skills by name, and that the skills guard blocks
four of the five today (hyperframes-creative scores dangerous).
brag-slim has no such dependency.

Docs: two generated pages, two catalog rows, two sidebar lines.

Credit: Shunit Haviv Hakimi (shunithaviv), upstream author.
2026-09-26 21:41:31 -05:00
Hermes Agent
9d09662993 docs(desktop): describe what --ignore-existing skips 2026-09-26 17:15:10 -05:00
Brooklyn Nicholson
959c7649fd fix(updates/win): click-session poll, Intel-Mac installer docs, cua-driver opt-in autostart
click-session flake (#97982): a bare scrollIntoView() smooth scroll could be dropped under load, and the script
slept a fixed 3000ms before reading state. Extract click-session-helpers.mjs:
instant centered scroll, a bounded poll-for-composer loop instead of the
fixed sleep, and correct nested CDP envelope unwrapping (the old read logged
undefined).

Intel-Mac installer docs (#99033): the Hermes-Setup.dmg bootstrap installer
is built for Apple Silicon only, so Intel Macs hit "not supported on this
Mac". The desktop release pipeline already builds a native darwin-x64
bundle, so the docs now scope the arm64 limit to the bootstrap installer and
name the darwin-x64 bundle (or the CLI plus `hermes desktop`) as the Intel
path, in the desktop README and the platform-support build-targets section.

cua-driver autostart opt-in (#97389): Windows installs registered the
cua-driver-serve scheduled task on every install, with no opt-out, and
treated the task as an install-readiness requirement. Gate the
install-ready check and _repair_cua_driver_autostart_windows on the new
computer_use.autostart config key (default false = on-demand, fails
closed), extract the registration PowerShell into a testable helper, and
document the opt-in (EN + zh-Hans). Windows-only code path: unit tests
cover the registration args and the config gate; live Windows
verification pending.
2026-09-26 17:09:16 -05:00
kshitijk4poor
393f03dbcb docs(plugins): say dependency dirs are not carried across catalog updates
The catalog guide said every symlink among untracked files stops the update.
Since guard-excluded dirs (.venv/, venv/, node_modules/, tool caches) are now
pruned from the carry, links inside them never stop an update and those dirs
are rebuilt rather than copied. State that exception next to the symlink rule
(the _carry_user_files docstring was updated with the code change).
2026-09-27 01:42:52 +05:30
kshitijk4poor
03fae2fac8 fix(plugins): refuse symlinked user files on git-checkout updates
cff26600d2 stopped following symlinks when carrying untracked/ignored
files into a staged catalog update: a link planted after the installer's
scan could point outside the plugin root, past the guard. But it did so
by skipping them silently. Base followed the link and kept the content,
so a user whose ignored config.yaml is a link into their dotfiles now
loses that config on repin with no warning. _carry_user_files promises
to fail before publication rather than drop user state.

In a git checkout, symlinked files or dirs in the ??/!! set now fail the
update before publication, and the error names every such path. Links
are still never followed. Links under node_modules/ are .bin shims that
a reinstall recreates, so they stay skipped rather than blocking every
JS plugin's update. The no-git branch is unchanged: there, links may be
upstream's own.

The existing ignored-data-dir git test gains the case: the update
refuses, names data/link.yaml, and the live plugin keeps its revision,
its link and its data. This also gives the no-follow rule a test that
fails if the link is followed.
2026-09-27 01:42:52 +05:30
kshitijk4poor
1362b95c04 fix(plugins): keep dashboard/ and JS module files revision-owned on no-git carry
A subdirectory install has no .git, so the carry cannot tell removed
upstream code from user files and relies on a deny-list. That list missed
the dashboard surface (web_server_dashboard loads dashboard/manifest.json)
and .mjs/.cjs/.jsx/.tsx, so an upstream that dropped its dashboard or a
hook script got the old files back.

Add dashboard/ to _NO_GIT_REVISION_DIRS and the four extensions as a
carry-local set. They stay out of tools.plugin_guard.CODE_FILE_EXTENSIONS
on purpose: that set exempts code files from env-secret scan patterns, so
widening it there would weaken scanning rather than broaden it.

The docs now list what is actually enforced, and the existing subdir
update test pins dashboard/manifest.json + hooks/run.cjs removed upstream
are not resurrected.
2026-09-27 01:42:52 +05:30
JoaoMarcos44
d05b33cb4d fix(plugins): contain user-file carry within staged tree
(cherry picked from commit 97be7ac043a95f20268628e1a3c2e6d22c9c74d1)
2026-09-27 01:42:52 +05:30
JoaoMarcos44
23e678aaf9 test(plugins): cover destructive update paths
(cherry picked from commit 7d5cfb5cff73661fc96a0737c1b187fc4594c8d6)
2026-09-27 01:42:52 +05:30
kshitijk4poor
1b00d0ad7b docs(gateway): state the real bound on concurrent turn threads
The comment claimed turn concurrency is bounded at admission by
max_concurrent_sessions, but that defaults to unset and auto-resume is
uncapped. The real bound is one live turn per session plus turns
abandoned by the inactivity timeout, whose threads keep running. Say so,
and note in the max_concurrent_sessions docs that it is the only cap on
concurrent gateway turns.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
2026-09-27 01:05:06 +05:30
IntrepidBytes
8682d5791b fix(gateway): stop launchd osascript idle CPU spin
(cherry picked from commit ef6f22810341601bb8de58c9fbba5423361fbe23)
2026-09-27 01:04:39 +05:30
John Paul Soliva
7d49b46e15 fix(sessions): prune keeps the compressed-away start of a chat still in use
Retention prune (the default-on startup auto-prune, `hermes sessions prune`
and the dashboard prune) aged every session row on its own. A conversation
that rotating compression split into segments has an ended, old root by
construction, so once that root passed retention_days it was deleted while
the conversation's live tip was still being written: the pre-compression
turns vanished from the resume/Desktop history and from session_search, and
the tip was orphaned.

Prune now deletes a compression ancestor only together with every
continuation after it (`whole_lineages`), so a lineage ages through its
newest segment and goes as a unit once the whole conversation qualifies.
Branch, delegate, reset and tool children do not count as continuations.
The CLI and dashboard prune previews pass the same flag, so they list what
prune deletes; bulk export keeps its current selection.

(cherry picked from commit b082fc6ffa02607cba219a5dfa361d4f3bf4c266)
2026-09-27 00:43:39 +05:30
kshitijk4poor
3b5945499e docs(sessions): a resumed memory-only row becomes repairable; the prompt is not healed 2026-09-27 00:40:21 +05:30
kshitijk4poor
e4154b88dd docs(sessions): memory-only repair-prompts rows self-heal on resume
They are not "never auto-repaired": resuming such a session re-pins the
full tool surface (restore_agent_tool_prefix), after which a scan sees
skill_manage without the Skill Safety guidance and clears the prompt.
2026-09-27 00:40:21 +05:30
kshitijk4poor
0c1a1036fc refactor(sessions): drop the repair-prompts pin clear; single-pass scan
A memory-only tools[] pin self-heals: restore_agent_tool_prefix appends every
fresh tool to the pin and persists it (merged != pinned) on the next turn. The
clear_pin flag was also unreachable in scan mode (findings need skill_manage).
So drop clear_pin/_HYGIENE_PIN_TOOLS and clear_system_prompt_for_rebuild, and
reuse update_system_prompt(sid, None), which already nulls prompt+hash and GCs
in one write.

The detector now parses each pin and checks the marker once per row, keyed on
prompt_builder.SKILL_SAFETY_HEADING instead of a hand-copied literal. The scan
classifies each compact_rows page as it arrives, keeps only finding dicts, and
dedupes ids that OFFSET paging can re-serve during concurrent inserts.
2026-09-27 00:40:21 +05:30
kshitijk4poor
8918e8a0fa fix(sessions): only auto-repair prompts with skill_manage evidence
The repair-prompts detector cleared two legitimate prompts:

- a pin with skills_list/skill_view but no skill_manage and zero skills
  installed: build_skills_system_prompt returns '' and SKILLS_GUIDANCE is
  only emitted with skill_manage, so the healthy prompt has neither marker;
- an exact memory-only pin, which is also a user toolsets=[memory] config;
  the healthy rebuild was re-flagged on every run (not idempotent).

Key the decision on the missing '## Skill Safety' guidance, which is
unconditional when skill_manage is in the pin, and require skill_manage in
the pin. Memory-only rows are now reported as unverifiable; the explicit
SESSION_ID override still clears them and their pin.

Docs: describe the tightened rule and note that a running gateway keeps
cached prompts in memory, so it must be restarted after --apply.
2026-09-27 00:40:21 +05:30
JoaoMarcos44
b09ed52c84 fix(sessions): make prompt repair atomic
(cherry picked from commit 55bcd8f6ccd667a5352769ad82bcbad1e73f82b1)
2026-09-27 00:40:21 +05:30
JoaoMarcos44
7b3dd14768 fix(sessions): make prompt repair evidence-safe
(cherry picked from commit b1eaaeabbeeb3500de390b77de62763bf33c86ae)
2026-09-27 00:40:21 +05:30
teknium1
63e44332f5 fix(kanban): a worker the dispatcher never recorded registers itself instead of being run twice
A dispatcher SIGKILLed between _call_spawn_fn and _set_worker_pid leaves a
live worker on a run with worker_pid NULL. release_stale_claims only extends
an expired claim for a recorded live pid, so on TTL expiry it reclaimed the
card and spawned a second worker beside the first: double billing, double
side effects, and a board showing one clean completed run (the first
worker's kanban_complete is refused as stale). Main CI hit it in
test_dispatcher_sigkill_mid_tick_never_destroys_or_duplicates_cards.

The worker now records its own pid on its run before the first model call
(adopt_worker_pid, worker_registered event, host-local claims only) and
exits without working the card when its run was already reclaimed. The
reclaim UPDATE also compares worker_pid so a registration landing between
the stale-claim SELECT and the UPDATE keeps the claim.

Repro: temporary sleep between spawn and pid record + kill 0.2 s after the
spawned event + slow first model reply -> 4/4 red on main with the CI
signature, 8/8 green here.

Fixes #121556
2026-09-26 11:16:45 -07:00
kshitijk4poor
f077152871 docs(cron): say the counter is reset to a valid count, not set to one 2026-09-26 23:00:54 +05:30
kshitijk4poor
7ca5cca50a fix(cron): an Infinity or negative repeat.completed no longer breaks load_jobs
06a495cc5b normalized non-int counters but caught only TypeError/ValueError.
json.loads turns a hand-edited Infinity / -Infinity / 1e999 into float inf,
int(inf) raises OverflowError, and that escaped load_jobs, so every job
(list, tick, mark_job_run, hermes cron list) failed, not just the bad one.
Catch OverflowError (-> 0), and also clamp a negative int count, which
granted extra runs. The docs sentence now says non-negative.
2026-09-26 23:00:54 +05:30
kshitijk4poor
66c5099b31 docs(sessions): note same-id recreate when deleting a live session
A session deleted while its chat is still running is recreated under the
same id with the full in-memory transcript on the next save (the gateway
session-key mapping expects the id to be stable). Say so next to
`hermes sessions delete`, in English and the zh-Hans mirror.

Fixes #123583

Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
2026-09-26 22:39:42 +05:30
kshitijk4poor
06a495cc5b fix(cron): normalize any non-int repeat.completed, not only null
A hand-edited "completed": "2" still crashed every recorded run ("2" + 1),
and 1.0 was stored as 2.0 ("2.0/3"). load_jobs now coerces any non-int
counter to a non-negative int (0 when unparseable). Document the load-time
repair next to the direct-edit tip.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:04:17 +05:30