The Nix desktop selected a legacy user venv before its explicit backend command. That runtime passed the version probe but failed agent initialization with a missing tools.kanban_toolset_context module.
Resolve the deployment command before the managed-install fallback. Add a headless packaged-desktop check with a competing legacy launcher; it fails on the old artifact and passes inside the Nix sandbox after the fix.
Keep desktop rollback in staged publication instead of restoring backups
inside a candidate directory that will be discarded.
Use the npm tag endpoint, let the downloader own archive verification,
and ensure CI toolchain roots without repeating dependency verification.
Remove an unused resolver input and replace overlapping tests with
fault-proven lifecycle coverage.
Verified with the scoped Python gate and before-pack tests. Native
Windows/macOS update execution and the full suite were not run.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):
- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
no duration are outside the unified video_generate surface and are dropped)
with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
live limits (nearest by value/height/ratio) and drops generate_audio/seed
for models that lack them (the API 400s otherwise); reference images ride
in input_references; local file inputs are refused (OpenRouter fetches
URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
never to a provider-supplied unsigned_urls host (kept from #103267)
Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.
Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
Run historical updater completion in a fresh interpreter so cached imports
cannot revive retired dependency installers. Share Git and ZIP completion,
carry receipt and recovery state, and preserve child exit status.
Route plugin admission, binary acquisition, desktop launch and build paths
through PM. Replace redundant helpers and tests with real worker, package,
publication and launch checks. Keep the shipped compatibility surface fixed.
Targeted Python and desktop checks pass. Native update journeys and fresh
production image qualification remain pending. This is a checkpoint before
those acceptance runs.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
Salvage follow-up to Xipong's #107736. Kept the core: `kanban` is a
configurable, default-off toolset whose check_fn answers the schema
build's own selection (ContextVar) instead of the legacy top-level
`toolsets` key, so `platform_toolsets.<platform>: [.., kanban]` — what
`hermes tools enable kanban --platform X` writes — actually reaches the
gateway agent's tool schema.
Dropped the `tui_gateway/server.py` change: turning an explicitly empty
CLI selection from "all" into "nothing" is a separate behaviour flip
already tracked by #107452, not part of this bug. The two TUI loader
tests that asserted `kanban` is auto-recovered onto a saved `[memory]`
list now assert the opposite: a configurable opt-in is never recovered.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
The warning fires once per process but was prefixed with the first
job's unit suffix, reading as a per-job notice for a host-level
condition. Drop the suffix; the message now says what applies to every
later dispatch.
Follow-up to the cherry-picked #102431 fix, addressing the review findings:
- The two real-helper scheduler tests ran the Linux-only helper unmarked
and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
scripts/check-windows-footguns.py --all (lint lane red). The remedy text
is now built once in the helper and passed into the warning, so the
"scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
once (helper level); default config still Popens externally and
`require_restart_safe_scope: true` raises (scheduler level, real helper).
Dropped the stubbed duplicate, the standalone config-raise test and the
in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
with the same `except Exception` guard as the sibling
`failure_nudge_threshold` read, so a config error no longer escapes
the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
`cron.require_restart_safe_scope` documented in the cron user guide.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).
Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.
The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.
Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.
Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
PM owns Python dependency generations. Shared frontend builders own Node
preparation and compilation. Route source updates and launchers through
these owners instead of separate repair ladders.
Remove obsolete live-venv holder gates and soft build-failure plumbing.
Preserve source validation, staged publication, fleet outcomes, and
historical relaunch hooks.
Verification: 956 tests passed in the combined focused run, with 43 skips.
After the final ZIP exit fix, 583 focused tests passed. Shared JavaScript
builder tests, Ruff, and diff checks passed. Native Windows/macOS and full
packaged-app builds were not run.
Replace the POSIX-only jobs-flock contention test (skipped off-POSIX,
~120 LOC of monkeypatched flock plumbing) with a single invariant test
that fails on pre-fix code in <1s: hold the per-job fire fence from a
worker thread, assert the heartbeat still returns True on the calling
thread, and that a takeover is still detected (False). The docstring on
heartbeat_fire_claim now records WHY it is not under the fence, so the
next refactor does not put it back.
Co-authored-by: Oliver Heckmann <46627487+oheckmann74@users.noreply.github.com>
Co-authored-by: salch-cred <141555468+salch-cred@users.noreply.github.com>
heartbeat_fire_claim only CAS-refreshes claim.at via _with_job; wrapping
_under_fire_fence across save_jobs let a blocked .jobs.lock pin the fence
and cause mark_job_run to fail closed on completed jobs.
Give native payload dependency builds two hours without changing the normal install timeout. Allow three hours for the standalone PM bundle job so setup, cache saves, and smoke tests fit around the build.
Verified seven focused tests, Ruff, and actionlint. Full CI builds were not rerun.
Resolve the default cache before isolating HOME so payload builds use the directory that CI restores and saves. Remove the standalone bundle workflow dependency on an unset cache variable.
Verified offline wheel reuse in isolated children, 20 focused tests, Ruff, and actionlint. Full native release builds and the separate source-build timeout remain unverified.
Replace the two stubbed tests with a single test that drives the real
CodexTransport.build_kwargs (the fixture agent already carries a tool), so
the test binds the summary body actually sent, not a hand-written dict. It
covers both the first attempt and the empty-summary retry; the previous
pair asserted the same three keys twice and reused one dict across
attempts, so the second-attempt check passed trivially after the first pop.
Add the WHY comment on the pops: the transport emits tools, tool_choice and
parallel_tool_calls as one block, and strict Responses backends 400 on the
controls without tools.
Iteration-limit summaries retry once when the first attempt comes back empty.
Both attempts share one builder closure, so assert the retry body is as
tool-free as the first: forces an empty summary, then checks every captured
codex request for tools / tool_choice / parallel_tool_calls.
Addresses the retry-coverage request on #32777.
Application errors (isError payloads) keep counting as breaker strikes:
that is #10447's point (a server answering errors made the model hammer
it 8x in 10s) and #109180 just reasserted it. What #11113 actually hit is
the open-breaker MESSAGE: after three rejected fetches the model was told
the server was "unreachable" and went to the user instead of fixing its
URL. Track whether the streak was all application errors and word the
pause accordingly; one transport strike restores the unreachable text.
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
them as `stopped`.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
The profiles guide promised "fresh sessions and memory" while
_CLONE_SUBDIR_FILES deliberately copies memories/MEMORY.md and USER.md
(curated identity, same tier as SOUL.md). State the real behaviour and how
to get a blank memory, so #10376's first half stops surprising people.
Refs #10376
Drop the categorical 'missing any step always means 200340' and
'200672/673 necessarily is an adapter bug' claims; the Interactive Card
toggle's role is unverified. Bring the Chinese guide in line (it still
pointed at the Event tab). Review follow-up.
Six issues (#10251, #10073, #13924, #38305, #25886, #8246) and four PRs
proposed json.dumps()-ing CallBackCard.data. The lark_oapi SDK types
`data` as Any and marshals the whole response with JSON.marshal(), so a
dict already reaches Feishu as the nested object the card-callback doc
shows in its response example; a string would nest a JSON *string* there
instead. Feishu documents 200340 as "the application has not configured
the card callback address or the configured request address is invalid"
— the click is rejected before delivery (nothing reaches Hermes), which
matches the reports (no gateway log, all four buttons identical).
The docs told users to subscribe to `card.action.trigger` under Event
Subscriptions; it lives under the separate Callback Configuration tab and
needs its own long-connection/URL mode and a published app version.
Spell out the four steps and map the sibling codes (200342/200343 =
unreachable URL; 200672/200673 = our payload).
Refs #10251
The check called the registry check_fn inside the updater's own process,
whose import caches predate the install just performed (and which may be
the outer Python entirely), so a healthy freshly installed SDK produced a
false "will fail to load" warning. Run the same registry check in the
target interpreter via the existing _venv_probe path used by the core
dependency verifier.
Found by independent review before merge.
When `.[all]` fails and the per-extra fallback also fails for e.g.
`feishu`, the update printed only "Skipped optional extras that still
failed" and finished green. The running gateway kept its already-imported
modules, so the loss surfaced hours later as "No adapter available for
feishu" on the next restart (#10651).
After the fallback, check every enabled+configured platform through its
registry `check_fn` (and MCP when `mcp_servers` is set) and print which
configured feature will fail to load, with its install hint. Unconfigured
extras stay a quiet skipped line.
Reworks PR #10733 (LeonSGP43) against the registry instead of a
hand-written platform->module table so plugin platforms are covered.
Fixes#10651
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Two regressions in the mirror support: (1) the canonical
https://openrouter.ai/api/v1 that `hermes setup` persists under
provider: openrouter was treated as a custom endpoint, dropping the
auth.json credential pool and returning an empty API key; (2) an
unrelated CUSTOM_BASE_URL (which outranks the config mirror) still
received OPENROUTER_API_KEY because key selection tested mirror
eligibility, not the endpoint actually selected. A config URL is a
mirror only when its host is not openrouter.ai, and the mirror key
branch fires only when base_url is the config URL.
Found by independent review before merge.
When config.yaml sets `model.provider: openrouter` together with a
`model.base_url` mirror/proxy, an explicit `--provider openrouter`
request ignored the mirror and sent traffic to the public OpenRouter
endpoint: the config base_url was only trusted for auto/custom, the
credential pool was still consulted (so a pooled key won over the
mirror), and even when the mirror URL was used its host failed the
openrouter.ai match so OPENROUTER_API_KEY was not selected for it.
Trust the config base_url for the explicit openrouter case, treat that
mirror as an OpenRouter context for key selection, and bypass the pool
like the other custom-endpoint cases already do.
Fixes#10622
send_document() accepted reply_to but never passed it down, so attachments
always threaded from the cached per-address context (or not at all) even
when the caller named the message to reply to. The plain-text path
(_send_email) already honored it; the attachment path now does too.
The metadata half of #10131 (send_image rejecting metadata=) was already
fixed on main by the adapter parity pass. Diagnosis from #10131 and the
explicit-reply_to-wins shape from PR #10321 (which targeted the
pre-plugin path).
Fixes#10131
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>