The serve a remote Desktop spawns over SSH is never restarted by the
host's updater (only its client holds the token and owner nonce), and
the idle watchdog stays quiet while that client is connected. After an
update it therefore kept running the pre-update code against the new
tree until the client happened to reconnect.
Poll the checkout sha against the sha this process loaded. On two
consecutive mismatches, with no update in flight, retire through the
existing retirement fence (which proves idle and closes admission), and
exit cleanly so the client respawns the backend on the new code. Unknown
shas, busy or unreadable ledgers and a live update all keep it up.
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.
Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
hermes serve --isolated (the backend another machine's Desktop spawns
over SSH) opts out of the host singleton on the CLI side, but its spawn
ledger row carried no structured marker, so the local Desktop's
attach-first discovery adopted it. Nothing on this host owns that
process, so a Desktop-driven update left the local app on stale code.
Record isolated=True in the ledger row and skip such rows in
parseSpawnLedger. An ordinary serve is still attached even when an
isolated record is newer.
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
The settled-flag fix covers a process that multiplexes itself. The update
and fleet processes replay a FOREIGN gateway's captured argv with no
settled flag of their own, so a selector-less argv fell back to the
ambient HERMES_HOME comparison — the exact coordinate the review rejects
(#93943): a host launched from a named profile was replayed as that
profile, donating the named credentials to the respawned host.
Both restart edges now consult, in order: this process's settled
multiplex verdict, then the live host gateway's published rendezvous
record (its SETTLED served set, proven live), and only then the
compatibility default-root comparison. The raw config re-read stays last
so no settled identity exists => unchanged compatibility behavior.
Regressions: selector-less replay from a named home with a live host
record is host; without one it stays profile-scoped; the restart watcher
env takes the default root and drops the named token when only the host
record proves hostness.
A profile-scoped parent donated its environ to the host multiplexer, so a
named launcher was treated as the primary adapter owner and its platform
token became the primary claim. Spawn the host with served_profile_child_env
for the default root, mark multiplex active before that primary load, and
name the env-derived side in a duplicate-credential refusal.
Addresses two P1 review findings on #121614.
1. `provider_model_ids`: a failed or empty relay probe fell through to the
canonical per-provider fetcher, sending the provider credential to exactly
the vendor host the user routed away from — recreating #121387 on the
failure path. A configured `model.base_url` relay is now TERMINAL for live
catalog egress and degrades to the local curated list instead. The curated
tail is extracted as `_static_catalog` and shared by both paths.
Fetchers that already resolve `model.base_url` themselves and degrade
locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
simple api-key fetchers) are excluded from interception via
`_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
produce a better-merged catalog.
2. `_try_anthropic`: an `explicit_base_url` that failed
`_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
the ambient/canonical host and continuing with the explicit credential —
a silent retarget of authority, not the refusal the PR body claimed. It now
returns unavailable before client construction.
Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).
Refs #121486
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
directly when the pool yields no token (no second uncached pool load,
no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
helper's error fallback reads the profile-scoped override, not the raw
process env.
- model setup flow: the confirm guards get the resolved Codex base, not
the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):
- the OAuth context-length probe (agent/model_metadata.py) and the
/model picker's live discovery (hermes_cli/codex_models.py) both GET
https://chatgpt.com/backend-api/codex/models with
Authorization: Bearer <gateway key> whenever model.context_length is
not pinned — the key is sent to a service it does not belong to and
cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.
Fix, mirroring the quota probe's existing gate in auth_codex:
- both catalog sites now decline to probe non-JWT credentials (real
Codex access tokens are JWTs; a gateway key is not one) and fall
back to the static table / offline sources — same outcome as the
doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
/models instead of chatgpt.com (catalog URLs are built from the
resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
same way the text client does.
Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.
(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
Follow-up to the ported picker fix (#121623):
- Ticking a row re-enables it even when plugins.disabled holds the manifest
name (e.g. telegram-platform, written by pickers before #40190). The save
only dropped the key and its bare leaf, so the gate still matched the
manifest name and the plugin stayed off. It now purges every alias like
`hermes plugins enable/disable` (_apply_activation, shared with
_set_plugin_enabled), with one discovery scan per save.
- The success line counts the rows actually turned on/off instead of treating
every unticked row as "disabled"; the unused new_enabled return is gone.
- _entry_status shares the per-row status call between the picker
preselection and `hermes plugins list --enabled`.
The bare `hermes plugins` picker preselected only rows listed in
plugins.enabled. Bundled platforms, backends and model providers are active
without a list entry, so they opened unticked, and the save on exit wrote every
unticked row into plugins.disabled: opening the picker and leaving without a
change disabled every messaging adapter on the next gateway restart.
Rows now open ticked by the load-time rule (_plugin_status), and only rows the
user flipped are written.
Ported onto the plugins_cmd_toggle sibling and the admission-authority save
(the original targeted the pre-split plugins_cmd.py). The nested
canonical-key composite test now ticks its row explicitly: it handed the menu
a checkbox state that contradicted the config and relied on the old
rebuild-everything save. The original helper-level tests are rewritten
against real admission in a follow-up commit.
(cherry picked from commit 3a0d5f1b8bce3b23b115ba4995b8c23dfbba9ad0)
An install's in-tree venv whose editable record names another checkout
(what project_venv_dir used to cause) turns <install>/venv/bin/hermes
into that checkout's CLI. Every update Desktop hands to the launcher
then pulls the other tree: the install never moves, the hand-off
reports success, and posix.sh swaps the install's stale release/ build
back over the app. Pressing Update can never get out of that loop.
`hermes update` now notices it is running on another checkout's in-tree
venv and re-runs itself with that install's own code (PYTHONPATH=<install>,
cwd=<install>). The install's updater pulls the install and reinstalls it
into its venv, which points the editable record home again, so the next
relaunch runs a fresh build. A wrapper that put the running checkout on
PYTHONPATH on purpose chose that tree and is left alone.
#122234 gave Desktop update steps NUL stdin, but the hand-off script
that runs is the one from the checkout being updated FROM. Every update
that starts on an older commit still runs the old script, which gives
steps the hand-off console as stdin and captures their stdout until
they exit. Only the steps after `hermes update` run new code.
So `gateway start --all` from the new checkout still saw an interactive
console, asked "Install it now so the gateway starts on login?" into
the captured stdout, and waited forever. The update never relaunched.
start() now asks only when stdout is a terminal too. Nobody can answer
a question they cannot see.
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.
Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
The setup stage installs the gateway service through
ensure_gateway_service. On Windows that asks the start-now, Scheduled
Task and UAC questions. The gateway stage then ran `hermes gateway
install`, which asked them all again.
`gateway install --if-missing` does nothing when a service is already
installed. Both installers' gateway stages use it, so they ask only
when setup did not install the service.
Under the updater's claim, sync whenever dependencies are not current
rather than only when nothing is committed. The tail's own children
and post-sync verification children are current and stay no-ops. A
stale generation still committed from the previous Python pin is the
same ABI trap as the pre-PM venv, and it now syncs too. If the sync
still leaves the tree out of date, raise instead of relaunching into
another sync.
prepare_launch returned early for any process running under the
updater's own claim, so it would not re-run the completion tail. That
also covered processes the updater spawns before PM commits a
generation (a restarted gateway), which then booted with no
environment: previously on the pre-PM venv, now refused.
Under the updater's claim with nothing committed, sync the dependency
generation (carrying the legacy venv's extras, as the first sync
always has), skip the tail since that belongs to the updater, and
relaunch on the store Python. The relaunched process sees the commit
and returns early as before, so the no-recursion guard still holds.
Desktop speak-stream sent type=fallback whenever the provider had no
chunked PCM API. Edge is that case, so the client waited for the full
reply and POSTed it. Cut sentences with the existing sync TTS tool and
stream that PCM. Fallback stays the last resort when synthesis produces
no audio.
Refs #91997
User installs failed with 'resvg-py is missing' because the web/desktop
source builds rendered icons on whatever python was on PATH. The default
brand outputs are now committed; source_build, apps/desktop build.mjs and
the npm/docusaurus pre-hooks consume them directly. Flavored release
bundles (canary/commit) still render into their own product dir.
icons-freshness-check now regenerates and fails on any byte diff.
prepare_launch finishes an interrupted update, or a hand-run git pull, by
syncing the venv at startup. It skipped the required-tool step that
`hermes update` now runs first, so a bumped ripgrep/ffmpeg/python pin stayed
uninstalled and activation warned on every start.
A normal `hermes update` only re-synced the venv. The sync pulls uv/python
in through its own dependency, but a ripgrep/ffmpeg/node/npm pin bump in
pm/lock.json was never installed, so PATH activation warned and skipped the
managed tool dirs on every CLI start and gateway boot. Only the takeover
route for historical releases ensured the tool roots.
Move the takeover's loop into pm.client.ensure_tools_for_sync() and call it
from both routes before the sync: update_completion._prepare (the CLI and
Desktop route, running from the new tree so the new lockfile applies) and
_update_takeover.prepare. It uses explicit=True like the takeover (an update
is an explicit user action) and a failed download fails the update.
post_update.step_provision_runtimes / MACHINE_STEPS stay: `python -m
hermes_cli.post_update --scope machine` and tests still reference them.
The ffmpeg docstring no longer claims that step re-ensures it.
A source update re-syncs only the venv, so installs created before
agent-browser became a default would never get it. Skipped for sealed
payloads and when lazy installs are disabled; honours the recorded
opt-out; a failed download warns and never fails the update.
A v2026.7.1 Linux Desktop runs `hermes update` as a piped child. The swap
already spares that Desktop's main process, because it is our ancestor. But its
zygote, renderer, GPU and network-service helpers run the same release exe and
are not our ancestors, so the swap still stopped them. E2E agent.log shows six
PIDs stopped before the staged app promotion. At that moment the renderer's
backend WebSocket dropped with 1006, and desktop.log and the screen recording
froze. The update receipt then succeeded, but the main process had no renderer.
It never relaunched and never quit, so launch-from-spec hung in the post-update
phase until the 20 minute self-timeout (exit 124).
Spare the whole process tree of a Desktop ancestor that runs from the release
tree. Only a Desktop ancestor's descendants are spared: every process descends
from init. Other Desktops running from the tree are still stopped (#109643).
Windows behavior is unchanged.
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.
The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.
Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
A v2026.7.1 Linux Desktop runs `hermes update` as a child with piped
stdout/stderr and relaunches itself afterwards. The build swap stopped every
process running from the release tree, including that ancestor, so the
update's next write failed with EPIPE and nobody relaunched the app.
Spare our own ancestors on POSIX, where the rename works with the app alive.
Windows keeps stopping them because the exe lock would fail the rename.
Importing hermes_cli on a host whose stdout is not UTF-8 (a cp1252 pipe on
Windows, a latin-1 locale on a Pi) set PYTHONUTF8/PYTHONIOENCODING in
os.environ as a side effect. Every library importer -- gateway.relay, the
compute host, agent.auxiliary_client -- then leaked that into each child it
spawned, which is what
test_library_imports_of_dual_use_entry_modules_stay_side_effect_free
catches on Windows (changed == [PYTHONIOENCODING, PYTHONUTF8]).
The package import now only repairs its own streams and records whether it
had to. hermes_cli.main.main() turns that record into the child-process
hint, so `hermes` still steers its Python children to UTF-8 on such a host.
On Windows the entry-point bootstrap and configure_windows_stdio() already
export both variables; the tool subprocess env builders set them for
children independently.
The test was unmarked, so the OS lanes (`-m "platforms and not
integration"`) deselected it and Windows never ran it. Mark it
platforms("any"): list_os_marked_tests.py lists the file for every lane and
the -m expression now selects the test.
About showed a bare "v0.21.5" for a checkout 1913 commits past the tag
(`hermes --version`: "v0.21.5+1913.gf83a9e9"), because /api/health only
reports the base release; the statusbar showed the version plus the exact
commit beside it.
VersionInfo.display_version is the short `<release>+<distance>` form and
/api/health carries it as displayVersion (older backends fall back to the
bare release). One renderer formatter, shortVersion, turns any version into
that form, and About, the updates overlay, the statusbar and the command
palette all label through it: "v0.21.5+1913". The full commit stays in the
statusbar tooltip and the expanded version details.
A source checkout on the main channel failed every update check with
"Could not resolve the main source channel: Channel object not found:
releases/channels/main.json" until R2 publishes that record, in the CLI
and in Desktop (which asks hermes_cli.source_check). main IS the source
branch -- its record can only add a retirement -- so a missing main record
now resolves to the main branch and the update continues via git. Other
channels, and transient read failures for main, still refuse.
Main-era installs selected [all] and then lazily installed opt-in extras
(FAL, messaging SDKs, ...) into the checkout's venv with no ledger. The
legacy -> PM migration (launch completion and the historical updater
takeover) selected only [all], so the first launch after `hermes update`
prompted "Extra 'fal' requires: fal_client" for a feature that already
worked.
pm.extras.legacy_selection reads the old venv's site-packages (never
imports it) and selects every non-umbrella, platform-supported extra whose
anchors it held. The install prompt that remains for genuinely new
features names the feature instead of Python module names.
uv, npm ci, icon generation and the desktop build printed hundreds of
lines on every interactive install, update and `hermes desktop`.
pm/progress.py holds one policy. CI (CI / GITHUB_ACTIONS) or
HERMES_VERBOSE=1 streams child output unchanged. Otherwise each child is
one "→ label…" line, rewritten in place with its latest output on a
terminal, or just the start and finish lines when piped. A failure
always prints the last 80 lines. HERMES_VERBOSE=0 forces containment.
Applied to PM's uv runs (no uv --verbose outside CI; quick steps stay
silent unless they fail), the source build scripts (node-deps, TUI,
icons, web; npm warn lines stay out of the live line), the desktop
build and package steps, and the Windows ARM64 vcpkg/OpenSSL provider.
The policy reads the parent's environment, so the CI=1 the builders
force on their node children does not switch it to verbose.
The fleet-restart-pending discharge parsed the marker inventory and
probed a gateway-less host inline (CC 36). _marker_owed_gateways now
turns the inventory into the owed set (raising ValueError, which the
caller's existing handler already maps to "keep the marker"), and
_discharge_gatewayless_marker owns the #118742 host probe. Every
verdict is unchanged; the caller is at CC 21.
_cmd_update_check (CC 30, 224 lines) did debris cleanup, channel
resolution, the scoped upstream/origin fetch, shallow-graft repair and
two verdict printers inline. Those steps move to the topical sibling
hermes_cli/update_cmd_check.py; the facade function (a frozen updater
surface) stays in update_cmd and orchestrates them at CC 10.
Output, exit codes and git argv are unchanged. The sibling reads
_capture_head_sha / _no_prompt_git_kwargs through the facade and
imports gitlock, source_releases, source_check and config per call, so
existing monkeypatch seams keep reaching production.
recover_if_needed (CC 33) and _restore_holding_claim (CC 30) carried
their failure bookkeeping and git put-back steps inline. Extract
_missing_environment, _count_failed_attempt and _put_back_paths so
both land at CC 24-25 with the same ordering: a failed restore still
skips the rm, and the attempt bump stays best-effort per marker.
The sealed-tree steward ladder in evaluate_update_admission was four
if-rungs that differed only in the install method whose command
remediates the tree. A steward -> method table plus one helper replaces
them; the refusal code, message and command are unchanged for docker,
nix, apt-termux, desktop-app and unknown future stewards.
Only a manual `hermes pm gc` reclaimed old app and PM runtime generations
on native installs; the Docker boot already collects after its refresh.
A successful venv sync (hermes update and launch-time completion) now runs
the same collectors, which keep selected, leased and day-young generations
and skip when another install holds the lock. Cleanup errors are logged and
never fail the committed update.
sync_venv took four mutually exclusive plugin kwargs (plugin_dirs,
extra_plugin_dirs, selection, staged_plugin) and venv_is_current two, with
runtime ValueErrors guarding the combinations. Replace them with a single
`plugins=` argument typed as one of pm.plugin_inputs.Members, Candidates,
Selection or StagedUpdate, so a conflicting request cannot be expressed.
The module also owns the worker wire encoding that client.py and worker.py
each duplicated.
pm.install.sync_venv is split into cohesive helpers (feature policy,
install lock, publication snapshot, target selection, commit) with the
same ordering, receipts and recovery; its CC drops from 54 to 16.
All in-tree callers and tests move to the new argument.