`hermes config edit` and `hermes doctor --fix` each carried their own
"copy cli-config.yaml.example, else save DEFAULT_CONFIG" block, and they had
already drifted: doctor's template copy skipped _secure_file, so the seeded
config.yaml kept the checkout's mode instead of 0600. Seeder drift is the
bug class behind #121230 (one seeder wrote DEFAULT_CONFIG verbatim and pinned
global display values over every messaging platform's defaults).
Both now call hermes_cli.config.seed_config_file(config_path, template=None),
which returns whether the template was used (doctor keeps its message).
Doctor still passes its own PROJECT_ROOT template so its tests' patching
keeps working. Drops doctor_config's now-unused shutil import.
The curl installer, the Windows installer, the Docker first boot and
`hermes doctor --fix` copy cli-config.yaml.example into config.yaml byte
for byte. The template had five display keys uncommented: tool_progress,
interim_assistant_messages, long_running_notifications, busy_ack_detail
and show_reasoning. The gateway reads config.yaml without a DEFAULT_CONFIG
merge, and resolve_display_setting takes a global display.<key> ahead of
_PLATFORM_DEFAULTS. So every seeded home ran with those values on every
platform. Telegram and Slack posted every tool call. Signal, email, SMS
and the other no-edit platforms got progress lines, heartbeats and
interim messages. Every messaging reply had the reasoning block prepended.
First-time `hermes setup` (quick and full) and Blank Slate setup also
wrote display.tool_progress: "all". That write was added as a Quick
Install recommended default (79aeaa97e6) nine days before the
per-platform tiers landed (#8006), and it has the same effect for
tool_progress on homes the template never touched.
`hermes config edit` on a home with no config.yaml wrote DEFAULT_CONFIG
unstripped, which pins show_reasoning (all 21 platforms),
interim_assistant_messages (12) and tool_preview_length (16). It now
seeds like the installer and `doctor --fix`: the template when the
checkout has one (a full file to edit, written owner-only), otherwise
DEFAULT_CONFIG with defaults stripped.
Measured through the real gateway loader across the 21 platforms in
_PLATFORM_DEFAULTS, a template-seeded home differed from a bare one on
tool_progress for 19 platforms, show_reasoning for 21, busy_ack_detail
for 14, long_running_notifications for 13 and interim_assistant_messages
for 12. With the pins commented out and the setup writes removed, the
diff is empty, and the same holds for both `config edit` seeds.
The CLI does not depend on these values. It defaults tool_progress to
"all" and show_reasoning to true when the keys are absent, and the TUI
defaults interim_assistant_messages to true.
Homes that were already seeded keep their values. A template value
cannot be told apart from one the operator chose, so there is no
migration. The messaging docs now say which lines to delete.
(cherry picked from commit 96450d4500613ab1ba45c7e972f31de570bc2d71)
migrate_config() and the docker boot script each parsed config.yaml twice:
check_config_version() coerced a missing `_config_version` to 0 and threw the
"was it present" bit away, so has_version_stamp() re-read the file to recover
it, guarded only by a call-order promise in its docstring. That promise did not
hold for the docker script, which used the tolerant check: a list-rooted
config.yaml read as "unversioned, not below the floor", ran the backup +
migrate_config() dance and exited 1 (base: floor warning, exit 0).
Factor the read into _read_config_version_stamp() -> (Optional[int], latest);
None means the mapping has no stamp. check_config_version() is a thin wrapper
(None -> 0) so its 10 callers see identical output. migrate_config() and the
docker script decide `unversioned` from that single read; the docker script
now does the strict read itself and leaves an unparseable or non-mapping file
alone with a warning and exit 0, matching its invalid-YAML posture.
has_version_stamp() is deleted. Docker tests that mocked the pre-check now
mock the new helper.
A config.yaml without _config_version reads as v0 and is exempt from the
support floor, so the first `hermes update`, profile clone,
`hermes doctor --fix` or docker boot ran every one-time migration step on
it. Installers seed config.yaml from cli-config.yaml.example, which had no
version, and targeted writers (`hermes config set`, /personality, the
TUI/Desktop config writers) never stamp one, so this is the normal state
of --skip-setup, non-TTY and Desktop (--non-interactive) installs. The
value- and absence-based steps then reset the personality, raised the
delegation caps, turned verify_on_stop off, shortened the curator windows,
dropped model_catalog.ttl_hours and enabled plugins the user had installed
but never enabled.
- A config with no _config_version now gets only the steps keyed on a
legacy key or identifier (LEGACY_KEY_STEPS), then the stamp.
- cli-config.yaml.example carries _config_version, so every seeded
config (install.sh, install.ps1, docker/stage2-hook.sh, doctor --fix)
starts at the current schema.
- docker_config_migrate.py no longer refuses a version-less volume with
the "predates version 12" warning; like migrate_config() it migrates
and stamps it.
(cherry picked from commit 97ba11e07009f633662b0c7fa8701aa5b441bd22)
The published image had no Xvnc/Xfce because nothing set the Dockerfile's
HERMES_BOT_DESKTOP argument, and a hosted instance (unprivileged, no sudo,
sealed /opt/hermes) cannot install at run time. The image layer is the only
delivery path.
- docker.yml: variant axis [slim, desktop]. :latest / :main / :v* stay the
image they are today; :latest-desktop / :main-desktop / :v*-desktop carry
the packages plus Playwright's headed Chromium. Slim owns the build cache
scope; one manifest per variant so a desktop publish failure never skips
slim's :latest.
- Dockerfile / stage2-hook.sh: XDG_RUNTIME_DIR=/tmp/hermes-runtime seeded
0700 as hermes (containers have no logind; the $HOME/.cache fallback was
the shared /opt/data volume), refused when foreign-owned; deterministic
Chromium discovery exporting the headless shell for ordinary browsing.
- bot_desktop: memory gate reads the cgroup working set (usage minus
inactive_file) so it cannot tighten over uptime and refuse to restart a
screen idle-stop just stopped; installable() gives three distinct dead-end
messages instead of a sudo line nobody there can run; env_for_agent
replaces a headless-shell pin so agent and dock share one Chromium.
Squash of IAvecilla/hermes-agent:bot-desktop-cloud-image (#112381, 13
commits), which GitHub auto-closed when its base branch merged as #108914.
Review fixes from pefontana (cache scope, per-variant merge, red browser
test) are included.
Co-authored-by: pefontana <pefontana@users.noreply.github.com>
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
A read error (EMFILE/EIO/EACCES) was never cached so the next load would retry
the file, but that also re-parsed the last-known-good backup (or rebuilt the
defaults) on every load_config() while the file stayed unreadable (~170x a
cache hit in review; 12-32 ms vs 0.2 ms per load here). The fallback is now
cached like a parse-error fallback, and a hit on a read-error fallback first
re-reads the raw file bytes (no parse): the fallback is served only while that
read still fails, so a cleared EMFILE (same file signature) reads the real file
at once.
The parse-failure banner, the active-failure record, FailedConfigRead and the
write refusals are one topic; config.py had grown past 4,100 lines with this
PR. Pure move (no behaviour change): config.py drops to 3,999 lines, below
origin/main. Callers outside config.py now import from the new module.
One EMFILE/EIO on an intact file recorded the file's signature in _CONFIG_PARSE_FAILURES,
and since nothing edits the file the record never expired: get_active_config_parse_failure()
kept returning the errno text for the rest of the process, so provider auto-resolution
refused with `corrupt_config` on every turn while the file read fine again. The good file
was also copied away as a `.corrupt` backup and the banner told the user to fix a
formatting error.
A successful read now clears the record (the flag means "the current file failed the last
attempt", whatever the cause, so the refusal still holds while the file is unreadable), a
read error keeps its intact file out of the `corrupt` backups, and the warning says the
file could not be read instead of pointing at YAML.
Live: PR head — one injected EMFILE, retry loads `skin: mono`, yet
get_active_config_parse_failure() == '[Errno 24] Too many open files' and
_refuse_env_adoption_if_config_corrupt() raises AuthError(corrupt_config); with this
commit the flag is None after the retry and no `.corrupt.*` copy is written.
Every fail-open config reader turned a read error on an intact file into
a stand-in ({} from read_raw_config and the TUI's _load_cfg_raw, defaults
or a possibly stale last-known-good from load_config). Writers then
mutated that stand-in and saved it; the writer re-reads the file, finds it
fine, and merges by deletion, so one EMFILE/EIO during a TUI config.set,
a dashboard save, a migration step or any load_config()->save_config()
caller replaced the whole file with the stand-in plus one key.
- Fallbacks from a failed read are now FailedConfigRead (a dict subclass
carrying the error). Readers are unchanged; save_config() and
atomic_config_write() refuse to persist one, so the error reaches the
caller and the file stays byte-identical. The subclass rides the
load->mutate->save round trip, so every writer is covered without
touching each of them.
- A read error's fallback is no longer cached (nor recorded as the next
last-known-good), so the retry reads the file again.
- PUT /api/config merges into a new dict, so it reads strictly instead.
- Gateway /verbose and /footer wrote the effective view back (fail-open
to {} on error, ${VAR} values expanded); they now round-trip the raw
file strictly.
_normalize_max_turns_config set config["agent"] unconditionally, so any
save of a raw document without an agent section (Desktop partial PUT,
every migration persist) appended `agent: {}` to config.yaml. Only set it
when the section exists or legacy root max_turns is being moved into it.
The v20 support-floor golden pinned that phantom: its `agent: {}` only
survived because the next migration step re-added it after the strip pass
emptied the section.
Use a distinct stripped-node sentinel so valid null values survive configuration saves and migrations. Keep existing default comparisons and explicit-path preservation rules.
Scope-risk: narrow
Tested: 14 new cases pass; 7 fail on unmodified main. Broader macOS run: 1266 passed, 7 skipped across 61 files.
Not-tested: Full repository suite and native Linux/Windows execution.
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.
Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
The cache-hit predicate was written twice per impl in hermes_cli/config.py:
once lock-free and once verbatim inside `_CONFIG_LOCK`. Extract the pure
lookup+predicate into `_load_config_cache_hit(path_key, cache_sig)` (sig
compare + ${VAR} env-snapshot check) and `_raw_config_cache_hit(path_key,
cache_key)` and call each from both sides. The helpers only look up and
compare; deepcopy-vs-identity stays with the caller, and the fast path keeps
its own try/except around the helper while the locked path stays
undefensive, so exception behaviour on each side is unchanged.
Not folded: passing the fast-path sig into the locked path to save the
second `_load_config_cache_sig` stat. The post-lock re-stat is the
double-checked-locking freshness guarantee; a write landing between the
fast path and lock acquisition would otherwise publish new contents under a
stale sig. Also the fast path only computes the sig when an entry exists, so
the double stat only happens on a sig-mismatch miss (a real re-parse).
PROOF: tests/hermes_cli/test_config_lock_free_cache_hit.py 3 passed on the
refactor. Mutation A (`_load_config_cache_hit` ignores the sig compare) ->
test_hook_timeout_memo_never_pairs_a_new_sig_with_an_old_value red (1 failed,
2 passed). Mutation B (`_raw_config_cache_hit` always returns None) ->
test_cached_raw_read_completes_while_another_thread_holds_the_config_lock red
("blocked 30.0s behind a held _CONFIG_LOCK"). Both reverted. ruff clean,
check-windows-footguns clean, real import under PYTHONSAFEPATH=1 ok.
- plugins.py: the `isinstance(cached_value, float)` guard defended against
nothing — `_resolve_hook_callback_timeout_uncached` only ever returns a
float and the hit test already requires `sig is not None`. Type the cache
value as `float`. `_reset_hook_callback_timeout_cache` is kept because
tests/hermes_cli/test_config_lock_free_cache_hit.py calls it; docstring
now says test-only (no production caller).
- config.py: the 16-line narrative comment in `_load_config_impl` restated
the raw-config comment and said 0.024us where the test says 0.024ms; cut
to 4 lines with the right unit and a pointer to `_read_raw_config_impl`.
- NOT folded: extracting the duplicated cache-hit check into a helper. The
fast path swallows every exception and the locked path does not, and the
load path also re-derives the sig, so a shared helper is not a pure
extraction at the same semantics; left in place.
PROOF: no behaviour change; tests/hermes_cli/test_config_lock_free_cache_hit.py
4 passed, probes/S3-fold-thrash.py still 2/100 uncached resolves.
`_read_raw_config_impl` had the same lock-on-hit shape #117440 removed
from `_load_config_impl`: a microsecond cache hit queued behind
`_CONFIG_LOCK`, which `save_config()` holds across an atomic YAML write,
so one background config write stalled every raw read (per-turn policy
checks on the gateway). `_RAW_CONFIG_CACHE` already publishes each entry
as one `(*sig, data)` tuple replaced wholesale, so the lock never
protected the read; check the signature first without it and fall
through to the locked re-parse on a miss.
A cache hit in `_load_config_impl` costs microseconds, but it was served
from inside `_CONFIG_LOCK` — which `save_config()` holds across an atomic
YAML write. Measured on a clean checkout, driving the real functions
against a temp HERMES_HOME:
cache-hit read, uncontended median 0.0237ms
the SAME cached read while another
thread holds _CONFIG_LOCK 10010.2ms
On a gateway this lands on the event loop. `invoke_hook` calls
`_resolve_hook_callback_timeout`, which reads config, and a gateway fires
hooks on every inbound message — so one background config write stalls
every message for the full duration of that write. The same probe showed
the hook path doing 100 config reads across 100 hook invocations.
Two changes:
1. `_load_config_impl` gets a lock-free fast path for cache hits. The lock
never protected the cache dict: CPython dict get/setitem are atomic under
the GIL, and the cached tuple is replaced wholesale rather than mutated
in place, so a reader observes either the complete old tuple or the
complete new one. The lock's real job is serializing the rebuild
(parse + merge + expand) and the writers. Worst case on a race is a
redundant rebuild, which the locked path re-checks and collapses. The
existing `_load_config_cache_sig()` helper is reused, so the fast and
locked paths cannot drift on freshness.
2. `_resolve_hook_callback_timeout` is memoized on that same signature. The
value only changes when config.yaml does; every other call is a dict
lookup. Validation and clamping move unchanged into
`_resolve_hook_callback_timeout_uncached`.
After: the blocked read returns in 0.0ms and the hook path does 1 config
read per 100 invocations.
Tests (tests/hermes_cli/test_config_lock_free_cache_hit.py, 11 tests) drive
the real functions against a temp HERMES_HOME — no mocks of the code under
test. They cover the blocked-read case for both the readonly and deepcopy
entry points, that the fast path still sees a changed file, that
`load_config()` still returns an isolated object, a 4-reader + 1-writer
concurrency arm asserting no torn observation, and the hook memo's
freshness plus its clamp/fallback/zero-disables contract.
RED/GREEN on this base, impl reverted via git stash:
without the change : 8 failed, 3 passed (blocked read: 30.0s)
with the change : 11 passed
Neighbours green: tests/hermes_cli/test_config.py, test_plugins.py,
test_config_loader_e2e.py, test_managed_scope_loaders.py,
test_read_raw_config_readonly.py — 266 passed, 4 skipped.
(cherry picked from commit b787427bdaef81d8b114194779fe1de3b4efe7bd)
`_platform_plugin_manifests()` scanned only the repo's `plugins/platforms/*`, so a
third-party platform plugin under `<HERMES_HOME>/plugins/` never reached
`OPTIONAL_ENV_VARS`: the Desktop Gateway form and `hermes config` showed bare
variable names with no prompt, description or password masking. It now also
walks `<HERMES_HOME>/plugins/platforms/*` and flat `<HERMES_HOME>/plugins/*`
manifests that declare `kind: platform`.
Slim redo of #46964 by @LeonSGP43 onto the refactored helper (the original
predates `_platform_plugin_manifests` and replaced `fast_safe_load`).
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.
Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.
- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
A list/mapping slot holding ONE quoted string (`plugins:\n enabled: '["a","b"]'`,
`model_catalog:\n excluded_providers: '["openai-api"]'`) is skipped by every
isinstance-gated reader (`plugins_cmd._config_name_set` -> set(), `plugins._names`,
`inventory.excluded_providers` -> []) while `config get` echoes it back, so user
plugins silently unmount and provider exclusions silently lapse. Neither `hermes
doctor` nor the startup `print_config_warnings` banner said a word.
Reader/doctor half: one schema-aware pass in `validate_config_structure`
(`_validate_quoted_containers`) walks `DEFAULT_CONFIG` (sections included) plus
`_KNOWN_CONTAINER_TYPES` and warns when the user value is a string that parses to
a list/mapping. The message names the key, the quoted value and the remedy
(`hermes config set <key> '<literal>'`). Finding only: the file is never
rewritten. String-typed keys (`approvals.mode: "[off]"`), the `model: <name>`
shorthand and the `parse_config_string_list`-read slots
(`agent.disabled_toolsets`, `skills.disabled`) are not flagged. Feeds both the
doctor "Config Structure" section and the startup banner.
Writer half (audit of every path that can put a string in a container slot):
- `hermes config set` (`hermes_cli/config.py::set_config_value`): already parses
bracket/brace literals (#88163) and refuses wrong-shaped values
(`_refuse_container_type_mismatch`), BUT the guard only knew slots present in
DEFAULT_CONFIG or `_KNOWN_CONTAINER_TYPES`. `plugins.enabled`,
`plugins.disabled` and `model_catalog.excluded_providers` are deliberately
absent from DEFAULT_CONFIG, so `config set plugins.enabled foo` /
`plugins.enabled a,b` / `model_catalog.excluded_providers openai-api` stored a
plain string (live repro on base). Added the three keys to
`_KNOWN_CONTAINER_TYPES`: those writes are now refused with the literal hint.
- `cli.py::save_config_value` + callers (cli_*_mixin, gateway/slash_commands,
gateway/run_busy): every caller passes a bool/enum string for scalar keys; no
list-slot caller. Not reachable.
- `hermes_cli/plugins_cmd.py::_save_plugin_sets` / `_write_config_value` and
`plugins_cmd_catalog` (via `_save_enabled_set`): write `sorted(set)` — real
lists. Not reachable.
- Dashboard `PUT /api/config` (`web_routers/config_env.py::update_config` ->
`_denormalize_config_from_web` -> `save_config`): schema-driven form;
`web/src/components/AutoField.tsx` splits list-typed fields into a real array
before the PUT. Not reachable.
- tui_gateway `config.set`: fixed `_CONFIG_SETTERS` table of scalar keys only
(out of this lane's files anyway). Not reachable.
Conclusion: the quoted shapes on real machines are leftovers of pre-#88163
`config set` runs plus the DEFAULT_CONFIG-absent slots fixed here.
Live repro (fake HOME/HERMES_HOME, both quoted shapes in config.yaml):
before: validate_config_structure() -> []; startup banner silent; doctor has
no Config Structure section; _config_name_set('plugins','enabled') -> set()
after: two warnings naming plugins.enabled / model_catalog.excluded_providers
with `hermes config set ... '["a","b"]'`; doctor prints them under
Config Structure; running the remedy yields real lists,
validate -> [], _config_name_set -> {'a','b'}
writer: `config set plugins.enabled a,b` wrote `enabled: a,b` on base; now
refused ("must be a list ... nothing was written").
Fixes#83308Fixes#105706
credit: @fangliquanflq #105725 (schema-aware detection in validate_config_structure; slim redo)
credit: @Luna161 #83313 (first report + per-key warning in plugins_cmd)
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
An unquoted YAML scalar (provider: 2) loads as int, and downstream
readers call (provider or "").strip() - a gateway turn dies before the
agent runs (#117345). Normalize the routing key at the load/save
chokepoint (_normalize_root_model_keys) with the existing
coerce_provider_id helper (#96505), so every reader sees a string and
the next save heals the persisted value. The early-exit condition now
also enters work when model.provider is a non-string, so a plain
model.provider config reaches the normalize; the key is only coerced
when present, so provider-less configs gain no empty key (and no
config.yaml churn on the next save).
Fixes#117345 (read-path half; the write-path half keeps `hermes config
set` from persisting the int in the first place)
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Review follow-up: get_managed_system, the container/chmod-skip check and
the HERMES_UID/GID chown (_resolve_hermes_uid_gid/_chown_to_hermes_uid) now
live only in hermes_constants — the import-safe module apply_secure_dir_policy
already needed them in — and hermes_cli.config re-exports them, so there are
no 'keep in sync' copies and _secure_file skips on the same canonical
_detect_container signal as _secure_dir/get_scratch_dir. The dead config
copies are deleted and test_ensure_hermes_home_uid.py drives the new
symbols; one invariant test pins the single implementation.
get_scratch_dir unconditionally chmod'ed <home>/cache/scratch to 0700,
bypassing the policy _secure_dir() applies everywhere else: managed
installs are left alone, containers only apply an explicit
HERMES_HOME_MODE, and HERMES_UID/HERMES_GID ownership is honored. On a
group-shared home (setgid 2770) that stripped the setgid bit and broke
the shared-group inheritance the operator configured (#77579, #117347).
Move the policy primitive to hermes_constants.apply_secure_dir_policy()
(import-safe, stdlib-only) and route both get_scratch_dir() and
hermes_cli.config._secure_dir() through it, so one implementation
serves all callers including the cron.jobs delegation chain.
Fixes#117347
The gateway.platforms.<p>.<field> -> platforms.<p>.<field> canonicalization was
applied to `config get` and `config unset` too, so a legacy config whose only
value lived under the nested block reported "Config key not set" (get) and
left the nested key in place (unset) while merge_platform_sections kept
honouring it. get now falls back to the nested key when the canonical one is
absent, unset removes both spellings, and set drops the shadowed nested
duplicate so the file keeps one source of truth.
`hermes config set gateway.platforms.telegram.enabled true` printed `✓ Set` and wrote a
nested `gateway.platforms` block, but `merge_platform_sections` gives the top-level
`platforms.<p>` block precedence on shared keys, so an existing top-level
`enabled: false` kept the platform off and nothing surfaced the disagreement.
Canonicalize the nested prefix onto `platforms.<p>.<field>` in
`_redirect_platform_display_key` (the same seam #71047 uses for display settings, so
set/get/unset all agree) and print a note naming the precedence. A nested display
setting (`gateway.platforms.<p>.streaming`) continues on to `display.platforms`.
Based on the analysis in #115226 (@chenzeyan54-commits) and the precedence
measurements by @KeyArgo on the issue thread.
Fixes#115212
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.
`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.
Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).
Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
The denylist in save_env_value is the fail-closed gate between the
dashboard's PUT /api/env surface (and every other env writer) and .env,
which lands in os.environ for every subprocess Hermes spawns. It covered
LD_PRELOAD, PYTHONPATH, PATH, EDITOR and GIT_SSH_COMMAND, but missed most
members of the class it states: env-driven git config injection
(GIT_CONFIG_PARAMETERS, GIT_CONFIG_COUNT, GIT_CONFIG_KEY_n/GIT_CONFIG_VALUE_n,
GIT_CONFIG_GLOBAL/SYSTEM), executed git helpers (GIT_SSH, GIT_ASKPASS,
GIT_EDITOR, GIT_SEQUENCE_EDITOR, GIT_PAGER, GIT_EXTERNAL_DIFF,
GIT_PROXY_COMMAND, GIT_TEMPLATE_DIR, GIT_DIR), credential-prompt helpers
(SSH_ASKPASS, SUDO_ASKPASS), shell init files (BASH_ENV, ENV, ZDOTDIR,
PROMPT_COMMAND, VIMINIT, EXINIT, MANPAGER), and interpreter/toolchain
injection (PERL5OPT/PERL5LIB/PERLLIB, RUBYOPT/RUBYLIB, PYTHONBREAKPOINT,
PYTHONCASEOK, CLASSPATH, JAVA_TOOL_OPTIONS/_JAVA_OPTIONS/JDK_JAVA_OPTIONS,
GOFLAGS, RUSTFLAGS).
GIT_CONFIG_KEY_n/GIT_CONFIG_VALUE_n pairs are unbounded, so a new
_ENV_VAR_NAME_DENY_PREFIXES tuple matches whole families by prefix:
LD_, DYLD_, GIT_CONFIG_. The check runs on the Windows-corrected policy
name, so mixed-case spellings are refused there while POSIX's inert
lowercase names stay writable.
git credential helpers and subprocess env construction already null
GIT_CONFIG_GLOBAL/SYSTEM and strip GIT_ASKPASS from child environments
for the same reason, so this closes the write surface those hygiene
paths assume.
Regression tests cover every denied name, the unbounded pairs, near-miss
names that must stay writable (GIT_COMMITTER_NAME, GIT_AUTHOR_NAME,
GIT_TERMINAL_PROMPT, POSIX-lowercase spellings), and Windows mixed-case
denial.
Custom providers already declare their wire protocol (api_mode / legacy
transport) on the entry, but nothing outside routing could read it: context
length resolution keyed Codex detection on the hostname, which a proxy on
127.0.0.1 never matches. get_custom_provider_api_mode() returns the canonical
transport of the first entry serving a base_url (None custom_providers loads
config, like the sibling context_length helper) so metadata decisions can
follow the route instead of the host.
Slim port of the helper from #116262 by @JoaoMarcos44.
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.
Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):
- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
a base_url is: the target's registry/plugin host, a `providers:` /
`custom_providers:` entry resolving to the target, or bare custom/local
aliases (configured BY base_url) -> owned; another known provider's host
or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
alone, with no base_url, is old-route wire state and goes too); keeps an
owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
prints what was cleared and why, or warns that an unrecognised base_url
still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
(openai-codex + chatgpt.com), a custom entry with that URL, and bare
`custom` are untouched.
Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
When a user runs `hermes config set model.provider <new>` to switch
providers, the old provider's `model.base_url` is left behind in
config.yaml. This causes API calls to go to the wrong endpoint.
The wizard flow (`_update_config_for_provider`) already handles this
correctly by clearing stale base_url on provider switch, but the
`hermes config set` path did not.
Now `set_config_value` detects when `model.provider` is being changed
and removes the stale `model.base_url`, allowing runtime
auto-detection to resolve the correct endpoint for the new provider.
Fixes#40862
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.
The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.
Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).
Part of #114471 (items 2, 5, 6).
Unioning _EXTRA_KNOWN_ROOT_KEYS into _known_top_level_keys() also widened the
wrong-prefix pre-write refusal in _validate_config_key: the extra roots include
the top-level forms of nested gateway settings bridged by gateway/config_loader.py
(filter_silence_narration, reset_triggers, always_log_local, ...), so
`config set gateway.filter_silence_narration false` flipped from main's
write-with-notice (rc 0) to a refusal (rc 1, 'Did you mean: filter_silence_narration').
Restrict the stripped-suffix root test to set(DEFAULT_CONFIG) | _OPEN_SUBKEY_TOP_LEVEL_KEYS
so those paths keep main's behaviour while gateway.discord.<field> / agent.gateway.strict
stay refused. Control param added to the written-with-notice test; drop the redundant
smart_model_routing.enabled known-key param so the PR adds two invariant params.
_known_top_level_keys() built the known-key set from DEFAULT_CONFIG plus
the open-subkey categories, but omitted _EXTRA_KNOWN_ROOT_KEYS — the
documented set of roots that are valid on disk yet intentionally absent
from DEFAULT_CONFIG (platform_toolsets, known_plugin_toolsets,
smart_model_routing, etc.).
Consequence: 'hermes config set platform_toolsets.cli ...' printed a
false '⚠ not a recognized config key' notice with a bogus suggestion
('Did you mean: platform_hints.cli') even though the key is read by the
setup wizard and tools_config save flow. The value was written, but the
misleading warning (and wrong suggestion) appeared on every such write.
Include _EXTRA_KNOWN_ROOT_KEYS in the union so all documented roots
validate as known, and extend TestValidateConfigKey.test_known_keys_pass
to cover them.
Salvage note (rebased onto current main): the original hunk targeted the
pre-rewrite multi-statement body; the fix is the same single union term.
Test params trimmed to two representatives (one toolset root, one other
extra root); 'session_reset' is not in _EXTRA_KNOWN_ROOT_KEYS on main.