They are not "never auto-repaired": resuming such a session re-pins the
full tool surface (restore_agent_tool_prefix), after which a scan sees
skill_manage without the Skill Safety guidance and clears the prompt.
The repair-prompts detector cleared two legitimate prompts:
- a pin with skills_list/skill_view but no skill_manage and zero skills
installed: build_skills_system_prompt returns '' and SKILLS_GUIDANCE is
only emitted with skill_manage, so the healthy prompt has neither marker;
- an exact memory-only pin, which is also a user toolsets=[memory] config;
the healthy rebuild was re-flagged on every run (not idempotent).
Key the decision on the missing '## Skill Safety' guidance, which is
unconditional when skill_manage is in the pin, and require skill_manage in
the pin. Memory-only rows are now reported as unverifiable; the explicit
SESSION_ID override still clears them and their pin.
Docs: describe the tightened rule and note that a running gateway keeps
cached prompts in memory, so it must be restarted after --apply.
`hermes import` only checked the central directory (`is_zipfile`,
`namelist`), so an archive with one member whose deflate stream or CRC is
rotten passed validation and blew up mid-restore with a zlib.error
traceback -- after config.yaml and everything before the bad member had
already been replaced, with later members never written (#121258).
Add a pre-flight pass in run_import that streams every member through
1 MiB reads (zipfile verifies the CRC at EOF) and collects every
BadZipFile / zlib.error / EOFError. If any member is damaged the command
prints a capped list and returns 1 with the home untouched. Stdlib
`ZipFile.testzip()` is deliberately not used: it lets zlib.error escape
and names at most the first bad member.
Widen the per-member catch from the previous commit with EOFError so a
member that rots between the two passes still becomes a "skipped" warning
+ `Import incomplete` / exit 1 instead of a traceback. Rework that
commit's test to drive the per-member path (the archive now never reaches
it with a corrupt member), and let `_break_member` serve the pre-flight
read before failing the restore's own read.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
`_import_members` catches the PermissionError/OSError of each member it
cannot publish and records it in `errors`. That includes the state.db the
live-safe restore refuses because a gateway or dashboard still holds it
(#100960, #110179). `run_import` printed those under "Warnings (N files
skipped)", then printed "Done. Your Hermes configuration has been
restored." and returned None. `cmd_import` did not forward a return value,
so the shell status was 0. A script chained on `hermes import --force`
carried on over a partial restore, and the dashboard's import action
showed the green "done" badge. For the refused state.db case, the user's
sessions were never restored.
`run_import` now returns 1 when `errors` is non-empty. It words both the
summary header and the final line as "Import incomplete", and `cmd_import`
forwards the return code, the same contract `hermes backup` has for an
incomplete archive. Runtime files the import deliberately keeps
(gateway.pid, SQLite sidecars) and the older-backup session warning stay
warnings. The gateway revive still runs, so the files that did land come
up as before.
Measured on main with a fresh HERMES_HOME through `main()`: a 3-file
backup with one read-only target directory, and a refused state.db (live
DB left at 3 sessions / 12 messages). Both used to exit 0 after "Done ...
restored". Both now exit 1 after "Import incomplete: 1 file(s) were not
restored". A clean import still exits 0 and prints "Done".
(cherry picked from commit a0f808ec9f868128f4174e4816e095ac17f18d28)
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Follow-up to the salvaged mirror commit.
- Opt-in (`mirror_to_session: true`, or `hermes webhook subscribe --mirror-to-session`),
default off like cron's `mirror_delivery`. The mirrored text lands with user authority in
the target chat, and on `deliver_only` routes it is the raw rendered payload, so a route
author has to ask for it. Only a real boolean `true` opts in (a YAML string "false" no
longer does).
- The mirror runs inside the routed profile's scope, so a `/p/<profile>/` delivery writes
into THAT profile's state.db. A Telegram DM chat_id is the user's id on every bot; an
unscoped mirror landed in the default profile's DM with the same person (live-probed).
- Tests trimmed to two invariants against a real state.db: an opted-in delivery lands in the
routed profile's chat session and not the default profile's; a route without the opt-in
(including `mirror_to_session: "false"`) never touches the target transcript.
- Docs: webhooks guide section "Replying to a delivery" (default, profile scope, trust note),
CLI reference row, bundled hermes-agent skill reference.
Follow-up to the salvaged wrapper: the argument layer now lives in a topical
sibling (agent/legacy_cli.py) instead of growing the run_agent facade, and
`python run_agent.py` (the installer's PATH launcher for hermes-agent) routes
through the same parser instead of fire, so `--version` and a bare invocation
no longer run the demo turn there either.
- --help/-h/--version use argparse's own exit path; options carry help text
- the runner is injected (`run=`) so `python run_agent.py` does not import
run_agent a second time
- tests trimmed to two invariants that resolve the console-script target
from pyproject exactly like pip's wrapper (red on main, green here)
- docs: hermes-agent section in the CLI reference
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.
Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
A list/mapping slot holding ONE quoted string (`plugins:\n enabled: '["a","b"]'`,
`model_catalog:\n excluded_providers: '["openai-api"]'`) is skipped by every
isinstance-gated reader (`plugins_cmd._config_name_set` -> set(), `plugins._names`,
`inventory.excluded_providers` -> []) while `config get` echoes it back, so user
plugins silently unmount and provider exclusions silently lapse. Neither `hermes
doctor` nor the startup `print_config_warnings` banner said a word.
Reader/doctor half: one schema-aware pass in `validate_config_structure`
(`_validate_quoted_containers`) walks `DEFAULT_CONFIG` (sections included) plus
`_KNOWN_CONTAINER_TYPES` and warns when the user value is a string that parses to
a list/mapping. The message names the key, the quoted value and the remedy
(`hermes config set <key> '<literal>'`). Finding only: the file is never
rewritten. String-typed keys (`approvals.mode: "[off]"`), the `model: <name>`
shorthand and the `parse_config_string_list`-read slots
(`agent.disabled_toolsets`, `skills.disabled`) are not flagged. Feeds both the
doctor "Config Structure" section and the startup banner.
Writer half (audit of every path that can put a string in a container slot):
- `hermes config set` (`hermes_cli/config.py::set_config_value`): already parses
bracket/brace literals (#88163) and refuses wrong-shaped values
(`_refuse_container_type_mismatch`), BUT the guard only knew slots present in
DEFAULT_CONFIG or `_KNOWN_CONTAINER_TYPES`. `plugins.enabled`,
`plugins.disabled` and `model_catalog.excluded_providers` are deliberately
absent from DEFAULT_CONFIG, so `config set plugins.enabled foo` /
`plugins.enabled a,b` / `model_catalog.excluded_providers openai-api` stored a
plain string (live repro on base). Added the three keys to
`_KNOWN_CONTAINER_TYPES`: those writes are now refused with the literal hint.
- `cli.py::save_config_value` + callers (cli_*_mixin, gateway/slash_commands,
gateway/run_busy): every caller passes a bool/enum string for scalar keys; no
list-slot caller. Not reachable.
- `hermes_cli/plugins_cmd.py::_save_plugin_sets` / `_write_config_value` and
`plugins_cmd_catalog` (via `_save_enabled_set`): write `sorted(set)` — real
lists. Not reachable.
- Dashboard `PUT /api/config` (`web_routers/config_env.py::update_config` ->
`_denormalize_config_from_web` -> `save_config`): schema-driven form;
`web/src/components/AutoField.tsx` splits list-typed fields into a real array
before the PUT. Not reachable.
- tui_gateway `config.set`: fixed `_CONFIG_SETTERS` table of scalar keys only
(out of this lane's files anyway). Not reachable.
Conclusion: the quoted shapes on real machines are leftovers of pre-#88163
`config set` runs plus the DEFAULT_CONFIG-absent slots fixed here.
Live repro (fake HOME/HERMES_HOME, both quoted shapes in config.yaml):
before: validate_config_structure() -> []; startup banner silent; doctor has
no Config Structure section; _config_name_set('plugins','enabled') -> set()
after: two warnings naming plugins.enabled / model_catalog.excluded_providers
with `hermes config set ... '["a","b"]'`; doctor prints them under
Config Structure; running the remedy yields real lists,
validate -> [], _config_name_set -> {'a','b'}
writer: `config set plugins.enabled a,b` wrote `enabled: a,b` on base; now
refused ("must be a list ... nothing was written").
Fixes#83308Fixes#105706
credit: @fangliquanflq #105725 (schema-aware detection in validate_config_structure; slim redo)
credit: @Luna161 #83313 (first report + per-key warning in plugins_cmd)
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):
- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
as catalog:official everywhere and marked the real entry "installed". Provenance
now lives on the installer-owned .install-metadata.json record (catalog block
written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
never reads the tree. Pre-fix installs are adopted once when the installer
record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
data and user patches with no warning. Untracked/ignored files are carried
into the new tree; edits to tracked files are copied to
<HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.
repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.
Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.
Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.
Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).
Completes #118273 (docs item). Part of #109417
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
Keep the in-process return-code matrix (issues / manual / --fix full and
partial repair) and the real-process exit-status check; drop the exception
passthrough, --ack and --live cases, which pin behaviour this change does not
touch. Document 0/1 exit status under `hermes doctor` in the CLI reference.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.
Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).
Co-authored-by: himanusia <himanusia@users.noreply.github.com>
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.
Part of #95794
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.
Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.
Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.
`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:
1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
generations, usage rows, system prompt) to the owning store, parents before
children so lineage survives, copy-then-delete so a crash leaves a duplicate
the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
-> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
import re-injects them into routing every boot).
Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).
Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).
Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.
`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.
Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.
Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):
- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
a base_url is: the target's registry/plugin host, a `providers:` /
`custom_providers:` entry resolving to the target, or bare custom/local
aliases (configured BY base_url) -> owned; another known provider's host
or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
alone, with no base_url, is old-route wire state and goes too); keeps an
owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
prints what was cleared and why, or warns that an unrecognised base_url
still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
(openai-codex + chatgpt.com), a custom entry with that URL, and bare
`custom` are untouched.
Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.
The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.
Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).
Part of #114471 (items 2, 5, 6).
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.
Docs: doctor reference and the network config section describe the check.
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.
agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.
Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.
Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.
Config-set atom of #114471.
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).
What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
(spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
`_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
`/proc/<pid>/environ` never reflects it. Without this, the common default-home install
(nothing exported) got a `--stop` that matched nothing — the regression flagged on
#113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
#113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
only a foreign backend prints "No hermes dashboard processes running for this profile."
instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
the (now filtering) finder. Docs: `--stop` row in the CLI reference.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).
resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.
Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).