Commit Graph

753 Commits

Author SHA1 Message Date
teknium1
4590ef8b58 fix(profiles): polled profile lists never walk skill trees; vanished skill dirs no longer abort enumeration
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.

- list_profiles(lazy_skill_count=True): skill_count is the last known value;
  a missing/aged entry schedules ONE background _count_skills per profile per
  60 s recheck window, so the request thread does zero skill-tree I/O and the
  refresh cadence is decoupled from the poll rate. The two polled callers and
  the REST fallback entry use it; the sync default (CLI, detail views) is
  unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
  os.walk with excluded/support dirs pruned) instead of rglob — half the
  fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
  uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
  projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
  skill add/remove immediately on the next refresh).

Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
  before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
          profiles.list 2189 + 2015, projects/tree 2177 + 2015;
          a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
  after:  GET /api/profiles 0 skill-tree stats/scandir on the request thread
          (counts land from the background refresh by the next poll),
          profiles.list 0, projects/tree 0; _count_skills (detail/control)
          still reports 200 with 203 scandir; the vanished-dir case returns
          199 and list_profiles() enumerates all 5 profiles.

Fixes #114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:18:40 -07:00
teknium1
72360ae1d2 fix(doctor): detect a dead IPv6 route and name network.force_ipv4
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.

Docs: doctor reference and the network config section describe the check.
2026-09-18 09:56:28 -07:00
teknium1
fc94ff56e1 fix(cli): expand HERMES_HOME once at process entry; sweep raw readers
Follow-up to the salvaged hermes_constants expansion (#109212): the CLI has
~30 raw `os.environ["HERMES_HOME"]` readers (fast --version path, early
display.interface probe, profile re-home, dotenv loader, subprocess-home
helper) that never go through get_hermes_home(). A literal `~` (fish, or
any quoted value) left them resolving `~/.hermes` against cwd while the
resolver now expands it, so the process would disagree with itself.

Normalize the env var once, at the earliest point of hermes_cli/main.py
(stdlib-only, before the fast paths), and expand it in the one raw
hermes_constants reader (`_profile_home_path`). A relative value that is
not tilde/variable-shaped is deliberately left alone: rejecting it would
break `HERMES_HOME=./tmp-home` in tests and CI for no user-facing gain.

Tests: real-CLI subprocess with HERMES_HOME='~/.x' under a fake HOME asserts
`config path` lands in the fake home and that no literal `~` directory
appears under cwd (the reporter's acceptance criterion, #114353), plus a
unit test for the normalizer. Docs: environment-variables reference.
2026-09-18 09:49:46 -07:00
teknium1
b65ec6a3f0 docs: list the DISCORD_MISSED_MESSAGE_BACKFILL_* env fallbacks
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
2026-09-18 09:44:59 -07:00
teknium1
9c088a7a78 fix(config): fix container type for unseeded roots, top-level lists and bare-name list slots
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.

agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
2026-09-18 09:41:02 -07:00
teknium1
0972b4819a fix(config): config set refuses a wrong-shaped value for a list/mapping key
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.

Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.

Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.

Config-set atom of #114471.
2026-09-18 09:41:02 -07:00
teknium1
46500f9c79 fix(dashboard): --stop and the update sweep resolve backend ownership tri-state, never guess
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).

What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
  (spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
  else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
  with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
  `_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
  `/proc/<pid>/environ` never reflects it. Without this, the common default-home install
  (nothing exported) got a `--stop` that matched nothing — the regression flagged on
  #113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
  #113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
  only a foreign backend prints "No hermes dashboard processes running for this profile."
  instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
  hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
  their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
  the (now filtering) finder. Docs: `--stop` row in the CLI reference.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 09:38:01 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
Ritesh Patel
9927199998 docs: drop opencode-free references (provider removed)
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
2026-09-18 15:40:37 +05:30
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
98b4efbf48 fix: clarify cards that cannot render are re-asked as plain text, never reported as user inactivity
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could
fail without the user ever seeing a question, and the agent then waited out the full
clarify_timeout and reported "[user did not respond within Nm]" (#112684):

- the platform rejects the card at once -> the wait aborted with a delivery sentinel but the
  question was never re-asked;
- the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts,
  a stale-thread retry) -> the runner never looked at the send future again and blocked for
  the whole timeout for a card that never posted;
- no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported
  timed_out=True with an empty notice.

gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers
move out of the gateway/run.py facade) now retries every definitive card failure once through the
adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress
DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a
possibly-delivered send so a late failure releases the waiter with
"[clarify prompt could not be delivered]". The no-surface case reports
"[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already
treats as a non-answer).

Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s):
card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and
the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt,
after the numbered prompt is sent and "2" resolves to the choice.
2026-09-17 09:06:17 -07:00
teknium1
f3bdcd0877 fix: dashboard stop sweeps the wedged descendants that outlive the SIGKILLed backend
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.

_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.

Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.

The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).

Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
2026-09-17 09:04:11 -07:00
teknium1
c61a11474b fix(desktop): update check falls back to the gh CLI login before going anonymous
Closes the last atom of #112615. A Desktop app started from the Dock, Finder or a
desktop launcher inherits a minimal environment, so the GITHUB_TOKEN / GH_TOKEN rung
landed in #113190 never fires for exactly the users the issue is about (shared exit
IP, 60/hour anonymous budget exhausted by neighbours). The update check now walks the
same ladder as the Python GitHub client (tools/skills_hub_github.py::GitHubAuth):

1. GITHUB_TOKEN, then GH_TOKEN, from the launch env (unchanged).
2. `gh auth token` — execFile with an argv array (no shell), stdin closed, 3 s
   timeout, windowsHide. `gh` is resolved on PATH plus the GUI-safe install
   locations backend-env already appends for the backend (Homebrew, /usr/local,
   ~/.local/bin, nix profile, the Windows GitHub CLI installer dirs). The outcome
   (token or none) is cached for the process lifetime so gh runs at most once.
3. Anonymous.

A rejected credential (401) still retries anonymously; the log line names the
source (env vars vs gh login) and never the token, once per source per process.
A rejected gh token also drops the cache so a re-login is picked up by the next
check. envTokenRejected is renamed githubTokenRejected since both rungs use it.

Live: with PATH=/usr/bin:/bin:/usr/sbin:/sbin and no env token, the module found
gh, resolved a credential in 135 ms and api.github.com reported a 5000/hour core
budget (anonymous: 60). A logged-out gh and a hanging gh both resolved to null
(anonymous) in 5 ms / 3001 ms.

Docs: the credential ladder is documented on the Desktop "Updating" page and in
the environment-variables reference.
2026-09-17 08:54:13 -07:00
kshitijk4poor
f537076738 docs(config): config set docs describe the wrong-prefix-only refusal
The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.

Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.
2026-09-17 21:19:22 +05:30
fangliquanflq
b632f20484 docs(mcp): clarify transport keepalive defaults 2026-09-17 21:18:03 +05:30
teknium1
8a8bd93ef2 docs: -z exit codes differ from chat -q/-Q on purpose; real turn_exit_reason example
The two one-shot entry points map the same outcome to different codes
(failed/partial/budget -> 2 on -z, 1 on -q/-Q; completed-with-no-text -> 1 vs 0)
and the reference documented both tables 60 lines apart without saying so.
Also `iteration_limit` is a metrics end_reason, not a turn_exit_reason; the
finalizer emits `max_iterations_reached(N/M)`.
2026-09-16 18:12:34 -07:00
teknium1
391a7007fb fix(cli): hermes -z exits non-zero for partial, failed and interrupted runs even when text was printed
`hermes -z` judged the run by whether stdout was non-empty: a turn that hit the
iteration budget, ended `partial`, failed on the provider, or was interrupted
still exited 0 as long as it printed an explanation, so scripts treated a
half-done job (or a 401 summary) as success. The `-q`/`-Q` paths already got
one result-derived contract in 23863ccbafe0d; this is the same class on the
scripted one-shot entry point (#111770).

- `hermes_cli/oneshot.py::_oneshot_exit_code`: 0 only when the turn completed;
  130 interrupted; 2 failed / partial / completed:false; 1 a completed turn
  with no text (unchanged message).
- `--usage-file` now also reports `partial`, `interrupted` and
  `turn_exit_reason`, so a pipeline can tell a budget stop from a Ctrl-C.
- Docs: `hermes -z` exit-code table + usage-file fields in cli-commands.md.

Live probe (run_oneshot with a stubbed _run_agent): partial-with-text 0 -> 2,
budget-with-text 0 -> 2, interrupted-with-text 0 -> 130, failed-with-text 0 -> 2;
completed 0 -> 0 (control), partial-no-text 2 -> 2 (unchanged).
2026-09-16 18:12:34 -07:00
teknium1
6671365fb2 test: one invariant suite for mid-turn session-switch refusals + docs
Replace the two salvaged per-command test files with a single suite that drives
the real /branch, /resume and /sessions <id> handlers against a real SessionDB
and asserts the invariant that matters: while a turn is running, the session id
(CLI and agent) is unchanged and the current row is not ended; when idle,
/branch still proceeds. Red on origin/main (3 busy cases), green with the guards.

The /branch MagicMock fixture gains an explicit `_agent_running = False`: a bare
MagicMock attribute is truthy and would trip the new guard.

Docs: /branch and /resume rows in the slash-command reference say they are
refused mid-turn in the classic CLI and why.

Co-authored-by: sonderhq <144971773+sonderhq@users.noreply.github.com>
Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 17:49:48 -07:00
teknium1
b3aef45987 fix(clarify): batch result carries the surface's no-answer notice; trim tests
Follow-up to the #112725 salvage. When a gateway clarify batch stops at an
unanswered question, the batch payload only said `timed_out: true` with blank
answers — so a card the platform REJECTED ("[clarify prompt could not be
delivered]") read exactly like a user who walked away, which is the misreport
#112684 describes ("a timeout must not be presented as user inactivity when the
prompt was never delivered").

- gateway/run_turn_runner.py::_clarify_batch_sync: the unanswered question's
  own response text rides along as `notice`.
- tools/clarify_tool.py::_run_batch/_batch_result: pass `notice` through into
  the result JSON beside `timed_out` (only when present); schema description
  mentions it. CLI/TUI batch callbacks send no notice -> byte-identical output.
- tests: fold the single-question re-arm pin into the existing bracket-answer
  test (<=2 new tests), parametrize the end-to-end test over an undeliverable
  Telegram-shaped adapter so the delivery sentinel is pinned as `notice`.
- docs: tools-reference.md describes `notice`.

Part of #112684
2026-09-16 17:47:50 -07:00
teknium1
da942e4483 fix(config): hedge the config get phantom-key notice — the schema walk has false positives
`_validate_config_key` is a DEFAULT_CONFIG walk, and several live keys are
deliberately unseeded so that a stored value counts as an explicit user pick
(browser.cloud_provider, stt.provider, cron.misfire_grace_minutes,
gateway.proxy_url / relay_url, dashboard.ssh_isolated_idle_grace_s,
computer_use.backend). The runtime reads all of them and Hermes' own wizards
write some of them, so `hermes config get browser.cloud_provider` asserting
"Hermes does not read it" was a false claim.

Match the set-path notice ("Hermes may not read it"): the key still gets
flagged so a leftover misspelling cannot pass unnoticed, but the CLI no longer
states something it cannot know. Docs updated to the same wording; one test
pins the hedge on an unseeded live key (red before this commit).
2026-09-16 17:42:32 -07:00
teknium1
efc0359051 fix(config): inline the phantom-key notice into get_config_value, trim to 2 tests, docs
Slim follow-up to the salvaged commit from #112350 (@DavidMetcalfe): the two
helpers (`_is_unread_nested_key`, `_print_unread_key_notice`) collapse into
~10 lines at the single call site in `get_config_value`, reusing
`_validate_config_key` for both the verdict and the did-you-mean suggestion
instead of calling it twice. `is_env_key` is dropped: `.env` keys and
UPPER_SNAKE settings are never a known top-level section, so the guard
already excludes them.

Tests: keep the two invariants (phantom nested key -> notice on stderr while
stdout/--json stays parseable; known / custom top-level / open-subkey keys
are not flagged). Dropped `test_missing_key_exits_without_a_notice` — the
missing-key exit precedes the print and is pre-existing behaviour already
covered elsewhere.

Docs: one sentence each in cli-commands.md (`config get` row) and
configuration.md (the set-path did-you-mean paragraph).
2026-09-16 17:42:32 -07:00
teknium1
c7d4d32800 fix(desktop): env-only GitHub token for the update check, honest rate-limit copy
Trim the salvaged ladder to its env rung and make the 403 copy actionable.

- github-api-auth.ts keeps githubTokenFromEnv / githubApiHeaders; the
  `gh auth token` rung (resolveGithubToken, readGhCliToken, the per-process
  cache and the execFile import in main.ts) is dropped. Whether a GUI app may
  silently spend the user's gh CLI login on a passive background check is a
  product call, not a bug fix; the pure seam makes re-adding it a ~15-line
  follow-up.
- fetchGitHubApi reads GITHUB_TOKEN / GH_TOKEN from process.env per request
  (nothing stored) and attaches x-ratelimit-remaining / x-ratelimit-reset plus
  whether the call was authenticated to the rejected error.
- describeUpdateCheckFailure moves out of the unimportable main.ts into
  update-api-check.ts. A 403/429 with x-ratelimit-remaining: 0 now says the
  anonymous budget is 60/hour per network address shared with everyone behind
  the same connection, gives the real reset time and names GITHUB_TOKEN as the
  remedy; a 403 without rate-limit headers is reported as a plain HTTP 403
  instead of a rate limit the user did not cause.
- Docs: desktop.md "Updating" and the GITHUB_TOKEN row in
  environment-variables.md.
- Tests trimmed to two invariants per module.

Why: behind a shared exit IP "try again in an hour" is advice the user cannot
act on, and the old copy hid the one thing that does help.
2026-09-16 17:22:39 -07:00
teknium1
08aa67e473 docs: state that HERMES_DEBUG_INTERRUPT=0/false/off keeps interrupt tracing off
Both readers now share utils.env_var_enabled(), so the reference table
spells out the falsy values instead of implying "any value enables".
Also registers the contributor email mapping for the salvaged commit.
2026-09-16 17:21:14 -07:00
teknium1
f41ad3bc0a fix(gateway,cron): key Matrix handoff/cron thread seeds the way Matrix replies are keyed
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).

The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).

- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
  `get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
  use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
  `group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
  asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
  contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.

Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before  create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
        cron seed matrix🧵… ≠ inbound
after   create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)

Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
2026-09-16 17:18:44 -07:00
teknium1
adb4323e75 test(tools): platform-aware TZ assertion, Windows live-child contract, docs
tests/tools/test_code_execution.py::test_timezone_injected_when_set asserted
child_env["TZ"] == "America/New_York" unconditionally, which encoded the
Windows defect as intended behaviour; on win32 it now asserts TZ is absent and
keeps the POSIX assertion.

Add a windows_only live-child test to test_code_execution_windows_env.py: with
timezone: configured, a real Windows child's time.timezone and astimezone()
offset must equal the test process's own OS-zone view — the reporter's exact
witness (time.timezone == 0, +01:00 instead of -07:00, tzname still correct).

Docs (environment-variables.md, configuration.md): HERMES_TIMEZONE / timezone:
reaches execute_code children as TZ on Linux/macOS only; Windows children keep
the OS zone because the Windows C runtime parses only POSIX-form TZ strings.

Sibling env builders swept: the remote-backend prefix in
code_execution_tool.py (TZ=... before `python3 script.py`) targets POSIX
backends (Docker/SSH/Modal) and is unchanged; no other spawner forwards
HERMES_TIMEZONE into TZ. HERMES_TIMEZONE itself stays stripped from the child
(adf23550f5: under the multiplexed gateway it holds only the default
profile's value).

Refs #112233
2026-09-16 17:14:28 -07:00
teknium1
8899aeff53 fix(oneshot): --usage-file ledger reports auxiliary LLM spend
`hermes -z --usage-file` copied only the main-loop result, so title generation,
vision, compression, web_extract and background-review calls — recorded per task
in session_model_usage — never reached the pipeline ledger the flag advertises
as "so pipelines can always account for spend". The Insights page already folds
those rows in (#23270); the ledger is now consistent with it.

- SessionDB.auxiliary_usage_by_task(session_id): per-task sums over the
  session's compression lineage (aux calls bill to the id the turn started with
  while compression mints child ids mid-turn).
- oneshot snapshots aux usage before the turn and attaches the delta after it,
  so a resumed session's earlier runs are not re-billed.
- The report gains `auxiliary` (totals + `by_task`) and
  `total_including_auxiliary`; every existing key keeps its main-loop meaning.
- The auto-title upgrade runs on a daemon thread and can still be writing when
  the turn returns: title_generator tracks in-flight upgrade threads and
  oneshot joins them (bounded) before reading — no sleep, no eager read.

Fixes #112848. Direction shared with #112852 (@KoNit-K), which folded aux into
the headline counters; this keeps them backward compatible instead.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:03:25 -07:00
teknium1
30bde46510 fix(profiles): profile create never deletes a live marker-less dir; dangling symlink markers count as identity
Review follow-up on the identity predicate:

- create_profile: the rmtree guard is back to "tombstoned AND no identity
  marker". A live profiles/<name>/ without a marker is invisible to
  `profile list` but may still hold user files (skills/, memories/,
  cron/jobs.json) — main refused it with FileExistsError; the widened guard
  silently rmtree'd it. Now fails closed with an error naming the stray dir.
- named_profile_has_identity: `is_file()` follows symlinks, so a legacy
  profile whose only marker is a dangling symlinked config.yaml/.env
  (clone/migration leftover) vanished from list/serve/-p. A symlink is an
  identity claim: accept `is_symlink()` too.
- Docs: faq.md still said "each profile is just a directory under
  profiles/"; faq.md + user-guide/profiles.md + hermes_cli/AGENTS.md now
  state the identity-file contract.
2026-09-16 16:48:42 -07:00
teknium1
c055eee1fe refactor(checkpoints): slim the profile-rename rekey and name it as a checkpoint_manager sibling
Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):

- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
  repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
  idempotent rekey: every step overwrites and the old project metadata is removed last, so a
  mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
  A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
  192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
  (`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
  and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
  assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.

Fixes #112973
2026-09-16 14:34:22 -07:00
xielevi
a41552fad4 fix(profiles): purge a deleted profile's session/routing identity on delete
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:

- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
  `_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
  also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
  namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
  (`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
  in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
  writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
  `_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
  has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
  partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
  profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
  `create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
  delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
  leaves of a profile's conversation record is `delete_profile`'s business — it removes the
  profile's own home, `state.db` included.

Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
2026-09-16 14:21:14 -07:00
xielevi
81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
teknium1
3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1
08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1
f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1
23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1
abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
teknium1
e133f3f607 fix: device OAuth login scans every advertised authorization server
`hermes mcp login <server> --flow device` took `authorization_servers[0]`
from the protected-resource metadata and failed when that entry was a
browser-only or issuer-inconsistent server, even though a later entry was
the issuer-bound device_code server meant for headless clients (Higgsfield
advertises exactly this shape: a PKCE server first, the device server second).

Discovery now tries each advertised server in order and binds to the first
whose metadata issuer matches its advertised URL and that offers device
authorization. Issuer validation (RFC 8414 / SEP-2468) is unchanged per
server; a single-server resource raises exactly the error it raised before,
and a multi-server resource with no usable entry reports every attempt.

The browser path (`tools/mcp_oauth_manager.py` pre-flight) is deliberately
left on the SDK's own first-entry selection: the SDK's 401-branch discovery
re-selects `authorization_servers[0]` itself, so a divergent pre-flight pick
would only desynchronise the cached metadata from what the SDK authorizes against.
2026-09-15 19:00:42 -07:00
teknium1
43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
teknium1
da9810387d docs(config): scope the .env routing claim to registered env settings
Only names in OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS and the platform *_HOME_CHANNEL /
*_ALLOWED_USERS suffix family route to .env; other documented ALL-CAPS names still land in
config.yaml as top-level scalars with a notice. Say so instead of 'every documented
environment variable' (#111848 stays open for the remaining names).
2026-09-15 18:28:49 -07:00
teknium1
0e63a1bc5c fix(config): refuse an unknown path under a known section before writing
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.

Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.

`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.

Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:28:49 -07:00
teknium1
70e4938c07 fix(config): route every registered env setting through .env from config set/get/unset
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.

Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.

Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
2026-09-15 18:28:49 -07:00
teknium1
0e0692240a test: fold the named-source control into one clone-all test; document the skipped trees
Trim the three contributor tests from #101340 to two invariants: the
root-only ignore test now also asserts that a named profile used as
source keeps its own models/ (the exclusion is gated on the default
root), and the end-to-end create_profile(clone_all=True) test stays.
Same assertions, two tests — the salvage bar.

Docs: profile-commands.md and profiles.md list models/, runtimes/ and
node/ among the --clone-all exclusions so users know why a clone from
the default profile does not carry the local-model weights.
2026-09-15 18:26:38 -07:00
teknium1
d1bd778a5e fix(checkpoints): clear-legacy trims to two real-path tests, aligns failure line with the log
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.

Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:24:22 -07:00
teknium1
fa915784cb fix: merge shipped skill categories instead of replacing them on update
The per-root merge in _copy_dist_payload treated the first level under
skills/ as the replace unit, but skills live at skills/<category>/<skill>.
A distribution shipping one skill inside a category rmtree'd the whole
category directory, wiping every skill the user (or hermes update) had
placed there — the exact scenario of issue #25120 that this PR claims to
fix. A shipped directory with no files of its own is now a container:
its children are merged one level down, and the replace unit becomes the
nearest directory that holds a file (a skill dir always holds SKILL.md).

A symlinked category (skills/devops -> shared dir) is a container too,
so the pre-write symlink check now covers it instead of silently
unlinking the link and dropping the shipped copy in its place.

Review finding: categorised skill payload rmtree'd skills/<category>/,
destroying sibling user and bundled skills; symlinked category was
unlinked rather than refused.
2026-09-15 06:22:15 -07:00
teknium1
ab7b97f55e docs(distribution): describe per-root merge of skills/ and cron/ on update
The user guide, command reference, `hermes profile update --help` and the
install-over-plain-profile warning all said skills/ and cron/ are overwritten
wholesale; they are now merged per skill / cron job, so say so, and document the
symlinked-container refusal.
2026-09-15 06:22:15 -07:00
teknium1
a9a8a3fa2e fix(config): hermes config get masks credentials on every path; --raw opts out
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.

`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.

Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).

Fixes #110758
Fixes #84106
2026-09-15 05:08:55 -07:00
teknium1
7bb52c0b74 feat: profile clone can opt into staying synced with its source (--sync-imports)
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.

`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
2026-09-15 04:44:09 -07:00
teknium1
5083d5f78e fix(gateway): honour allow_all_users from config.yaml by bridging it to GATEWAY_ALLOW_ALL_USERS
`gateway.allow_all_users: true` (and the top-level spelling) in config.yaml
was a silent no-op: `_TOPLEVEL_BRIDGE` never forwarded it, GatewayConfig has
no field, and every allow-all reader (authz mixin default-deny branch,
startup access check, own-policy adapters, Discord/Matrix/Email plugin
gates) consults the GATEWAY_ALLOW_ALL_USERS env var only.

Bridge the YAML key into that env var in `bridge_core_env_settings`, the
one seam every reader already shares, instead of threading a new config
attribute through ten readers:
- first-writer-wins: an explicit env var beats YAML (matches every other
  {PLATFORM}_* gate);
- skipped inside a multiplexed secondary profile's scope (#80099 class):
  the secondary's config.yaml must not become the default profile's policy,
  and the isolation test now asserts GATEWAY_ALLOW_ALL_USERS stays unset;
- a startup warning names config.yaml as the grant source, because the key
  was inert until now and a forgotten `true` flips the posture to open.

Tests: both spellings authorize a stranger through `_is_user_authorized`;
`false`, absent key, and env=false over YAML=true all stay denied.
Docs: security guide, env-var reference, gateway internals.

Fixes #110690
2026-09-15 04:37:41 -07:00
kshitijk4poor
569b4242a3 fix(auth): a Portal-returned inference host is accepted only when the operator named it
#111809 accepted any *.nousresearch.com https host once the operator's Portal override
pointed at a non-production Portal. That let a network-provenance value — the Portal's
refresh response — pick any Nous-owned host as the bearer recipient, including hosts that
are not inference gateways. Owning the DNS suffix is not the same as being an authorized
recipient, and the validator's threat model (an injected refresh response) is exactly the
case a suffix rule fails to bound.

The recipient is now the operator's own NOUS_INFERENCE_BASE_URL: a non-production host
returned by the Portal is accepted exactly when it equals that override's host, otherwise
the strict production set stands. The Portal override grants nothing by itself. What the
operator gains over plain use of the override is that the Portal's value is then persisted
and used for the pricing scope and proxy instead of being healed to production, and the
per-turn "refusing inference URL host" warning stops. No environment is named in code.

Raised on #111809 review. Tests: recipient match accepted, unrelated Nous host refused, no
override refused, Portal override alone grants nothing, the match follows the profile scope
under multiplexing; the widening cases are red on main. Docs row for NOUS_INFERENCE_BASE_URL.

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-09-15 16:55:43 +05:30
Teknium
398279fd6a feat(skills): add ip-as-logo optional skill (minimal cute IP mascot marks)
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.

Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
  (square aspect, main-prompt constraints mode — no negative_prompt
  parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
  own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
  proposal-round skip for pre-authorized batches, dimensions-reporting
  rule when the backend returns only a URL, limbless-subject note

Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.

Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).

Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
2026-09-15 04:20:15 -07:00
teknium1
b531023622 fix(gateway): drop the redundant whole-block lease; one invariant test per atom; document the watchdog env vars
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.

Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.

Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
2026-09-15 04:19:19 -07:00