Background process/delegation notifications re-enter the conversation as
role=user turns. Two failure modes (#52694):
- The injected MessageEvent reused evt.message_id — the id of the user
message that ARMED the watch, hours stale by delivery — as its reply
anchor, so the gateway posted the system notice as a reply to an old
user message (Discord reports). Post fresh instead; topic routing
stays intact via source.thread_id.
- The model-facing text carried no machine-provenance marker, so the
model read the notice as something the human said. Append an explicit
[INTERNAL NOTIFICATION — not a user message] footer (trailing, so
start-anchored consumers keep matching) and set
metadata.notification_origin=process_registry_synthetic.
Fixes#52694
Background-process completions and watch events re-enter the gateway as
synthetic MessageEvent(internal=True) turns carrying the id of the message
that STARTED the process (captured at spawn time via
HERMES_SESSION_MESSAGE_ID, persisted as watcher_message_id, replayed on the
queued event). By delivery time that message is old news: the user has
typically continued elsewhere, and _reply_anchor_for_event made the finished
job's reply quote it on every reply-anchoring platform — on Discord the
completion visibly answered a stale DM message from a different topic
(#52694).
The synthetic event is not a reply to the trigger message. Drop the id from
the event and from the restored origin source (the same strip
run_goals._synthetic_prompt_event applies to goal/loop prompts), keep it in
event metadata as original_trigger_message_id for debugging, and leave
routing untouched: thread/topic lanes carry thread_id and the anchor-less
synthetic-send branches are already covered (#87051).
Fixes#52694
Co-authored-by: Hermes Agent <agent@hermes.local>
Pending async-delegation completions were replayed only at process start
(restore_undelivered_completions from ProcessRegistry.__init__ and the gateway's
secondary-profile restore). A result persisted by a process that then died (a
desktop reload) sat at delivery_state='pending' until some process restarted.
The gateway async-delegation watcher and the TUI notification poller now sweep
each profile home they serve every 30s: abandoned in-flight rows are classified
by recover_abandoned_delegations, and terminal pending rows whose owner fails
the shared start-time liveness check, idle for 60s and not under a live claim,
are offered once to this process's completion queue. Delivery still goes
through claim_completion_delivery, so two processes offering the same row never
both deliver it. Rows past _MAX_DELIVERY_ATTEMPTS or the replay age converge to
dropped.
Fixes#97202
A plugin registered on transform_terminal_output only ever saw foreground
`terminal` output: tools/terminal_tool_result.py::_apply_output_transform_hook
runs from finalize_foreground_result and nowhere else. Background output
reached the model through a different seam — process_manage poll/wait/log/kill
results, `list` previews and the completion/heartbeat/watch notifications all
pass through tools/process_registry.py::_redact_process_result — which redacted
but never transformed, so a fleet redaction or summarising plugin silently did
nothing for backgrounded commands.
Apply the same hook helper at that shared seam (one new
transform_process_output wrapper) and at the two gateway agent-notify sites
that read session.output_buffer directly. The order matches the foreground
path and teknium1's review note on #71401: hook first, redaction after, so a
replacement the plugin returns is still masked. returncode is None while the
process runs; env_type is not recorded per process and is passed empty.
Not changed: the spawn acknowledgement ("Background process started") carries
no command output, and the persistent local shell already goes through
_run_foreground and was transformed — the issue's reading of that branch was
wrong; the real gap was the process_manage/notification seam.
Fixes#70760
Slim redo of #71401 (Christopher-Schulze): same seam and ordering, without
the ANSI-stripping relocation and render helper.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
(cherry picked from commit 84661de52f78ccb84054e1ded92b07166e40546d)
An /update run leaves a marker naming the chat to notify. When that platform's
adapter is not connected as the update finishes, _send_update_notification
defers and keeps the markers for a later retry — correct for the restart that
/update itself triggers, where the adapter reconnects seconds later.
Nothing bounds that wait. If the platform is not configured at all, no adapter
will ever appear and the defer never resolves. The startup path reschedules the
watcher whenever the markers are still on disk, so the marker outlives every
restart: it re-logs "adapter not connected yet" each poll_interval (2s), in
every process, indefinitely. A marker written months ago was still doing this
on one install, filling gateway.log with ~19k duplicate lines.
Bound the wait using the timestamp the marker already carries. Once the marker
is older than the cap the notice is undeliverable by any retry, so log it once
at WARNING, clear the markers, and return True — the definitive answer that
stops the caller rescheduling. Markers with no timestamp (written before that
field existed) keep the previous retry behavior rather than guess an age.
Tests cover the three behaviors: a stale marker is dropped and clears every
marker file, a recent marker is still held for the reconnecting adapter, and a
marker without a timestamp still retries.
Host-scoped update-restart obligation (061195fac1 / 953b6f6f08 / 3be255eca6)
lands on the PM model: the obligation record, its readers and the legacy
per-home marker compat come in as-is. The catch-up restart path
(`_apply_pending_fleet_restart_catchup` / `_run_pending_fleet_restart`) is
retired here (the fleet restart rides the completion owner), so main's
per-host restart-once guard on that path is not carried; its unit→live
MainPID collapse IS ported into the live post-update systemd pass
(`_restart_systemd_gateway_units`), with the two collapse tests rewritten
against that function (red on the pre-port tree: `_unit_main_pid` absent).
Tests that exercised only the retired catch-up path are dropped.
Desktop: main's shared log-rotation planner replaces the inline constants
in main.ts; the merge keeps our machine-profile import beside its import.
utf-8 → utf-8-sig on the three new BOM-intolerant reads (footguns lint).
Review findings on the host-scoped update→restart obligation.
- update_cmd_fleet: an unwritable host state dir (read-only HERMES_GATEWAY_LOCK_DIR,
container UID that does not own $HOME) made write_host_obligation return False and the
caller ignored it, so an interrupted update left ZERO obligation — stale code, no
warning, no catch-up restart. The return is now propagated: the legacy per-home marker
(still read by every reader) carries it, and a host that can write neither says so.
- update_cmd_fleet::_obligation_fields: a PRESENT but unparseable/foreign-versioned host
record no longer falls through to the legacy marker; terms nobody can read cannot be
discharged by another record's terms.
- update_cmd_fleet::_restart_identity_sha: zip/pip/Docker installs resolve no checkout
SHA, so the restart-once stamp was "" and could never match — every profile's update
re-killed the one shared multiplexer. Falls back to the record's expected_sha, then the
receipt's post-update identity.
- update_host_obligation: any main_pid probe error is unproven identity (keep its own
restart), never an aborted restart pass.
- run_notifications: the online notice dedupes per home CHAT, so two served profiles
sharing one chat get one message (accounting stays per profile); transport resolution is
isolated per profile, so one broken adapter no longer starves the rest of the fan-out.
- run_adapters / run_profile_reconcile: _profile_configs is pruned with the served set, so
a failed or removed profile no longer owes a notice nothing can deliver and
.restart_pending.json is unlinked.
Tests cover the new format's own hazards: unwritable record dir, foreign-version record
with a legacy marker present, non-git install, shared home chat, broken adapter, pruned
config, plus a parity test for the duplicated host-state-dir resolver.
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
main folded the auto-archive housekeeping into per-profile state.db maintenance; the plugin
update-check chore stays. The consolidated-away TestWorkdirParallelPool stays removed. Two new
utf-8 reads in gateway/host_rendezvous.py read utf-8-sig.
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:
- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
now enter the SESSION's profile scope, the same chokepoint _finalize_session
already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
profile's scope (mirrors startup discovery and the reconcile chore), with a
trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
eviction and state/memory handle release now run inside the deleted
profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
profile no longer means "no scope". launch_profile_scope_if_multiplexed()
binds the launch profile once the process multiplexes; before activation it
is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
assignee's secret AND terminal scope for toolset resolution and the spawn-env
build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
when the host has more than one servable profile home, instead of lazily on
the first ?profile= request after earlier work already ran unscoped.
Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The queued-follow-up lane treated "the send call did not raise" as "the text reached the
chat". `_deliver_queued_first_response` returns normally when the adapter reports
`SendResult(success=False)` (flood control, retries exhausted): `_send_queued_final_text`
only logs the failure. The result was then marked `already_sent`, the completion path
suppressed the only remaining send, and the user got nothing at all — the opposite failure
to the duplicate this lane fixes, on the same code path.
`_deliver_queued_first_response` now reports whether the text is in the chat, and the mark is
gated on it. A refused send returns False (and skips the attachment upload, so the completion
send replays text and files together); a connector egress DECLINE still returns True, because
that destination is not approved and must not be re-sent.
Attachments were also uploaded twice on the single-delivery path: the queued lane uploads the
response's MEDIA: files, then the completion path's `already_sent` rescan uploaded them again.
The result now records `media_already_delivered` and the rescan skips them.
Live (real gateway, real IRC adapter, stand-in server/model, temp home): follow-up refused by
the reference guard with the first send refused — pre-fix head delivered nothing, this head
delivers the answer once.
The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).
Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
After a restart the routing index rebuilt every lane from `SessionEntry.origin`,
which carries the runtime profile (key namespace) but not the bot that received
the conversation. Delivery then fell to `_is_shared_bot_satellite`: a lane owned
by a secondary bot whose runtime profile is ALSO a satellite of the default bot
was handed to the default bot, and authorization read the wrong allowlist.
- `SessionEntry.transport_profile` (routing JSON) + nullable
`sessions.transport_profile` (SCHEMA_SQL, reconciled by the existing column
path; `agent:main` keys untouched, standalone gateways write nothing). Stamped
from the pinned `RoutingIdentity` at create, reset/switch, DB recovery and
every peer refresh; compression forks inherit it like the other routing columns.
- `session_identity.restore_identity()` re-pins a `RoutingIdentity(transport=None)`
from the persisted transport profile; `authz_mixin._restored_source(entry)` is
the one seam every revive path uses (auto-resume, heartbeat restore, plugin
injection, background-process events).
- `_adapter_for_source` / `_adapter_profile_for_source` honour a restored identity:
the persisted bot's adapter or None — never the default bot by heuristic.
Entries written before the column exist keep the old chain.
Phase 5 of #88715.
Mechanical migration of the `_adapter_for_source` call sites in this sibling. Intake policy sites take `_intake_adapter_for`; every send/edit/typing/pending-slot site takes `_delivery_adapter_for`. Part of #88715 (phase 4).
A served (multiplexed) profile's api_server turn binds the raw session id
as its session key, so its watch/completion event names no profile and the
wake path self-posted it unprefixed with the primary key — resuming the
session in the DEFAULT profile's store — while an event whose source did
name a route-only served profile resolved no adapter and was deferred
forever.
_self_post_api_server now proves ownership through the served profile's
own session store (the rung the Kanban notifier already applies), scans
served stores when the raw event carries no hint, runs the wake under that
profile's scope via deliver_wake(profile=...), and fails closed for a hinted
profile that does not own the session. The ownership helper moves to
gateway/wake.py so both callers share one function.
The non-compression branch of _resolve_async_delegation_session awaited the
spawning-session row lookup and then unconditionally called switch_session(),
so a run invalidated (/stop) or a route replaced (/new, /resume) while the
lookup was pending was overwritten by the stale completion's pin.
- Snapshot the routing key's run generation before the await and re-check it
before mutating the route; invalidated -> drop the injection, route untouched.
- switch_session(expected_session_id=...) turns the /resume primitive into a
CAS for this caller, mirroring advance_compression_session: a route that
moved past the snapshot wins over the late completion.
- _current_session_run_generation extracted from _is_session_run_current.
Slimmer redo of #113692 (@KoNit-K) and #113716 (@wangtaotaotao95): same
direction, without the second routing-authority lock and authorize callbacks.
Fixes#113690
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: wangtaotaotao95 <wangtaotaotao95@users.noreply.github.com>
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
gateway/run_notifications.py::_send_session_db_warning_notifications already
computed profile_arg but left `hermes doctor` (fts_index branch) and
`hermes gateway restart` (default branch) bare; agent/turn_finalizer.py had a
bare `hermes doctor` in the error fallback used when the explainer produced
no text. Same class as the previous commit. Also reflows the explainer
strings so continuation lines break at clause boundaries.
`terminal(background=true, notify_on_complete=true)` appended its watcher descriptor to
`process_registry.pending_watchers`, which only the post-turn hooks drain. A process that
finished while the turn that launched it was still running (an agent sleep-polling for
hours) had no watcher task at all: the completion_queue entry sat inert, nothing was
injected, and the chat stayed mute until that turn ended (#112033).
- `_register_completion_watcher` arms the watcher on the live gateway loop at registration
(`GatewayRunner.arm_process_watcher`, via the existing `_gateway_runner_ref` /
`_gateway_loop` seam that send_message and cron already use); `pending_watchers` stays
the fallback while the gateway is not serving (checkpoint recovery at startup, shutdown).
- The agent-notify branch of `_run_process_watcher` keeps its design (the agent's next turn
is the user-facing report) but, when the launching turn is still active at process exit,
the injection only queues a follow-up — so the concise receipt is sent to the chat right
away instead of never. The busy check is taken before injection because the injected turn
itself installs the adapter's session guard.
Live probe (real process, real GatewayRunner loop, fake telegram adapter, busy session):
before — pending_watchers=1 after exit, 0 watcher tasks, 0 injections, 0 receipts;
after — pending_watchers=0, watcher task armed at launch, 1 injection, 1 concise receipt.
Control (idle session): 1 injection, 0 receipts, unchanged.
Slimmer redo of #112038 by @KoNit-K: same two gaps closed, without a second scheduler
registry / loop attribute on ProcessRegistry and GatewayRunner.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Slim the salvaged #112111 mechanism (issue #112109) while keeping its
behaviour: the boot pass and the reconnect hook both run
_replay_pending_planned_restart_notification, which sends to every home
channel still owed an online notice, records delivered targets in
.restart_pending.json and unlinks the marker only when every owed target
(configured home with gateway_restart_notification=true) has been reached.
Dropped from the contributor diff:
- the per-target on_delivered checkpoint callback and pending_targets
field: delivered targets are written once after the pass. Residual is a
benign duplicate notice only if the process dies mid send-loop.
- getattr-based lazy lock -> class attribute default, same idiom as
run_profile_reconcile._reconcile_lock.
- _clear_planned_restart_notification in gateway/run.py: no production
caller remained; the roundtrip test unlinks the path directly.
- tests trimmed to two invariants: offline-at-boot is replayed once on
reconnect (with live-at-boot control), and partial delivery is persisted
so a fresh process does not re-notify and an opted-out home never keeps
the marker alive.
Live probe (temp HERMES_HOME, Discord home, adapter absent at boot then
reconnected): base consumed the marker with 0 sends; fixed head retains it
and sends the online notice exactly once on reconnect, then clears it.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.
Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
The raw-output watcher modes (all/result/error) and the interim running
update sent the bracketed debug wrapper with the internal process id
(`[Background process proc_… finished with exit code N~ Here's the final
output: …]`) to Telegram/Discord/Slack chats. Reuse the concise one-line
status header for every mode and append the bounded, ANSI-stripped output
tail in a code block; the running update gets the same shape.
Salvaged from #54266 (rebased onto the post-#102117 run_notifications
sibling; the concise mode had landed in between, so the header is shared
rather than reimplemented). Also covers #13122 (ANSI stripping).
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:
- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
timeout were bare os.getenv → the default profile's values; a job without a
model silently ran on the default's HERMES_MODEL instead of refusing.
cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
kanban worker inherited the launch profile's non-credential .env settings
and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
HERMES_MODEL) — X's worker ran in the default's docker image on the
default's model. tools/environments/local.py::strip_launch_profile_env drops
them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
assignee (toolset probes call get_secret without a scope → swallowed
UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
writing the completed tick into the DEFAULT profile's state.db and leaving
X's row awaiting_response forever; the --until judge ran with the default's
aux credentials.
- background processes: a secondary's processes.json (scope-relative since
adf23550f5) was never read at startup; its processes were not re-adopted
and notify_on_complete notices were lost. Startup recovers every served
home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
drain for the ambient profile (default's mode for everyone; X's `off`
dropped a sibling's `all` event), recovered watchers used the default's
mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
_deliver_platform_notice used the default's GatewayConfig so a secondary's
notice_delivery: private went public.
Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
- gateway/platforms/yuanbao.py::AutoSetHomeMiddleware: the first authorized DM
to a SECONDARY Yuanbao bot wrote YUANBAO_HOME_CHANNEL into os.environ, making
that tenant's chat the default profile's cron/notification home. The write
now only happens unscoped; reads go through the scoped reader + config.
- tools/env_passthrough.py::_config_passthrough: one module slot froze the
first profile's terminal.env_passthrough for every profile's sandbox children;
keyed by hermes_home_key().
- gateway/run.py::_slack_ignored_channels_from_gateway_config: the runner-level
fail-safe only had the DEFAULT profile's GatewayConfig, so a secondary Slack
bot's traffic was judged by the default's ignored list. It now takes the
source's routed adapter (whose extra is the secondary's own config) and reads
the env fallback through the scoped gate reader.
The supervised `_async_delegation_watcher` and startup-recovered process
watchers run under the ROOT scope, so a secondary profile's completion was
classified against the DEFAULT profile's state.db (row absent → "terminal" →
"permanently-gone session" warning, delivery dropped) and every durable-ledger
op (`claim`/`complete`/`release`) hit the default's ledger, stranding the real
row `pending` forever.
`_deliver_completion_notification` and `_deliver_async_delegation_group` now
run their whole pre-flight + claim + inject + settle sequence inside
`_completion_event_scope(evt)` — the runtime scope of the profile the event's
session belongs to. At multiplex startup `_restore_secondary_completion_ledgers`
replays every secondary's pending rows, which the process registry (launch
home only) never saw.
Builds on the classify-only half of #107247.
Inside a routed satellite's turn the ambient scope is the satellite's, whose
.env has no token or allowlist. Five sites still called `_is_user_authorized()`
directly there — `/topic`, the sibling-thread `/stop` grant, plugin message
injection, Discord voice transcripts and startup auto-resume — so the shared
bot's owner was refused ("not authorized to use /topic") and a satellite that
DID copy an allowlist widened who may drive those commands. They now go through
`_is_user_authorized_for_source`, and `_under_authorization_profile` derives the
transport home from the delivering adapter's owner when ingress did not stamp
one (restored/cached sources).
`_adapters_for_profile` (the resolver behind `_authorization_adapter`,
`_adapter_for_source` and now `_resolve_injection_adapter`) returns the primary
map for a shared-bot satellite: a served profile with the `{}` startup
placeholder, no reconnect pending, targeted by a default-bot route. Kanban and
cron already applied that rule; the gateway's own resolvers returned None, so
heartbeats, process completions, goal notices and delegation results for such
profiles were undeliverable after a restart. A secondary that owns a credential
(connected on any platform, or queued for reconnect) still fails closed.
_parse_session_key() only accepted the agent:main namespace, so under
multiplex_profiles named-profile keys (agent:<profile>:...) parsed to None and
notification/shutdown routing fell back to the LRU source cache; and
_sibling_thread_run_keys() hardcoded agent:main, so a per-user thread /stop
never matched a sibling run on a named profile. Accept any valid profile-id
namespace slot and report it as profile (main keys keep their exact historical
shape), build the sibling prefix from _session_key_namespace(source.profile),
and drop the manual agent:main re-wrap in _build_process_event_source().
(cherry picked from commit 562792015a60da3c675498edc2552a6234015457)
The recovery commands rendered on structural corruption — the turn explainer's
`session_persistence_failed`/corrupt body, the gateway's home-channel state.db
warning, and hermes_state_repair._persistent_repair_exhausted_error — already
interpolate the active profile's state.db path, but every `hermes ...` verb in them
was bare. A bare `hermes` follows the sticky `active_profile` file, so an operator
running the pasted `hermes doctor --fix` (or `hermes sessions recover` with a
relative source) from a named-profile incident could inspect or repair a different
profile's database (#105887).
hermes_constants.profile_cli_selector() renders `-p <name> ` for a named profile
home (default home and custom roots outside the profile tree render nothing: the
default is what a bare `hermes` already means, and a custom root is only reachable
via HERMES_HOME). Every command in the three guidance sites now carries it, and
the new `fts_index` guidance inherits the same interpolation.
Live check with HERMES_HOME=<root>/profiles/research and active_profile=other:
before `1. Run \`hermes doctor --fix\`` (targets "other"); after
`1. Run \`hermes -p research doctor --fix\`` and `hermes -p research sessions
recover --source <root>/profiles/research/state.db --inspect-only`.
Refs #105887
Reported-by: Cuttingwater
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.
One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".
Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).
Fixes#97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
_send_session_db_warning_notifications() broadcast the error recorded at startup
without asking whether it was still true. A startup `database is locked` routinely
clears while the adapters are still connecting (another profile's open, a `hermes
sessions` one-shot, a slow SMB lock release), so the home channels were told the
store was unavailable when it had already healed — and a warning that is wrong once
is ignored the next time it is right.
The RecoverableHandleCache opener already clears `_session_db_init_error` on
recovery; the broadcast now drives one open attempt through it first and only warns
when the error is still standing. A store that is genuinely still down warns exactly
as before.
Fixes#108031.
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved
its live adapter by bare platform from runner.adapters — the DEFAULT profile's map —
so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy,
every plugin platform), its "Gateway shutting down/restarted" and /update notices,
its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left
through the default bot (Telegram DMs landed in the user's chat with the other bot).
Every such door now resolves through the profile-aware, fail-closed resolver already
used by the inbound reply path (authz_mixin: _adapters_for_profile /
_authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear
error, never the default bot.
- tools/send_message_senders.py::_live_adapter — resolve via
runner._authorization_adapter(platform, get_active_profile_name()); shared by
_send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and
the WeCom standalone sender.
- gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay-
aware resolve_delivery_transport callers); _authorization_adapter reuses it.
- gateway/run_shutdown.py — shutdown/restart notice for a running session uses the
session's source transport / agent:<profile>: key lane, never self.adapters.
- gateway/slash_commands.py + run_notifications.py — /restart and /update markers
persist `profile`; the restart notice, update result and update prompt resolve the
requester's own adapter (legacy markers fall back to the session_key lane).
/goal, /heartbeat, /approve, /deny confirmations use the source's own transport.
- gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its
route; the wakeup watcher scans every served profile's store under its own scope
(same shape as _handoff_watcher) and fires through that profile's adapter map.
- plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in
the owning profile (its adapters and its home channels).
- plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter.
- docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity).
Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py,
tests/gateway/test_multiplex_notice_egress_profile_adapter.py.
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract
The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:
- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
covers auxiliary work) and skips Nous for vision, which the welcome model does not take.
- The structured 429 body was never read. The classifier now parses `reason` /
`retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
are rate limits that honour `retry_after` and never rotate the free tier's only credential.
The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
fall back instead of retrying or re-exchanging. The terminal paths say what happened and
name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).
- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
beside the rate-limit and credits headers; the next call moves the session, and the config
default when it still names `nous/welcome`, to the backing model the gateway named.
- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
outside the host allowlist, routing defaulted to inference-api, where every request is a
400. A guest now defaults to the welcome literal at the exchange, in the shared store's
shape, and in effective routing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)
* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use
A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.
nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)
* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit
Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.
The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)
* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use
`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).
The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.
Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.
(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)
* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent
When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.
`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.
(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)
* fix(auth): the free tier outranks implicit host credentials in provider resolution
On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.
The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.
Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.
Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.
(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)
* fix(auth): review follow-ups for the free-tier rung (NS-829)
- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
the free tier off; its contract is the boto chain, and the free tier now
sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
before consulting the resolver, so a gateway boot on a machine with AWS
credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
six (parametrized ladder cases; a failed mint that returns None or raises
falls through to Bedrock).
scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.
(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)
* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone
The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.
`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.
`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.
Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).
Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.
Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.
* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read
Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.
Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.
`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.
Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.
`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).
Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.
Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.
Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.
* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"
A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.
Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.
Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.
* fix(copy): free-tier text stops promising a connector transfer and never names the config key
Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.
The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).
The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.
zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.
* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn
The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.
`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.
Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.
The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).
Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.
* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll
The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.
`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.
`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.
Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.
* fix(cli): the banner names the free tier's model instead of "no model configured"
The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.
`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").
Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.
* fix(aux): vision on the free tier uses nous/welcome too
The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)
* fix(gateway): hermes gateway run is a boot owner of the free tier too
Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).
GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.
Live, real GatewayRunner.start against a fake portal in a fresh home:
gate on -> 1 create, identity persisted, resolve_runtime_provider=nous,
/login precondition sees the identity
gate off -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.
Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.
---------
Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>