reply_expected was read from the event that opened the turn, so an addressed
message the same turn ended up answering could still end on a bare silence
marker and vanish:
- a queued chain paired the terminal turn's display kind with the opener's
flag; the recursive run now receives the pending event's flag, persists
it, and returns it as queued_terminal_reply_expected beside
queued_terminal_display_kind for the outer shaping to read;
- pending-message merges (merge_pending_message_event, text batching, busy
debounce) and an active-turn redirect folded the new message in but kept
the old flag; MessageEvent.absorb_reply_expected now folds it: an
addressed message wins, then an unknown one.
Also: reply_expected is persisted only when the adapter set it, so rows on
other platforms carry no null key; the turn_params pop and getattr reads
go (turn_params already flows into TurnContext); silence_allowed is a
module import; crash recovery keeps the diagnostic-mute check on machinery
turns and applies silence_allowed only to the silence verdict; the DEBUG
line no longer fires for machinery turns.
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.
`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.
Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.
(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
Review follow-ups on the crash-left reply adoption:
- A persisted reply is now judged the way live delivery would have judged
it. A bare silence marker ([SILENT] / SILENT / NO_REPLY ...) on an
internal turn, or the reply to a diagnostic wake whose chat policy mutes
diagnostics (read in the routed profile's scope, as the adapter does), is
owed nothing: the marker is cleared and nothing is sent or resumed. A
human turn's bare silence marker becomes the same "returned only a silence
marker" notice the live path sends. Before, the raw "NO_REPLY" reached the
user as a "Recovered reply".
- The active-turn marker start is written as aware UTC and compared as
epoch seconds (startup adoption and recover_interrupted_turns). A naive
local wall clock read by a process in another zone (DST, container vs
unit TZ) was hours off: a fresh in-flight turn was dropped as stale, or a
previous turn's reply could be adopted as this one's. updated_at stays
naive local for the older binary's recency heuristic; a pre-upgrade naive
marker still reads as local time.
- Reply timestamps go through coerce_epoch instead of float(), so one odd
transcript row cannot abort the whole recovery pass.
On an unclean start, suspend_recently_active(120) marked every session
touched in the last 120 s resume_pending/restart_interrupted, so startup
auto-resume ran a fresh model turn for chats whose turn had already
finished and been delivered: one kill re-answered 52 chats in the C12
delivery suite. The durable active-turn markers already name the exact
in-flight turns, so the recency sweep is removed.
That sweep also hid a real window: _handle_message cleared the turn
marker in its finally BEFORE the adapter recorded the delivery
obligation, so a kill in between left neither marker nor ledger row and
the persisted reply was never sent. The adapter now owns the marker for
turns it delivers and clears it right after record_delivery_obligation
(or once nothing more is owed). At unclean startup a marked turn whose
final reply is already in the transcript has that reply adopted into the
delivery ledger (unowned, 'attempting': sent once, marked as a possible
duplicate) instead of being regenerated; a marked turn with no reply
resumes once, as before.
The boot sweep now moves every claimed row to attempting. If the platform went
fatal between the claim and the send (restart notification, flood sleep),
_obligation_adapter skipped the row without releasing it, so it sat in
attempting owned by this live process: the reconnect sweep only takes failed
rows and the boot sweep skips live owners, so the reply waited for the next
restart (and then carried a false duplicate marker). Release any undispatched
claim, not only runtime ones.
Found by independent review of #120450.
Two invariants on the boot outbox sweep, both red on the previous ledger:
- a pending row whose boot redelivery was interrupted after the platform
accepted it (or that an older build claimed) is redelivered with the
recovered marker, never a second plain copy;
- a boot-claimed failed row is not re-claimable by the runtime reconnect
sweep while the boot send is in flight (it used to stay 'failed' and
owned by this process, so a reconnect could send it a second time).
Also corrects the startup-claim comment in _obligation_adapter and the
messaging docs bullet on mid-send recovery.
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:
1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
with a per-plugin activation summary (hermes_cli/plugins_activation.py):
activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
plugins.manage install/toggle/update, dashboard REST install, tool-triggered
force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
discover_plugins() short-circuits on _discovered, which is why reload.mcp after
a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
factories not yet wired on the live native client (keyed (plugin, qualname);
a force reload hands back new function objects). Telegram hoists late handlers
ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
the first match per group) and re-wires on the transient-init rebuild; Slack
dedupes register_slack_action_handler per AsyncApp. The gateway runner
subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
hint and plugins.manage results (activation, gateway_reloaded,
restart_required only when no gateway answered) say exactly that.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The line did int(os.getenv("HERMES_MAX_ITERATIONS", "500")): agent.max_turns: unlimited
made int() raise inside suppress() so no line was logged at all, and the invented 500
default never matched what the turn loop enforces. Resolve through resolve_turn_limit
like _current_max_iterations does. The "bypass" half of #116888 is already config-first:
_bridge_max_turns_to_env writes agent.max_turns over the env slot unconditionally (PR
#18413) and routed profiles read their own config (819517fbac).
Fixes#116888
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
write_runtime_status(wait_timeout=, reload_existing=) are part of the
call contract, not private state: drop the leading underscores at the
definition and every call site. flush_runtime_status_async no longer
polls the writer every 20 ms; a daemon waiter parked on the writer's
Condition relays the settled generation into a loop future, so the loop
wakes on the event and a wedged write past the timeout cannot hold
interpreter exit.
The cherry-pick flushed the writer at each of the ten return points in
GatewayRunner.start(). Move the body to _start_impl() and flush once in a
finally in start(), so every path (early aborts included) bounds startup
on the latest snapshot without repeating the call.
An API-server-only (or webhook/local-only) gateway logged "No env user
allowlists configured. Messaging platforms default to pairing/allowlist
policies..." on every boot even though no messaging adapter exists to
gate, producing unactionable monitoring noise (#115439).
_start_check_access_policy now emits the reminder only when at least one
enabled platform outside {LOCAL, API_SERVER, WEBHOOK} is configured —
the same non-messaging set run_shutdown already uses. The open-policy
refusal check that follows is unchanged.
Supersedes #115449 (@whyyagswhy) whose enabled_platform_count gate still
counts api_server, and #115577 (@Baophan00) whose config-based gate this
follows in a slimmer shape.
_wait_bounded_or_release logged its timeout warning synchronously on the
asyncio loop; a contended logging.Handler.lock parked the loop long enough
for the loop-liveness watchdog to exit 75. Dispatch it via
asyncio.to_thread, matching #111498.
Fixes#115516
Extend the contributor fix (#116516) so _startup_restore_in_progress is
cleared even when the bounded resume wait or the warm-up await raises,
not only the drain: any raise before the flag flips leaves every
non-internal inbound queued forever (run_inbound reads the flag first).
_drain_startup_restore_queue awaited adapter.handle_message with no
containment, and _finish_startup_restore cleared
_startup_restore_in_progress only after the drain returned. A single
raising replay propagated out of both, leaving the flag stuck True —
and since run_inbound consults it before anything else, every
non-internal inbound message queued into _startup_restore_queue forever.
The events behind the raising one were never replayed either.
Wrap the per-event body (adapter resolution plus the replay await) so a
bad event logs and the drain continues, and move the gate release into a
finally so it opens even if the drain itself raises. Non-adapter raises
still propagate — those are internal bugs and should surface.
Regression tests cover a raising first replay (gate opens, remaining
queue drains, failed replay counted out) and a post-drain inbound that
dispatches instead of re-queueing.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
After a restart the routing index rebuilt every lane from `SessionEntry.origin`,
which carries the runtime profile (key namespace) but not the bot that received
the conversation. Delivery then fell to `_is_shared_bot_satellite`: a lane owned
by a secondary bot whose runtime profile is ALSO a satellite of the default bot
was handed to the default bot, and authorization read the wrong allowlist.
- `SessionEntry.transport_profile` (routing JSON) + nullable
`sessions.transport_profile` (SCHEMA_SQL, reconciled by the existing column
path; `agent:main` keys untouched, standalone gateways write nothing). Stamped
from the pinned `RoutingIdentity` at create, reset/switch, DB recovery and
every peer refresh; compression forks inherit it like the other routing columns.
- `session_identity.restore_identity()` re-pins a `RoutingIdentity(transport=None)`
from the persisted transport profile; `authz_mixin._restored_source(entry)` is
the one seam every revive path uses (auto-resume, heartbeat restore, plugin
injection, background-process events).
- `_adapter_for_source` / `_adapter_profile_for_source` honour a restored identity:
the persisted bot's adapter or None — never the default bot by heuristic.
Entries written before the column exist keep the old chain.
Phase 5 of #88715.
Mechanical migration of the `_adapter_for_source` call sites in this sibling. Intake policy sites take `_intake_adapter_for`; every send/edit/typing/pending-slot site takes `_delivery_adapter_for`. Part of #88715 (phase 4).
Gate review: `degraded` became a whole-life serving state but four
sibling predicates only accepted `running`: derive_gateway_busy /
derive_gateway_drainable (the NAS drain gate reported a degraded gateway
with in-flight turns as idle and undrainable), the stale-heartbeat
detectors in `hermes gateway status` and the dashboard, and the Windows
doctor probe. All four accept `degraded`; a dead watchdog-stamped
`degraded` is still excluded by `gateway_running=False`.
The parked-platform ERROR pointed at `/platform resume`, which only
resumes platforms in the retry queue - a non-retryable failure never
enters it. The remedy is `hermes gateway restart`.
Test: a degraded gateway with active agents is busy; not live => not
drainable.
Gate review on the stack:
- `hermes gateway restart` under systemd only accepted `gateway_state == "running"`
as proof of the replacement; a boot with a parked platform stamps `degraded` for
its whole life, so every restart/update on such a host waited out the 60 s+
budget and reported a false failure while the gateway was serving. The verifier
now accepts `degraded` as restarted and prints one DEGRADED warning.
- Drain release and scale-to-zero wake re-stamped `running` unconditionally, wiping
the parked-platform signal after the first `.drain_request.json` cycle. Every
"we are serving" stamp now goes through `_serving_state()`; the mixed fatal +
retryable boot path sets the flag too (it fell through as a plain run before).
`_startup_parked_platforms` is a bool — the joined error text was only ever logged.
- `_wait_for_tcp_port_free`: an unresolvable or unreachable configured host raised
a non-refused OSError on every probe and burned the full 10 s wait; only a
connect timeout means "listener alive", anything else means nothing to wait for.
- Windows `restart()` replaces its blind `time.sleep(1.0)` "let Windows release the
port" with the same configured-address wait.
- The api_server bind retry rebuilds the AppRunner per EADDRINUSE attempt instead of
calling aiohttp's private `_unreg_site`; attempt count is a named constant.
Test: a `degraded` replacement is reported as restarted (red on the previous head).
E2E re-run: predecessor releases the port 0.5 s after the first bind → bound on the
third attempt (+0.61 s), connect() True.
The shared startup gate returned early whenever any sibling platform connected,
so a non-retryable fatal platform failure left one WARNING and a 'running'
runtime state: the gateway served messaging while every API client was gone,
with nothing retried and no observable trace.
Park the failures, log at ERROR and stamp gateway_state 'degraded'. Retryable
peers are untouched (the reconnect watcher heals them). Behaviour and exit code
unchanged; only logging and health reporting. Port-race half of the report is
left to the open PRs #92060/#96419.
(cherry picked from commit edfb98f86911b70aa247ab0f718f0218002e154e)
Gate review: releasing an undispatched runtime claim kept its pre-claim
error ("503 ..."), so the redelivery timer re-armed and claimed/released
the row every tier for as long as the adapter was absent - the loop the
previous commit closed for `send_path_degraded` rows only. Only a flood
row keeps its error (the platform wait must be honoured); every other
row is released as reconnect-only and re-claimed by the reconnect sweep.
`is_reconnect_only()` replaces the three spellings of that predicate.
A final reply the platform rejected with anything other than flood control
(a 5xx, a transient parse error, an unclassified failure) sat in the ledger
as `failed` until the next gateway restart, because the runtime sweep only
claimed `send_path_degraded` and flood rows and no timer was armed for it
(#91653). The deadline-driven redelivery worker main already runs for flood
rows now covers every failed row: `retry_not_before` keeps the platform's
wait for flood rows, retries allowlisted reconnect errors at once, backs
any other rejection off by attempts (30 s, 2 min, 10 min) so an outage is
not hammered, and returns None for a whole-chat death (blocked bot, deleted
chat) so a dead target is never resent. `pending_flood_retries` becomes
`pending_retries`; the adapter arms the timer on every rejection.
`release_runtime_claim`'s attempts decrement is untouched: it refunds an
attempt a claim never spent, which the backoff relies on unchanged.
design via #91655 by @dasgltd
Co-authored-by: Daniel Silva <daniel@dasg.ltd>
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).
The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).
- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
`get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
`group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.
Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
cron seed matrix🧵… ≠ inbound
after create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)
Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
Same bug class as the warm-up: `_start_free_tier_bootstrap` ran on a bare executor thread, and
`anon_auth.ensure_portal_identity` resolves the Portal through `_nous_portal_env_override`, so
under multiplex a non-production HERMES_PORTAL_BASE_URL read as absent and the identity would be
minted against the production Portal. Route the hop through `_run_boot_probe_in_launch_scope`.
Gated behind guest onboarding, so latent on today's fleet; one test pins the binding.
Hoist the "enter the launch profile's scope under multiplex, hop to the executor with
copy_context" into GatewayStartupMixin._run_boot_probe_in_launch_scope so the warm-up has one
call site instead of two executor branches, and the sibling boot probes can share it.
_scoped_operator_override now takes the candidate names: _nous_portal_env_override used to call
it twice (HERMES_PORTAL_BASE_URL, then the NOUS_PORTAL_BASE_URL alias), so one lost-scope event
logged two identical WARNINGs; one try around the list logs once and names both.
Tests: drop the dead UnscopedSecretError arm in the probe (the helper swallows it), fold the
no-scope-left-behind assertions into the multiplex test, keep the environ-semantics contract for
single-profile. Docstrings keep the WHY; the incident chain lives in the PR.
Scope entry that cannot be built falls back to the unscoped executor call (the
load_gateway_config_for_runner shape) instead of aborting startup. The no-scope-left-behind
assertions run in the warm-up's own task — asyncio.run copies the context, so asserting from the
caller proved nothing.
_warm_turn_prerequisites ran _warm_turn_machinery_sync on a bare executor thread with no
profile scope. get_tool_definitions runs every check_fn; the vision probe resolves live Nous
runtime credentials (resolve_vision_provider_client -> _resolve_nous_runtime_api ->
resolve_nous_runtime_credentials -> _nous_effective_routing). Under GATEWAY_MULTIPLEX_PROFILES
get_secret fails closed there, so HERMES_PORTAL_BASE_URL / NOUS_INFERENCE_BASE_URL read as
absent, the Portal allowlist heals to production, and a staging deployment's refresh token is
POSTed to portal.nousresearch.com: invalid_grant, "Nous OAuth state quarantined", ~10 s after
every boot and before any inbound turn. Reproduced live on hermes-agent-stg-gg-probe-test-0062
with the override present in the profile .env (four NAS re-seeds, four boot-time deaths); the
unscoped call stack was captured on the box.
The warm-up now enters the launch profile's _profile_runtime_scope and carries the context
into the executor thread with copy_context().run — the same shape _discover_gateway_mcp_tools
uses (#95518). Single-profile gateways are unchanged.
_scoped_operator_override logs which override was unreadable and why; the downstream
"ignoring invalid portal_base_url" warning never named the caller that lost its scope.
Tests: scoped warm-up resolves the .env override on the executor thread (red on main), no
scope leaks past the warm-up, single-profile keeps environ semantics.
A /handoff into a Telegram forum supergroup created a topic and bound the CLI session under
telegram🧵<chat>:<topic>, while the Telegram adapter keys every topic reply
telegram:group:<chat>:<topic> — the same restart-orphaning shape as the Slack case. Non-private
Telegram homes now use chat_type group; private-chat DM topics are unchanged.
`/handoff slack` bound the CLI session under `slack🧵<channel>:<ts>` while every inbound
reply in that thread resolves (via the Slack adapter's source shape) to
`slack:dm|group:<team>:<channel>:<ts>`. After a gateway restart the reply key found no
binding and the gateway opened a fresh empty session, orphaning the handed-off one (#111896).
The handoff destination now mirrors the adapter: chat_type `dm` for a D… home channel, else
`group`, plus the workspace scope_id (home channel provenance, falling back to the adapter's
channel→team map). Channel handoffs were equally affected (`thread` vs `group`), so the fix
covers both, not only DMs. Discord and Telegram destinations are unchanged.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
#108319 moved hooks.outbound[].secret_env from os.environ to get_secret, which fails closed
outside a profile scope while multiplexing is on. The launch profile's hooks: block is
registered from start(), before any turn scope exists, so the first secret_env target raised
UnscopedSecretError, register_from_config propagated it, and _register_config_hooks swallowed
it at DEBUG — every launch-profile outbound webhook (signed or not) was silently dropped on a
multiplexed gateway. Secondaries were unaffected: they register inside their own scope.
The launch profile now gets the same treatment: _register_launch_profile_config_hooks enters
_profile_runtime_scope(get_process_hermes_home()) when multiplexing is active, then registers.
Not the environ fallback — a scopeless multi-profile read has no authority over the launch
env; the spawn site is where the scope belongs. The swallow logs at WARNING: a dropped hook
block is not debug noise.
Two invariant tests: multiplex on, no scope, secret in the launch .env -> both targets
register, the signed one with the launch secret (0 register on main); a registration failure
surfaces at WARNING.
faulthandler.register(SIGUSR2, ..., chain=True) writes the stack dump and
then re-raises the signal to its previous handler. SIGUSR2's default
disposition is "terminate", so the diagnostic hook added for #70344 kills the
very process an operator is trying to inspect (rc = -12), which is what the
#110437 reporter hit while introspecting a long-lived gateway. Nothing else
in the gateway installs a SIGUSR2 handler, so there is nothing to chain to.
Test: a child interpreter runs the real _start_install_faulthandler, receives
SIGUSR2, and must still be alive with a dump in gateway_faulthandler.log.
Red on origin/main (rc=-12), green with chain=False. Docs: the stack-dump
signal is now documented next to the event-loop watchdog.
Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.
- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
naming the reason (shared process-level ingress owned by the default) and
the remedy (enable it on the default profile, or disable it here), plus a
`<profile>:<platform>` runtime-status stamp (state=disabled,
error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
`whatsapp: not served under multiplex (shared ingress owned by default)`
from that stamp; /api/status excludes `disabled` entries from the
platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:
- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
timeout were bare os.getenv → the default profile's values; a job without a
model silently ran on the default's HERMES_MODEL instead of refusing.
cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
kanban worker inherited the launch profile's non-credential .env settings
and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
HERMES_MODEL) — X's worker ran in the default's docker image on the
default's model. tools/environments/local.py::strip_launch_profile_env drops
them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
assignee (toolset probes call get_secret without a scope → swallowed
UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
writing the completed tick into the DEFAULT profile's state.db and leaving
X's row awaiting_response forever; the --until judge ran with the default's
aux credentials.
- background processes: a secondary's processes.json (scope-relative since
adf23550f5) was never read at startup; its processes were not re-adopted
and notify_on_complete notices were lost. Startup recovers every served
home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
drain for the ambient profile (default's mode for everyone; X's `off`
dropped a sibling's `all` event), recovered watchers used the default's
mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
_deliver_platform_notice used the default's GatewayConfig so a secondary's
notice_delivery: private went public.
Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
Inside a routed satellite's turn the ambient scope is the satellite's, whose
.env has no token or allowlist. Five sites still called `_is_user_authorized()`
directly there — `/topic`, the sibling-thread `/stop` grant, plugin message
injection, Discord voice transcripts and startup auto-resume — so the shared
bot's owner was refused ("not authorized to use /topic") and a satellite that
DID copy an allowlist widened who may drive those commands. They now go through
`_is_user_authorized_for_source`, and `_under_authorization_profile` derives the
transport home from the delivering adapter's owner when ingress did not stamp
one (restored/cached sources).
`_adapters_for_profile` (the resolver behind `_authorization_adapter`,
`_adapter_for_source` and now `_resolve_injection_adapter`) returns the primary
map for a shared-bot satellite: a served profile with the `{}` startup
placeholder, no reconnect pending, targeted by a default-bot route. Kanban and
cron already applied that rule; the gateway's own resolvers returned None, so
heartbeats, process completions, goal notices and delegation results for such
profiles were undeliverable after a restart. A secondary that owns a credential
(connected on any platform, or queued for reconnect) still fails closed.
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract
The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:
- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
covers auxiliary work) and skips Nous for vision, which the welcome model does not take.
- The structured 429 body was never read. The classifier now parses `reason` /
`retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
are rate limits that honour `retry_after` and never rotate the free tier's only credential.
The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
fall back instead of retrying or re-exchanging. The terminal paths say what happened and
name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).
- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
beside the rate-limit and credits headers; the next call moves the session, and the config
default when it still names `nous/welcome`, to the backing model the gateway named.
- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
outside the host allowlist, routing defaulted to inference-api, where every request is a
400. A guest now defaults to the welcome literal at the exchange, in the shared store's
shape, and in effective routing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)
* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use
A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.
nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)
* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit
Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.
The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)
* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use
`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).
The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.
Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.
(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)
* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent
When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.
`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.
(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)
* fix(auth): the free tier outranks implicit host credentials in provider resolution
On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.
The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.
Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.
Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.
(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)
* fix(auth): review follow-ups for the free-tier rung (NS-829)
- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
the free tier off; its contract is the boto chain, and the free tier now
sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
before consulting the resolver, so a gateway boot on a machine with AWS
credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
six (parametrized ladder cases; a failed mint that returns None or raises
falls through to Bedrock).
scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.
(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)
* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone
The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.
`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.
`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.
Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).
Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.
Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.
* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read
Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.
Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.
`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.
Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.
`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).
Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.
Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.
Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.
* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"
A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.
Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.
Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.
* fix(copy): free-tier text stops promising a connector transfer and never names the config key
Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.
The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).
The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.
zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.
* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn
The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.
`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.
Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.
The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).
Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.
* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll
The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.
`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.
`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.
Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.
* fix(cli): the banner names the free tier's model instead of "no model configured"
The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.
`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").
Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.
* fix(aux): vision on the free tier uses nous/welcome too
The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)
* fix(gateway): hermes gateway run is a boot owner of the free tier too
Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).
GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.
Live, real GatewayRunner.start against a fake portal in a fresh home:
gate on -> 1 create, identity persisted, resolve_runtime_provider=nous,
/login precondition sees the identity
gate off -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.
Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.
---------
Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>