Commit Graph

33468 Commits

Author SHA1 Message Date
nftpoetrist
8099745a4e fix(sms): scope TWILIO_PHONE_NUMBER per profile
__init__ and _standalone_send read TWILIO_PHONE_NUMBER via raw os.getenv,
right next to the already-scope-aware TWILIO_ACCOUNT_SID/TWILIO_AUTH_TOKEN
(_get_scoped_secret, established by merged PR #76664). Under
gateway.multiplex_profiles, a secondary profile with its own Twilio account
would silently send replies from the default profile's bridged
TWILIO_PHONE_NUMBER instead -- Twilio rejects a from-number not owned by
that profile's account (auth failure), or the message gets attributed to
the wrong sender.

__init__'s path is currently unreachable via normal secondary-profile
adapter construction ("sms" is in gateway/config.py's
PORT_BINDING_PLATFORM_VALUES, so _start_one_profile_adapters refuses to
construct a port-binding platform for a secondary profile) -- fixed anyway
for consistency with its sibling reads and to avoid a latent bug if that
guard's scope ever changes. _standalone_send's path IS reachable: it's the
generic cron/home-channel out-of-process delivery hook, which runs inside
_profile_runtime_scope.

Swaps both reads to the adapter's own existing _get_scoped_secret() helper,
matching its sibling account_sid/auth_token reads exactly -- no new
mechanism needed.

Adds a TestMultiplexProfileScope test class to tests/gateway/test_sms.py
mirroring the established Buzz/Discord/Telegram/WhatsApp/LINE/DingTalk/
Teams coverage for this bug class.

(cherry picked from commit 2aeb3148367ecf87598d4458221a5fbf30c0af01)
2026-09-10 18:13:53 -07:00
Teknium
08830efd96 fix(secrets): secret-source re-pull no longer latches an empty snapshot or wipes sibling profiles
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).

Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (1aa62ceb45 isolated the routed reload but not the reset).

Change:
- env_loader: snapshot every name a source supplied (provenance + skipped_existing) from
  the home's effective environment, so a shadowed re-apply keeps the values it had.
- reset_secret_source_cache(hermes_home=None): optional per-home reset; global clear kept
  for tests/config edits.
- cron per-fire re-pull and plugins._refresh_secret_sources_after_discovery reset + reload
  only the home they resolve to.
- Two invariant tests (red on base) in tests/test_env_loader_secret_sources.py; docs note in
  secret-source-plugin.md.

Reported-by: luochen1990
Addresses #102041
2026-09-10 18:13:16 -07:00
Teknium
580322ef1e fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for
every name tagged in the process-global `_SECRET_SOURCES` map. That map is
filled by EVERY served profile's secret-source hydration, while `os.environ`
only ever holds the LAUNCH (default) profile's values — so once any profile's
1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio
MCP server was started with the default profile's token.

Resolve those names through the active profile's secret scope (`get_secret`)
instead: the routed profile's value, or omitted when that profile has none.
Under multiplex `get_secret` never falls through to environ; single-profile
runs keep the .env overlay + environ behaviour, so the existing "vault vars
reach MCP subprocesses" contract still holds there. `secret_source_names()`
exposes the tagged NAMES only — values are never read from the shared map.

Docs: the multi-profile guide's "MCP subprocesses only see their own profile's
secrets" claim is now true for source-injected names too; say so explicitly.
2026-09-10 18:12:39 -07:00
Nick Seelert
89fb1d028e fix(mcp): retain profile secret scope during discovery
(cherry picked from commit f344095fcf16c5d154ef2d207cce9b13339bf2e4)
2026-09-10 18:12:39 -07:00
Teknium
2952dc62bc fix(feishu): drive-comment turns and WS-thread callbacks run under their own profile, not the launch profile
Under `gateway.multiplex_profiles`, a Feishu adapter is built and connected inside
`_profile_runtime_scope` (HERMES_HOME override + secret scope as contextvars), but two
hops started from an EMPTY context and so executed under the LAUNCH profile:

- `feishu_comment.handle_drive_comment_event` ran the whole comment AIAgent turn on a bare
  `loop.run_in_executor(None, ...)`: model/credential resolution raised
  `UnscopedSecretError` (silent empty reply), or — when the default profile held the same
  key — used the default profile's config/model/state for a secondary profile's doc.
- `FeishuAdapter._connect_websocket` ran the lark WS client on a bare executor thread. The
  SDK fires every event/card callback on that thread and they hop back to the adapter loop via
  `run_coroutine_threadsafe`, which copies the CALLER's context — so all pre-handler work
  (inbound media caching, `.update_response` marker, FEISHU_REACTIONS env, drive comments)
  ran unscoped. The pending-inbound drainer thread spawned from that callback had the same
  shape.

Carry the scope across each hop with `contextvars.copy_context().run`. For the SDK-owned
thread the snapshot is taken once at `_connect_websocket` (inside the profile scope; the
restart supervisor task inherits it too), so no per-callback re-scoping is needed.

Supersedes the `_submit_on_loop` re-scoping approach of #63962 (nateEc, earliest fix):
scoping the WS thread at its source covers every callback without rebuilding the secret
scope per call.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
2026-09-10 18:12:02 -07:00
Teknium
2b4deeb32b fix(auth): key sibling per-process credential memos by profile home under multiplex
Same class as the resolve_nous_access_token memo: three more process-wide
memos carried a credential resolved under one profile's HERMES_HOME override
into another profile's turn for their TTL.

- hermes_cli/nous_billing.py::_token_cache (30s (token, base) memo for the
  charge poll loop) was a single unkeyed slot -> dict keyed by
  hermes_home_key(); invalidate_cached_token() clears the dict.
- agent/moa_loop.py::_runtime_cache carried api_key/base_url/api_mode keyed
  only (provider, model) for 5 min -> (hermes_home_key(), provider, model).
- agent/auxiliary_client.py::_client_cache_key had no profile component, so
  callers that omit api_key (pool / Nous auth.json paths) could be handed a
  client built with another profile's bearer -> hermes_home_key() leads the key.

WHY hermes_home_key(): it reads the per-turn HERMES_HOME override the
multiplex gateway sets (falling back to the env var), and it is
symlink-stable, so the memo key is exactly the credential home the
resolution itself read from. Profiles stay independent islands; the
default-profile process env never leaks into a secondary's turn.

Tests: one invariant per site, proven red on origin/main.
2026-09-10 18:11:25 -07:00
pierrenode
173105ce6f fix(auth): scope the resolve_nous_access_token memo to the active profile
resolve_nous_access_token()'s 5s startup-burst memo (#76930) cached the
resolved Nous Portal access token in a single module-level slot keyed
by nothing but wall-clock time. The underlying resolution is
profile-scoped: _auth_file_path() reads get_hermes_home(), which
checks the context-local _HERMES_HOME_OVERRIDE ContextVar before
falling back to the HERMES_HOME env var — gateway/run.py and
tui_gateway/server.py set that override per-profile for multiplex
concurrency.

In a multiplex gateway serving two profiles with different
authenticated Nous accounts, if profile A's context resolves a token
and profile B's context calls resolve_nous_access_token() within the
next 5 seconds, profile B received profile A's cached token — used to
authenticate against the managed tool gateway / relay self-provisioning
under the wrong account.

Key the memo by str(get_hermes_home()) instead of a single slot, so
each profile's context reads only its own cached token. The lock
around read/write is unchanged; only the cache's shape moved from a
single (timestamp, token) tuple to a dict keyed by resolved home.

(cherry picked from commit 9d8846b88cfd2ea74c6958d5f8f28a50880dda75)
2026-09-10 18:11:25 -07:00
Teknium
3b044261b6 fix(gateway): media denylist covers every profile's credentials, not just the launch home
Under `gateway.multiplex_profiles` one process serves every `<root>/profiles/*`,
but `_media_delivery_denied_paths` expanded `_ROOT_CREDENTIAL_PATHS` only under
the import-time `_HERMES_HOME` / `_HERMES_ROOT`. A `MEDIA:<root>/profiles/<B>/.env`
(or auth.json, state.db, config.yaml, sessions/, mcp-tokens/) emitted in ANY
profile's turn — including B's own, whose HERMES_HOME override was never
consulted — passed validation and was natively uploaded to the chat. The ALLOW
side (`_profile_cache_roots`) already enumerated profiles at check time; the
DENY side did not.

`_credential_home_roots()` now yields the active `get_hermes_home()`, the shared
root and every `<root>/profiles/*` at check time (shared `_profile_dirs()` with
the allow side), and the denylist is built from that. Profile cache artifacts
and plain agent-written files under a profile stay deliverable.

Live repro (/tmp/mux_audit/fix-media-denylist/repro.py): before, all five of
profiles/B/{.env,auth.json,state.db,config.yaml,sessions/s1.json} validated as
deliverable while <root>/.env was blocked; after, all None, cache/images/gen.png
and report.pdf still deliverable.

No prior report. Write-side analogue: #107327 / #107335 (memo keying, different
mechanism — left as is).
2026-09-10 18:10:48 -07:00
Teknium
942973ae90 fix: every Bedrock client rebuild lands on the startup wire (Claude SDK, Converse region, guardrails)
Follow-up to 564aef2946 (Mantle SigV4 on /model). The same class of bug covered the
other two Bedrock wires and two more rebuild paths:

- Claude on Bedrock (anthropic_messages): startup builds an AnthropicBedrock SDK
  client (SigV4 via boto3). /model, fallback-to-Bedrock and fallback restore built a
  plain Anthropic client with api_key="aws-sdk" against bedrock-runtime → 401/403.
- Converse models (bedrock_converse: Nova, DeepSeek, Llama): only agent_init set
  _bedrock_region / _bedrock_guardrail_config. After a rebuild the transport fell
  back to us-east-1 and guardrail_config=None, so eu-/ap- users hit the wrong region
  and configured Guardrails silently dropped. switch_model also built a pointless
  OpenAI client against bedrock-runtime.

Introduce bedrock_adapter.bind_bedrock_runtime(agent, base_url, api_mode) as the one
place that puts an agent on a non-Mantle Bedrock wire, and call it from agent_init
(replacing the two duplicated bodies), _build_switched_client, _rebuild_primary_client
and _swap_fallback_clients. try_recover_primary_transport had a hand-copied version of
_rebuild_primary_client's ladder (and would have hit the same gap); it now calls the
shared helper. The region/guardrail parsing that lived in agent_init and
runtime_provider_backends is now bedrock_region_from_runtime_url /
bedrock_guardrail_config in the adapter.

Live probe on origin/main across switch_model, restore_primary_runtime and
_swap_fallback_clients for both wires: 0/6 correct before, 6/6 after.
2026-09-10 18:08:16 -07:00
g3org3yo
6c3d4a4af7 fix(desktop): resolve plugin SDK namespaces lazily, not at module scope
runtime.ts captured the SDK namespaces (the plugin SDK, React and the two jsx
runtimes) in a module-scope object literal, and that module sits in an import
cycle: sdk/index -> @/contrib/* -> contrib/runtime-loader -> sdk/runtime ->
sdk/index.

In an unbundled (dev) graph the namespace object is live, so the capture works.
In a production bundle the bundler emits the SDK namespace as a hoisted `var`
whose assignment lands AFTER the literal that reads it, so the captured value
is `undefined` (no TDZ error) and `Object.keys(GLOBALS[globalKey])` in
`shimUrl()` throws "Cannot convert undefined or null to object".

That throw happens inside `loadRuntimePlugin()` -- via `unsupportedImports()`
-> `sdkImportMap()` -> `shimUrl()`, which run for every source before it is
evaluated -- so it is content-independent: EVERY plugin loaded from
$HERMES_HOME/desktop-plugins/<id>/plugin.js (and the desktop/plugin.js half of
a unified package) fails to load, showing status "failed" in
Capabilities -> Plugins.

Resolve the namespaces at call time instead: `installPluginSdk()` and the shim
builder only ever run once the app is up, so reading them there is always safe
and statement ordering can no longer matter.

Repro (production build only): build apps/desktop for production, drop any
plugin.js into ~/.hermes/desktop-plugins/<id>/, start the app ->
[plugins] runtime load failed (<id>) TypeError: Cannot convert undefined or
null to object (.../assets/sdk-<hash>.js:5).

Note: a vitest unit test cannot catch this (dev module graph keeps the
namespace live); the faithful guard is a production-build smoke test that
loads a fixture plugin through loadRuntimePlugin().
2026-09-10 15:35:31 -07:00
Siddharth Balyan
4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Siddharth Balyan
cbcf7b72f7 feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop

* feat(gateway): /signin signs the free tier into a Nous account from a DM

* feat(cli): chat surfaces name /signin as the sign-in verb

* fix(auth): review follow-ups for the shared sign-in flow and /signin

* fix(i18n): carry the /status free-tier line in every locale catalog

* refactor(cli): the chat sign-in command is /login

* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
2026-09-11 03:45:33 +05:30
Siddharth Balyan
3b01b4ce0f feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors

The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status
answers has_guest / enabled / carries_inference / notice_pending with zero network, and
free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check
reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and
billing.state carry free_tier (billing answers the free tier locally instead of a portal call that
can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or
locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector
transfer and returns its code and consent URL; the poller waits for the transfer before the token
grant, persists the account, runs settle_after_upgrade, and the poll response gains reason,
account_email and model.

* feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog

The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch
intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding
overlay opens on a ready screen when the free tier carries inference, else a one-time strip above
the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model /
Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the
tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that
drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens;
Done settles billing, model options, providers and re-homes a session still on nous/welcome. The
picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes.

* fix(desktop): free_tier.status starts the free tier's background setup when no identity exists

A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free
tier was never set up on the desktop: no connectors, no notice strip. The first status read now
starts the same one-attempt background setup; the call itself never waits.

* fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected

The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in).
The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free
tier, instead of Nous Portal · Connected.

* fix(desktop): Settings > Providers never files the free tier under Connected

* fix(desktop): the intro's shape is keyed on the route, not on the identity

free_tier.status reports available (an identity exists and the tier is on); whether inference
runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready
screen shows when that route is the free tier; the composer strip when the user's own provider
carries inference. An own-key install used to get the ready screen.

* docs(desktop): say what the free-tier chip is keyed on

* fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds

* fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro

Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the
transfer wait, after the token grant, and once more under the session lock together with the
save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an
attempt generation that every continuation checks after each await, so a poll from a closed
attempt cannot publish over the one on screen (and its backend session is cancelled). The ready
screen comes down only after the backend recorded the acknowledgement. A composer still mounted
takes over the notice claim when its owner unmounts. One thin test per hole.
2026-09-11 03:45:32 +05:30
Siddharth Balyan
04a76c4109 fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller

Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a
sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host
serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs
after the account is persisted: a config on the free tier's route moves to the account's inference
host and the recommended default for the account's plan, through the same config write a plain
Nous login uses; a config on the user's own model is left alone. The pick is the one
GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model
so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default.

* fix(auth): a sign-in completion with no eligible recommendation leaves no default model

The static provider-wide default is not narrowed by the account's plan or org policy, so writing it
as a fallback could persist a model the account may not use. When the recommendation cannot yield a
model, the route still moves to the account's host but model.default is left unset; the CLI says so
and points at `hermes model`.

* docs(free-tier): say what happens when no recommendation is available after sign-in

* fix(auth): sign-in completion moves the host and clears the default in one config write

Two writes could fail between them and leave the account host paired with nous/welcome.
_update_config_for_provider gains clear_default so the caller with no model to offer removes
model.default in the same atomic write that sets the host.
2026-09-11 03:45:31 +05:30
Siddharth Balyan
a2db110ccc feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping

A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.

* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin

* docs(user-guide): free tier and signing in

New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.

* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account

Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.

* fix(model): switching off the free tier points at signing in, never hops providers

* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once

* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token

* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check

Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.

* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors

The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.

* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off

* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check

The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.

* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure

A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.

* fix(auth): decide the Nous model together with the route on every credential-pool swap

The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.

* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start

* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died

The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.

* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential

The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.

* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only

A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
2026-09-11 03:45:31 +05:30
SHL0MS
2ddeba9e17 Merge pull request #99523 from somewheresy/justin/e-1047-route-hermes-actual-provider-through-chat-completions-with
fix(providers): use chat completions for all Actual routes
2026-09-10 17:25:53 -04:00
Teknium
564aef2946 fix: /model onto Bedrock Mantle keeps SigV4 auth instead of 401ing
Startup (agent_init._init_openai_client) ran configure_bedrock_openai_client_kwargs,
so the aws-sdk sentinel became a SigV4-signing httpx client. Every later client
rebuild — switch_model, fallback restore, credential rotation, request-scoped
clients — went through create_openai_client with bare {api_key, base_url} kwargs,
so the OpenAI SDK sent "Authorization: Bearer aws-sdk" and Mantle answered 401
"Invalid bearer token". Symptom: `hermes --provider bedrock --model
openai.gpt-5.6-terra` works, `/model openai.gpt-5.6-terra` inside the CLI/TUI fails.

Install SigV4 in create_openai_client itself (the single chokepoint every primary
OpenAI-wire client passes through) whenever the base_url is a Mantle host, so all
rebuild paths inherit the fix rather than each remembering to call the adapter.
A real AWS_BEARER_TOKEN_BEDROCK key is left alone (the adapter only rewrites the
aws-sdk / no-key-required placeholders).

Live repro (local HTTP sink capturing the Authorization header after switch_model):
before "Bearer aws-sdk", after "AWS4-HMAC-SHA256 Credential=...".
2026-09-10 12:18:44 -07:00
Justin Bennington
8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Justin Bennington
4135933ec9 fix(providers): preserve Actual routing during startup and key reload (E-1047) 2026-09-10 15:01:17 -04:00
Justin Bennington
8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington
d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
jonny
872bafd58d Merge pull request #107576 from NousResearch/fix/107109-probe-settle
fix(desktop): settle the login-shell PATH probe and kill the probe child on timeout
2026-09-10 20:48:35 +02:00
Teknium
f98cb00a8b fix(vault): 2FA review follow-ups — per-digit fill only for an unmistakable maxlength=1 widget; honour otpauth digits/period/algorithm
Reviewer findings on #107585:
- build_otp_fills split any >=4 code-like controls into digits. A page with
  promo/zip/referral 'code' inputs next to the real OTP box would have had a
  digit sprayed across unrelated fields. Split now requires exactly len(code)
  controls that are all maxlength=1, same form, adjacent in DOM order
  (inspection JS exports maxLength); anything else fills ONE field, the
  best-scoring one. Verified on the real Browser Use stack: 6-box widget gets
  one digit each; scattered page fills only the one-time-code input.
- normalize_otp_secret dropped digits/period/algorithm from otpauth:// URIs,
  so an 8-digit or SHA-256 authenticator would get wrong codes. Non-default
  parameters are now stored as seed|digits|period|algo and honoured (RFC 6238
  SHA-256 8-digit vector added); hotp:// is rejected explicitly.
- rebase on main (prompts.ts conflict) + prettier.
2026-09-10 11:48:01 -07:00
Teknium
d9ca9c974d feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped
the agent cold: the login classifier excludes one-time-code fields on
purpose (a password must never land in an OTP box) and there was no tool
for the second step, so the only move was to ask in chat.

browser_vault_enter_code
  Fills the one-time code the current page asks for. Two sources, same
  invariant as passwords (the code goes to the page over the supervisor
  socket and never enters model context):
  - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib,
    verified against the RFC test vectors), 1Password `op item get --otp`,
    Bitwarden `bw get totp`. Nobody is asked.
  - no seed: the surface prompts "Verification code for {site}"; the user
    types what their phone/email/app shows. Enter on empty / Skip declines
    and the tool returns code_declined ("do not ask again this turn").
  no_code_field tells the model the site wants a passkey / hardware key /
  app approval: hand it to the user's device and wait for navigation.
  Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order.

Surfaces
  CLI: sudo-style panel, code shown as typed (not a secret worth masking,
  typos must be visible), Enter submits, ESC/empty skips.
  Desktop: "Verification code for {site}" card via vault.code.request /
  vault.code.respond (gateway), owner-routed like the other vault prompts.
  Settings → Passwords & Logins: optional "Authenticator key" field on the
  add form (base32 or otpauth:// link); items with one show a "2FA auto"
  badge. `hermes vault add` asks for the same optional key.
  browser_vault_fill's result now says what to do next ("if the site asks
  for a verification code, call browser_vault_enter_code with this handle").
  Six locales.

Verified live (real model, local 2FA site that checks the TOTP; CLI PTY):
  A. login saved with authenticator key → signed in through 2FA, zero
     prompts, code/password absent from the transcript
  B. login without key → code panel → user types code → signed in
  C. panel dismissed → agent stops and explains, never asks in chat
Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit
spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip).
2026-09-10 11:48:01 -07:00
Teknium
77e55b4d1f fix(desktop): show each in-app tip once, never lap the catalog again
The idle tip rotation walked the catalog as a ring: after the last tip it
wrapped to the first, so a user who had already seen every tip kept
getting "Start fresh", "Teach it once", ... again every six hours for as
long as they used the app. Only the X stopped a tip, and letting a bubble
time out (the normal way it leaves) counted for nothing.

The walk is now one lap. nextTip also steps over every tip in the seen
ledger ($tipShownAt, which already recorded every catalog tip that
reached the screen), so a tip shows once however it left, and the
rotation runs dry once every tip has had its moment. Settings > Reset
clears the seen ledger and the cursor as well as the retired set, and its
button counts what a Reset would actually bring back (shown or closed,
counted once). Agent tips carry no catalog id and are untouched.

Live repro (Playwright against the worktree's Vite renderer, all nine
tips seeded as seen, clock fast-forwarded past settle + cooldown):
origin/main re-showed "Start fresh"; fixed renderer shows nothing;
a fresh user (nothing seen) still gets the first tip.
2026-09-10 11:18:16 -07:00
kshitijk4poor
1e7d29a081 test(gateway): exercise the shared overflow classifier instead of a copy
test_7100 replicated the phrase list inline, so it tested its own copy rather
than production; point it at is_context_overflow_failure_result. Drop a dead
`error=` parameter from the normalizer test helper.
2026-09-10 23:45:03 +05:30
kshitijk4poor
949724b321 fix(gateway): one context-overflow verdict for the reply and the transcript skip
Moving the `if response` guard below the failed branch exposed the normalizer's
loose overflow predicate (bare "token"/"exceed"/"context"/"payload", or any 400
on a long session) to failed turns that carry real text: billing, rate-limit,
auth and content-policy replies were rewritten to "Session too large / /compact".

Hoist run_turn's stricter classifier (compression_exhausted, multi-word phrases,
400 on history > 50) into a module-level `is_context_overflow_failure_result` and
use it for both the #1630 transcript skip and the user-facing rewrite, so the two
can never disagree. Populated text is only rewritten when it is the bare provider
envelope (`_looks_like_gateway_provider_error`) on an overflow turn; curated agent
text (compression-timeout guidance, /compress hint) survives. Replaces the
sanitizer-wording test with the passthrough invariants that catch the regression.
2026-09-10 23:45:03 +05:30
xxxigm
613d1f7c03 fix(gateway): keep session-too-large reply when a failed 400 still has text 2026-09-10 23:45:03 +05:30
xxxigm
722bb1a674 test(gateway): pin overflow reply when a failed 400 still has text 2026-09-10 23:45:03 +05:30
yoniebans
8762a08b67 test(desktop): verify PATH probe child kill on timeout 2026-09-10 20:08:03 +02:00
KoNit-K
f2b33fb4a2 fix(desktop): kill the PATH probe child on timeout 2026-09-10 20:08:03 +02:00
Siddharth Balyan
d5aaaa4a1b fix(tui-gateway): a hidden seed row stays out of search, a partial seed copy is rolled back, live resume counts the wire (#107562)
Two independent reviews of the seeded-create change found three more
places where the newly durable hidden row, or the new create-time copy,
was not handled by the same rule as the rest of the path:

- Message search (dashboard search and the session_search tool) had no
  display_kind filter, so a hidden opening row matched a query the
  person never saw. The shared search predicate now skips hidden rows.
- _seed_row left the fresh session row behind when the transcript copy
  failed after the row was committed. The first prompt's retry copies
  the whole seed, so a kept partial copy would be duplicated. The row
  is now deleted when the copy did not complete, the compensation
  _persist_branch applies to branch children; the first prompt then
  starts clean.
- _live_session_payload (a resume that reuses a live session) reported
  message_count as the raw history length while its messages array was
  filtered. It now follows _resume_response: the stored size when
  messages are omitted, else the wire count.

Tests: the two seeded-create tests now drive the first-submit path
through _persist_session_row_for_submit, the function prompt.submit
calls, and assert search and the reuse-live count; a third test pins
the rollback (no row after a failed copy, one copy after the retry).
2026-09-10 18:05:33 +00:00
Siddharth Balyan
c22a8d8e3f Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once

session.create accepts opening messages. Three defects sat in that path:

- A seeded session without a parent was never persisted at create, so a
  restart before the first prompt lost it and session.resume answered
  4007. Only branch children (#93959) were persisted up front. The
  same rationale applies to any seeded create: seeded content is
  intent, not an abandoned draft. Parentless seeds now persist their
  row, transcript and client title at create; empty drafts stay lazy.
- _coerce_seed_history dropped display_kind, so a seeded row tagged
  "hidden" (model-facing scaffolding) rendered as a user bubble. The
  coercion keeps "hidden" and only "hidden"; every other kind is
  stamped by the gateway at turn time and is not accepted from the wire.
- A branch child's seed was written twice: _seed_branch_row copied it at
  create but never marked it persisted, so the first prompt's
  _persist_branch_seed appended the copy again. The create path now
  sets _branch_seed_persisted, and the gate is a create-time `seeded`
  stamp instead of parent_session_id, so a resumed session (whose
  history comes from the DB) can never re-append its transcript.

Two invariant tests, both red on main: a parentless seed survives a
gateway restart with the hidden row kept out of the wire transcript and
not re-written by the first-submit path; a branch child's seed is stored
exactly once. The reasoning-fields fixture stamps `seeded`, the flag
session.create sets.

* fix(tui-gateway): a hidden seed row stays out of the list preview and the create count

Live-testing the seeded create on every surface showed two places where
the newly durable hidden row (display_kind="hidden") still surfaced:

- session.list built a session's preview from its first user row with no
  display_kind filter, so a hidden opening row (model-facing scaffolding
  the gateway never paints) became the sidebar preview. The preview
  predicate now skips hidden rows, in every listing query that shares it.
- session.create reported message_count as the raw seed length while its
  messages array already filtered the hidden row (2 vs 1). It now counts
  what is on the wire, the same rule session.resume applies.

Both are covered by the existing seeded-create test: the create count
equals the wire transcript, and the preview of a session whose first
user row is hidden is its first visible user row.

* fix(tui-gateway): a live unpersisted resume counts the wire transcript

session.resume on a live session that has no row yet reported message_count as
the raw history length while its messages array was already filtered, the same
mismatch the previous commit fixed on session.create. Count the wire, as the
cold, deferred and reuse-live resume paths already do.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-10 17:55:24 +00:00
Teknium
438a313500 docs(honcho): document a2aSessions and per-author writes on the site page 2026-09-10 10:45:57 -07:00
Teknium
d55f1ed0a8 test(honcho): trim the author-peer suites to their invariants
Keep the isolation contracts (bot never on the human peer, bot turn never in the
human session, writes refused mid bot-turn, one join per author, signature busts
the cache) and drop the alias/prefix/sanitize enumerations. 795 -> 454 lines.
2026-09-10 10:45:57 -07:00
Teknium
a42388b3a9 feat(agent): a2a_key names a bot author's turns
Dropped from the #103888 salvage for lack of a consumer; the honcho a2a
session key is that consumer.
2026-09-10 10:45:57 -07:00
Erosika
d74f13e4a0 fix(honcho): include a2aSessions in identity_signature
sync_turn reads a2a_sessions from the config bound when the provider was built. A cached gateway provider kept the old value after honcho.json flipped it, because the signature that busts that cache did not carry the flag.
2026-09-10 10:45:57 -07:00
Erosika
431cd9084b fix(honcho): bound the joined author peer memory by session count
_joined_author_peers kept an entry for every honcho session the manager ever wrote to. It now holds at most _SESSION_CACHE_MAX_SIZE sessions and drops the oldest past that, so a forgotten session's authors rejoin on their next write. A failed join no longer leaves an empty entry behind.
2026-09-10 10:45:57 -07:00
Erosika
bcac7e9465 fix(honcho): read an author join's observation flags through one manager method
The join read the manager-wide user_observe_me and user_observe_others directly. It now asks _join_observation_flags(honcho_session_id), which returns the same values today. #103889 stores the effective flags per session and replaces the body of that method.
2026-09-10 10:45:57 -07:00
Erosika
2ac7fcddf8 fix(honcho): a bot author never lands on the session's human runtime peer
_generated_runtime_peer_id takes a reserved set, and the bot path passes the session's human peer ids: each runtime id and the peer _resolve_user_peer_id returns for the key. bot:coder with a runtime human coder and no runtimePeerPrefix now gets the digest suffix. _explicit_user_peer_ids keeps its meaning for prefixed runtime users.
2026-09-10 10:45:57 -07:00
Erosika
a2c65ace0a test(honcho): cover bot:<connection>/<profile> authors end to end
Two senders named coder on different connections get different peers and different a2a sessions, and a userPeerAliases entry keyed by the full connection-qualified id wins.
2026-09-10 10:45:57 -07:00
Erosika
d96a6d9ab9 fix(honcho): include the workspace in identity_signature
identity_signature now carries cfg.workspace_id. The gateway folds these values into its agent cache key, and a workspace change in honcho.json reused a cached agent that was still bound to the old workspace.
2026-09-10 10:45:57 -07:00
Erosika
8704e9ca4c fix(honcho): put this agent's aiPeer in the a2a session key
_a2a_session_key now names the session <session>:a2a:<aiPeer>:<sender id>-<digest>. Two profiles that share a workspace and a session key wrote one sender's DMs into one Honcho session. The recipient peer comes from the same aiPeer derivation the session builder uses, moved into session_peers.assistant_peer_id_for so the two cannot drift.
2026-09-10 10:45:57 -07:00
Erosika
b4d7a33b51 fix(honcho): derive a bot author's peer from its full id with the runtime digest rule
_peer_id_for_runtime_id now looks up userPeerAliases by the full bot id and otherwise passes everything after bot: through _generated_runtime_peer_id. A digest suffix is added when sanitizing changed the id or the result equals peerName or an alias target, so bot:eri never resolves to the operator's peer and bot:a.b stays apart from bot:a-b. The docstring and README no longer claim a cloned profile's aiPeer defaults to the profile name.
2026-09-10 10:45:57 -07:00
Erosika
bdeb6f1b77 refactor(honcho): name the a2a session from core's a2a_key
The plugin spelled the `a2a:` prefix itself. `agent.turn_author.a2a_key` is the shared name
for a bot author's turns, so the session key now derives from it and every reader that files
bot turns apart agrees on the prefix. The resulting key is unchanged.
2026-09-10 10:45:57 -07:00
Erosika
170589dd10 fix(honcho): every bot author gets its own peer or its turn is skipped
A gateway platform marks a bot sender with its raw user id and a bot flag, never a `bot:` id.
`resolve_author_peer_id` treated that author as a human, so `pinUserPeer` collapsed a bot onto
`peerName` and an unresolved peer opened the a2a session under the human's peer. The resolver now
takes `is_bot` and gives every bot its own peer. `sync_turn` skips the turn when no peer resolves
or when the peer equals this agent's `aiPeer`.

The a2a session key carries an eight-character digest of the author id, so two ids that sanitize
alike stay in separate sessions. During a bot-authored turn `honcho_conclude` and `honcho_profile`
refuse writes and the built-in memory mirror is skipped, because conclusions and cards describe
the human. The README paragraph on bot DMs now matches the code.
2026-09-10 10:45:57 -07:00
Erosika
9f2a9384dc fix(honcho): pinUserPeer collapses the operator's accounts, not bot authors
with pinUserPeer on, resolve_author_peer_id returned None for every author,
so a bot dm's words were written under the human's pinned peer inside the
a2a session. the pin exists to unify one person's platform accounts. a bot
is not one of them.

bot: authors now resolve to their peer before the pin check, so a pinned
operator still gets bot speech attributed to the bot.
2026-09-10 10:45:57 -07:00
Erosika
df1513b728 docs(honcho): document a2aSessions and bot dm attribution 2026-09-10 10:45:57 -07:00
Erosika
f7d5ac3230 feat(honcho): write bot dms into their own a2a session
A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.

The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.

a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.

get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
2026-09-10 10:45:57 -07:00
Erosika
88fc9402d8 feat(honcho): declare identity_signature and drop the gateway's honcho keys
The gateway agent cache read honcho.json itself through a honcho-named
block in gateway/run.py and gateway/run_agent_cache.py. Every other memory
provider had no way to bust the cache when its identity mapping changed.

HonchoMemoryProvider.identity_signature() now returns the same values under
provider-neutral keys: user_identity, agent_identity, pin_user_identity,
runtime_identity_prefix, user_identity_aliases, session_prefixing. The
gateway files them under memory.<key> through the MemoryProvider hook. The
hook reads config only, memoizes on the file's mtime and size, and returns
an empty dict when the file cannot be read.

The honcho-specific extractor, its memo and its key tuple are gone from the
gateway. The pinPeerName cache-busting test now asserts on
memory.pin_user_identity.
2026-09-10 10:45:57 -07:00