Files
hermes-agent/gateway/run_agent_cache.py
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30

834 lines
46 KiB
Python

"""Agent cache, session model overrides, turn leases, run generations and conversation-scope reset
for GatewayRunner (MRO mixin). ``gateway.run`` internals are imported lazily inside method bodies
(import cycle), so ``patch("gateway.run.X")`` keeps intercepting them at call time."""
from __future__ import annotations
import importlib
import logging
import threading
import time
from contextlib import nullcontext, suppress
from typing import TYPE_CHECKING, Any, Dict, List, Optional
from agent.interrupt_compat import _accepts_keyword
from gateway.config import Platform
from gateway.session import SessionSource, build_session_context_prompt
from hermes_cli.config import cfg_get
if TYPE_CHECKING: # string annotations only; never imported at runtime (cycle)
from gateway.run import GatewayRunner # noqa: F401
from gateway.run_turn_runner import TurnRunner # noqa: F401
# Log-record parity with the origin module.
logger = logging.getLogger("gateway.run")
# Override fields layered onto runtime kwargs when non-None (partial overrides don't clobber defaults).
_OVERRIDE_APPLY_KEYS = (
"provider", "requested_provider", "api_key", "base_url", "api_mode", "credential_pool", "capabilities", "max_tokens",
)
def _first_agent(entry: Any) -> Any:
"""Unwrap a cache entry (``(agent, sig, ...)`` tuple or bare agent) to its agent."""
return entry[0] if isinstance(entry, tuple) and entry else entry
def _tuple_agent(entry: Any) -> Any:
"""Agent of a ``(agent, sig, ...)`` cache tuple; None for any other entry shape."""
return entry[0] if isinstance(entry, tuple) and entry else None
class GatewayAgentCacheMixin:
"""Agent cache, session model overrides, turn leases, run generations and conversation-scope reset for GatewayRunner."""
@classmethod
def _extract_cache_busting_config(cls, user_config: dict | None) -> dict:
"""Values that must bust the cached agent, as a flat dict keyed by 'section.key'. Missing keys /
non-dict sections yield None (still enters the signature). Includes the live tool registry
generation: MCP reloads mutate the registry without touching config.yaml."""
out: Dict[str, Any] = {}
cfg = user_config if isinstance(user_config, dict) else {}
for section, key in cls._CACHE_BUSTING_CONFIG_KEYS:
section_val = cfg.get(section)
if section == "checkpoints" and isinstance(section_val, bool):
# Legacy ``checkpoints: true``: a live toggle must still rebuild the cached agent.
out[f"{section}.{key}"] = section_val if key == "enabled" else None
else:
out[f"{section}.{key}"] = section_val.get(key) if isinstance(section_val, dict) else None
try:
from tools.registry import registry
out["tools.registry_generation"] = getattr(registry, "_generation", None)
except Exception:
out["tools.registry_generation"] = None
for key, value in cls._memory_provider_identity_signature(cfg_get(cfg, "memory", "provider")).items():
out[f"memory.{key}"] = value
return out
# Kept for the process lifetime: loading a provider imports its plugin module, and this runs on every inbound message.
_MEMORY_IDENTITY_PROVIDER_MEMO: dict[str, Any] = {}
@classmethod
def _memory_provider_identity_signature(cls, provider_name: Any) -> dict[str, Any]:
"""The active memory provider's ``identity_signature()``. ``{}`` when there is no provider,
it fails to load, or the hook raises."""
if not isinstance(provider_name, str) or not provider_name.strip():
return {}
name = provider_name.strip()
try:
instance = cls._MEMORY_IDENTITY_PROVIDER_MEMO.get(name)
if instance is None:
from plugins.memory import load_memory_provider
instance = load_memory_provider(name, register_skills=False)
if instance is None:
return {}
cls._MEMORY_IDENTITY_PROVIDER_MEMO[name] = instance
signature = instance.identity_signature()
return dict(signature) if isinstance(signature, dict) else {}
except Exception:
return {}
@staticmethod
def _agent_config_signature(
model: str, runtime: dict, enabled_toolsets: list, ephemeral_prompt: str,
cache_keys: dict | None = None, user_id: str | None = None, user_id_alt: str | None = None,
skip_context_files: bool = False,
) -> str:
"""Stable key from agent config: change → cached AIAgent rebuilt; unchanged → reused (frozen
prompt + schemas for cache hits). ``user_id`` / ``user_id_alt`` participate because Honcho
freezes them at init; omitting them in shared-thread keys would cross-attribute messages.
``user_id`` and ``user_id_alt`` are the runtime user identities carried by the current message's
gateway source. They participate in the cache key because the Honcho memory provider freezes them
into ``HonchoSessionManager`` at first-message init (see
``plugins/memory/honcho/__init__.py::_do_session_init``). Without them in the signature, a
shared-thread session_key (one in which ``build_session_key`` intentionally omits the participant
ID, e.g. ``thread_sessions_per_user=False``) would reuse the cached AIAgent across distinct users,
causing the second user's messages to be attributed to the first user's resolved Honcho peer. This
broke #27371's per-user-peer contract in multi-user gateways. Per-user agent rebuilds in shared
threads trade prompt-cache warmth for correct memory attribution.
"""
import hashlib, json as _j
# Fingerprint the FULL credential, not a short prefix: OAuth/JWT-style tokens often share a
# common prefix (e.g. "eyJhbGci"), so a prefix would give false cache hits across auth switches.
_api_key = str(runtime.get("api_key", "") or "")
blob = _j.dumps(
[
model,
hashlib.sha256(_api_key.encode()).hexdigest() if _api_key else "",
runtime.get("base_url", ""), runtime.get("provider", ""),
runtime.get("requested_provider", ""), runtime.get("api_mode", ""),
sorted((runtime.get("capabilities") or {}).items()),
sorted(enabled_toolsets) if enabled_toolsets else [],
# reasoning_config excluded — set per-message on the cached agent; no prompt/tool effect.
ephemeral_prompt or "",
sorted((cache_keys or {}).items()),
str(user_id or ""), str(user_id_alt or ""),
# skip_context_files changes the agent's frozen system prompt (context files in vs out):
# a toggled edit must rebuild the cached agent, not silently reuse it.
bool(skip_context_files),
],
sort_keys=True, default=str,
)
return hashlib.sha256(blob.encode()).hexdigest()[:16]
def _session_model_override(self, session_key: str) -> Optional[dict]:
"""Current in-memory /model override for ``session_key`` (None when absent)."""
state = self._peek_session_state(session_key)
return state.conversation.model_override if state else None
def _rehydrate_session_model_override(self, session_key: str) -> None:
"""Lazily restore a persisted /model override after a gateway restart: non-secret parts
(model/provider/base_url) are written through on /model and read back on first use; api_key
is never persisted and is re-resolved. No-op when an in-memory override or nothing exists."""
from gateway.run import _resolve_runtime_agent_kwargs_for_provider
store = getattr(self, "session_store", None)
if self._session_model_override(session_key) is not None or store is None:
return
try:
persisted = store.get_model_override(session_key)
except Exception:
logger.debug("Failed to read persisted session model override", exc_info=True)
return
if not persisted:
return
override: Dict[str, Any] = {k: persisted.get(k) for k in ("model", "provider", "base_url")}
provider = persisted.get("provider")
if provider:
# Re-resolve credentials for the persisted provider. On failure (e.g. credentials removed
# since the switch) keep the credential-less override — _resolve_session_agent_runtime
# falls back to env resolution and layers model/provider.
try:
runtime = _resolve_runtime_agent_kwargs_for_provider(provider)
for k in ("api_key", "api_mode", "credential_pool", "requested_provider", "max_tokens"):
override[k] = runtime.get(k)
override["request_overrides"] = dict(runtime.get("request_overrides") or {})
override["capabilities"] = dict(runtime.get("capabilities") or {})
if not override.get("base_url"):
override["base_url"] = runtime.get("base_url")
except Exception:
logger.debug(
"Credential re-resolution failed for persisted override "
"(provider=%s); using credential-less override", provider, exc_info=True,
)
self._session_state(session_key).conversation.model_override = override
logger.info(
"Rehydrated persisted /model override for session=%s: model=%s provider=%s",
session_key, override.get("model"), provider or "",
)
def _apply_session_model_override(self, session_key: str, model: str, runtime_kwargs: dict) -> tuple:
"""Apply /model session overrides (precedence over config.yaml defaults; ``None`` fields skipped
so partial overrides don't clobber defaults), returning (model, runtime_kwargs)."""
from gateway.run import _credential_pool_for_provider
override = self._session_model_override(session_key)
if not override:
return model, runtime_kwargs
model = override.get("model", model)
for key in _OVERRIDE_APPLY_KEYS:
val = override.get(key)
if val is not None:
runtime_kwargs[key] = val
# request_overrides reflects the switched-to provider; apply whenever the override recorded
# it (even as None) so switching to a provider without configured overrides clears a stale
# value left by the default provider's runtime resolution.
if "request_overrides" in override:
ro = override.get("request_overrides")
runtime_kwargs["request_overrides"] = dict(ro) if isinstance(ro, dict) and ro else ro
if (
runtime_kwargs.get("api_key")
and runtime_kwargs.get("credential_pool") is None
and override.get("provider")
):
runtime_kwargs["credential_pool"] = _credential_pool_for_provider(override.get("provider"))
return model, runtime_kwargs
def _snapshot_session_model_override(self, session_key: str) -> dict:
"""Capture a gateway session override before a one-turn switch."""
override = self._session_model_override(session_key)
return {"had_override": override is not None, "override": dict(override) if override is not None else None}
def _restore_session_model_override(self, session_key: str, snapshot: dict) -> None:
"""Restore the session override captured before a one-turn switch."""
if not session_key:
return
if snapshot.get("had_override"):
self._session_state(session_key).conversation.model_override = dict(snapshot.get("override") or {})
elif (state := self._peek_session_state(session_key)) is not None:
state.conversation.model_override = None
self._evict_cached_agent(session_key)
def _is_intentional_model_switch(self, session_key: str, agent: Any, config_model: str) -> bool:
"""True when *agent* running a model other than *config_model* is deliberate: a /model session
override names that model, or the Nous gateway moved the session off the ``nous/welcome``
alias that *config_model* still carries (``anon_auth.apply_model_switch``)."""
override = self._session_model_override(session_key)
if override is not None and override.get("model") == agent.model:
return True
# Exactly the recorded move (alias -> backing): a later fallback onto some other model is
# ordinary drift and still evicts.
return getattr(agent, "_nous_model_switch", None) == (config_model, agent.model)
def _release_running_agent_state(
self, session_key: str, *, run_generation: Optional[int] = None
) -> bool:
"""Pop ALL per-running-agent state for ``session_key`` (call at every site that ends a running
turn); True when cleared. Persistent state (model overrides, voice mode, approvals) is NOT
touched. With ``run_generation``, only clear if still current — a stale async unwind bumped
by /stop or /new must not clobber a newer run (returns False)."""
if not session_key or (
run_generation is not None and not self._is_session_run_current(session_key, run_generation)
):
return False
state = self._peek_session_state(session_key)
if state is not None:
if state.turn.lease is not None:
try:
state.turn.lease.release()
except Exception:
logger.debug("Failed to release active session slot", exc_info=True)
# One structured reset instead of a drifting pop-list. Turn-lease tokens are deliberately NOT
# cleared here — _release_turn_lease owns them.
state.turn.clear()
# Turn boundary: a running-agent slot was just released; persist the new (lower) in-flight count
# so the dashboard readout stays current. Preserves gateway_state (see _persist_active_agents).
self._persist_active_agents()
return True
def _held_turn_lease(self, session_key: str, run_generation: int):
"""Return ``(registry, turn)`` when ``session_key`` holds a lease token for ``run_generation``, else None."""
registry = getattr(self, "_turn_leases", None)
state = self._peek_session_state(session_key) if session_key and registry is not None else None
if state is None or state.turn.lease_token is None or state.turn.lease_generation != run_generation:
return None
return registry, state.turn
def _release_turn_lease(self, session_key: str, run_generation: int) -> bool:
"""Release the turn lease acquired by (``session_key``, ``run_generation``). Keyed by (routing
key, run generation) so a stale unwind pops only ITS token; the registry's identity check
refuses it if a newer turn holds the lease. Idempotent."""
held = self._held_turn_lease(session_key, run_generation)
if held is None:
return False
registry, turn = held
token, turn.lease_token, turn.lease_generation = turn.lease_token, None, None
try:
return registry.release(token)
except Exception:
logger.debug("Failed to release turn lease", exc_info=True)
return False
def _rebind_turn_lease(self, session_key: str, run_generation: int, new_session_id: str) -> bool:
"""Follow a mid-turn session_id rotation (compression) with the held turn lease, or an alias
key resolving the new id could start a concurrent turn the lease never sees. Call at every
mid-turn reassignment; no-op if no token."""
held = self._held_turn_lease(session_key, run_generation) if new_session_id else None
if held is None:
return False
registry, turn = held
try:
return registry.rebind(turn.lease_token, new_session_id)
except Exception:
logger.debug("Failed to rebind turn lease", exc_info=True)
return False
def _clear_conversation_scope(self, session_key: str, *, reason: str) -> None:
"""THE single conversation-boundary funnel (/new, /resume, suspension replacement,
compression-exhausted reset). New conversation-scoped dicts go in _CONVERSATION_SCOPED_STATE
so every boundary picks them up. Turn-scoped state (_running_agents/_ts, slot leases, turn-
lease tokens) is owned by _release_running_agent_state and NOT cleared. Idle agent-cache
eviction is NOT a boundary (a resumed turn rebuilds from these). getattr-guarded.
Why a funnel: these boundaries used to each carry a hand-copied pop-list of the per-session dicts,
and the lists drifted every time a new dict was added (#48031, #58403, #10702, #35809 were all
"boundary X forgot dict Y" bugs — e.g. /new cleared the /model override but not the /model --once
restore snapshot). Adding a new conversation-scoped dict now means adding its attribute name to
_CONVERSATION_SCOPED_STATE below; every boundary picks it up automatically.
"""
from gateway.run import _CONVERSATION_SCOPED_STATE
if not session_key:
return
state = self._peek_session_state(session_key)
if state is not None:
state.conversation.clear()
# Legacy plain-dict stores still in _CONVERSATION_SCOPED_STATE (not yet folded into
# SessionState), e.g. _pending_model_notes. SessionState-backed names resolve to MutableMapping
# views (not dict), so the isinstance(dict) guard skips them — already handled above.
for attr in _CONVERSATION_SCOPED_STATE:
store = getattr(self, attr, None)
if isinstance(store, dict):
store.pop(session_key, None)
self._clear_session_boundary_security_state(session_key)
logger.debug("Cleared conversation scope for %s (%s)", session_key, reason)
def _clear_session_boundary_security_state(self, session_key: str) -> None:
"""Clear per-session control state that must not survive a boundary switch."""
if not session_key:
return
pending_skills_reload_notes = getattr(self, "_pending_skills_reload_notes", None)
if isinstance(pending_skills_reload_notes, dict):
pending_skills_reload_notes.pop(session_key, None)
state = self._peek_session_state(session_key)
if state is not None:
state.persistent.approvals = None
state.persistent.update_prompt_pending = False
for mod, attr, what in (
("tools.slash_confirm", "clear", "slash-confirm"), ("tools.approval", "clear_session", "approval"),
):
try:
clear = getattr(importlib.import_module(mod), attr)
except Exception:
continue
try:
clear(session_key)
except Exception as e:
logger.debug("Failed to clear %s state for session boundary %s: %s", what, session_key, e)
def _begin_session_run_generation(self, session_key: str) -> int:
"""Claim a fresh, monotonically increasing run generation token (NEVER reset): a late result
from a worker /stop or /new invalidated is recognized and dropped."""
if not session_key:
return 0
persistent = self._session_state(session_key).persistent
# Monotonic by design (#28686): incremented here, NEVER reset.
persistent.run_generation = int(persistent.run_generation) + 1
return persistent.run_generation
def _invalidate_session_run_generation(self, session_key: str, *, reason: str = "") -> int:
"""Invalidate any in-flight run token for ``session_key``."""
generation = self._begin_session_run_generation(session_key)
if reason:
logger.info("Invalidated run generation for %s → %d (%s)", session_key, generation, reason)
return generation
def _is_session_run_current(self, session_key: str, generation: int) -> bool:
"""Return True when ``generation`` is still current for ``session_key``."""
if not session_key:
return True
state = self._peek_session_state(session_key)
current = state.persistent.run_generation if state is not None else 0
return int(current) == int(generation)
def _bind_adapter_run_generation(self, adapter: Any, session_key: str, generation: int | None) -> None:
"""Bind a gateway run generation to the adapter's active-session event."""
if not adapter or not session_key or generation is None:
return
with suppress(Exception):
interrupt_event = getattr(adapter, "_active_sessions", {}).get(session_key)
if interrupt_event is not None:
interrupt_event._hermes_run_generation = int(generation)
async def _interrupt_and_clear_session(
self, session_key: str, source: SessionSource, *, interrupt_reason: str,
invalidation_reason: str, release_running_state: bool = True,
) -> None:
"""Interrupt the current run and clear queued session state consistently."""
from gateway.run import _AGENT_PENDING_SENTINEL, _reap_gateway_turn_processes, request_hard_interrupt
if not session_key:
return
state = self._peek_session_state(session_key)
running_agent = state.turn.agent if state else None
_process_task_id, _process_baseline = "", None
if running_agent and running_agent is not _AGENT_PENDING_SENTINEL:
request_hard_interrupt(running_agent, interrupt_reason)
_process_task_id = getattr(running_agent, "_gateway_turn_process_task_id", "")
_process_baseline = getattr(running_agent, "_gateway_turn_process_baseline", None)
# Bump the generation BEFORE scheduling the reap thread and capture the post-bump value:
# task_id is session-scoped, so a replacement turn spawning before the reap runs bumps it
# again and the closure sees a stale generation and skips — the replacement's own baseline
# covers its cleanup, so nothing stays unreaped.
_generation_at_interrupt = self._invalidate_session_run_generation(session_key, reason=invalidation_reason)
if _process_task_id and _process_baseline is not None:
threading.Thread(
target=_reap_gateway_turn_processes,
args=(_process_task_id, _process_baseline),
kwargs={
"source": "gateway_turn_interrupt",
"is_still_current": lambda: self._is_session_run_current(session_key, _generation_at_interrupt),
},
name=f"gateway-turn-reaper-{_process_task_id[:12]}",
daemon=True,
).start()
adapter = self._adapter_for_source(source)
interrupt_session_activity = getattr(type(adapter), "interrupt_session_activity", None)
if adapter and callable(interrupt_session_activity):
metadata = self._thread_metadata_for_source(source)
if _accepts_keyword(interrupt_session_activity, "metadata"):
await adapter.interrupt_session_activity(session_key, source.chat_id, metadata=metadata)
else:
await adapter.interrupt_session_activity(session_key, source.chat_id)
if adapter and hasattr(adapter, "get_pending_message"):
adapter.get_pending_message(session_key) # consume and discard
if state is not None:
state.persistent.pending_command_text = None
if release_running_state:
self._release_running_agent_state(session_key)
# Evict the cached agent: ``_interrupt_requested`` is only cleared by the turn finalizer,
# so on a hung/still-draining run the flag survives and silently kills the session's NEXT
# message (interrupted=True, api_calls=0, empty response). Like /new and /model, the next
# message rebuilds from history; the old agent keeps its flag so a hung drain still dies.
# See #44212.
self._evict_cached_agent(session_key)
async def _refresh_agent_cache_message_count(self, session_key: str, session_id: Optional[str]) -> None:
"""Re-baseline a cached agent's stored message_count after THIS turn — the coherence guard
rebuilds on mismatch, so without this every turn would rebuild and destroy prompt caching.
Only the count is refreshed, only if the same agent is still cached. DB errors leave the
snapshot as-is (one spare rebuild).
But the snapshot is taken at agent-BUILD time — before this turn writes its own user + assistant (+
tool) rows — and the cache entry is never rewritten on a reuse. See #45966.
"""
from gateway.run import _AGENT_PENDING_SENTINEL
_cache_lock = getattr(self, "_agent_cache_lock", None)
_cache = getattr(self, "_agent_cache", None)
if self._session_db is None or not session_id or not _cache_lock or _cache is None:
return
try:
_sess_row = await self._session_db.get_session(session_id)
_live = _sess_row.get("message_count", 0) if _sess_row else None
except Exception:
return
if _live is None:
return
with _cache_lock:
cached = _cache.get(session_key)
# Only re-baseline a live 3-tuple entry; skip pending sentinels, legacy 2-tuples (they opt
# out of the guard), and entries evicted/rebuilt mid-turn. A snapshot taken for a different
# session_id (same session_key, different conversation) is a different DB row — leave it.
if not (isinstance(cached, tuple) and len(cached) > 2 and cached[0] is not _AGENT_PENDING_SENTINEL):
return
_snapshot_sid = cached[3] if len(cached) > 3 else None
if (_snapshot_sid is not None and _snapshot_sid != session_id) or cached[2] == _live:
return
# Legacy 3-tuple keeps its 3-element shape for callers indexing ``cached[2]``.
_cache[session_key] = (cached[0], cached[1], _live) + (() if _snapshot_sid is None else (_snapshot_sid,))
def _set_pending_turn_sidecar_notes(self, session_key: str, notes: List[str]) -> None:
"""Stage per-turn must-deliver notes for the next agent run (one-shot)."""
if not session_key or not notes:
return
self._session_state(session_key).conversation.sidecar_notes = list(notes)
def _consume_pending_turn_sidecar_notes(self, session_key: str) -> List[str]:
state = self._peek_session_state(session_key) if session_key else None
if state is None:
return []
staged, state.conversation.sidecar_notes = state.conversation.sidecar_notes, []
return list(staged) if isinstance(staged, list) else []
def _voice_channel_sidecar_note(self, event, source: SessionSource, session_key: str) -> Optional[str]:
"""``[Voice channel now: ...]`` note when VC state changed; ``None`` when unchanged so per-turn
member/speaking churn can't touch the prompt."""
if source.platform != Platform.DISCORD:
return None
adapter = self.adapters.get(Platform.DISCORD)
guild_id = self._get_guild_id(event)
if not (guild_id and adapter and hasattr(adapter, "get_voice_channel_context")):
return None
try:
vc_now = adapter.get_voice_channel_context(guild_id) or ""
except Exception:
logger.debug("voice-channel context read failed", exc_info=True)
return None
vc_prev = None
if session_key:
_vc_state = self._session_state(session_key)
vc_prev, _vc_state.conversation.vc_last = _vc_state.conversation.vc_last, vc_now
if vc_now == (vc_prev if vc_prev is not None else ""):
return None
return f"[Voice channel now: {vc_now or 'not connected to a voice channel'}]"
def _pinned_session_context_prompt(self, context, redact_pii: bool, session_key: Optional[str]) -> str:
"""Session-context prompt pinned per session: key hit → pinned bytes reused VERBATIM (immune
to renderer nondeterminism); key miss → re-render and re-pin (rename, topic edit, /sethome)."""
_eph_key = self._ephemeral_change_key(context, redact_pii)
_pin_state = self._peek_session_state(session_key) if session_key else None
_eph_pin = _pin_state.conversation.ephemeral_pin if _pin_state else None
if _eph_pin is not None and _eph_pin[0] == _eph_key:
return _eph_pin[1]
text = build_session_context_prompt(context, redact_pii=redact_pii)
if session_key:
self._session_state(session_key).conversation.ephemeral_pin = (_eph_key, text)
return text
@staticmethod
def _ephemeral_change_key(context, redact_pii: bool) -> str:
"""Hash the exact inputs ``build_session_context_prompt`` renders. Invariant
(test_prompt_tail_freeze.py): any input whose change alters the rendered bytes MUST appear
here — omission means a stale pinned prompt; extras only re-render."""
import hashlib
src = context.source
def _s(v) -> str:
return str(v or "")
discord_ids: tuple = ()
discord_tools = ""
if src.platform == Platform.DISCORD:
from gateway.session import _discord_tools_loaded
discord_tools = "1" if _discord_tools_loaded() else "0"
# message_id: only PRESENCE is rendered (the id itself arrives per-turn in the user
# message) — keying on the value would re-render every message for zero byte change.
discord_ids = (
_s(src.guild_id), _s(src.parent_chat_id), _s(src.thread_id), _s(src.chat_id),
"1" if src.message_id else "0",
)
# Slack's capability-aware platform note is gated on _slack_tools_loaded() — the gate state must
# be in the key (same parity contract as the Discord gate above) so a config / MCP-registration
# flip re-renders once instead of serving a stale pinned note for the rest of the session.
slack_tools = ""
if src.platform == Platform.SLACK:
from gateway.session import _slack_tools_loaded
slack_tools = "1" if _slack_tools_loaded() else "0"
try:
from hermes_constants import display_hermes_home
home_display = str(display_hermes_home())
except Exception:
home_display = ""
key_tuple = (
src.platform.value if src.platform else "",
_s(src.chat_id), _s(src.thread_id), _s(src.chat_type), _s(src.chat_name), _s(src.chat_topic),
_s(src.user_name), _s(src.user_id), _s(getattr(src, "profile", None)),
bool(context.shared_multi_user_session), discord_ids, discord_tools, slack_tools,
tuple(p.value for p in context.connected_platforms),
tuple(
(p.value, _s(getattr(hc, "name", "")), _s(getattr(hc, "chat_id", "")))
for p, hc in context.home_channels.items()
),
bool(redact_pii), home_display,
)
return hashlib.sha256(repr(key_tuple).encode("utf-8")).hexdigest()
def _evict_cached_agent(self, session_key: str) -> None:
"""Remove a cached agent (/new, /model, ...) and soft-release its LLM client pool (AIAgent
holds reference cycles; without it RSS grows across /new). Soft = frees clients and child
subagents but PRESERVES terminal sandbox / browser / bg processes since the session may
resume; true boundaries call ``_cleanup_agent_resources`` first. Cleanup runs on a daemon
thread so ``_agent_cache_lock`` never spans slow socket teardown.
Pops the entry AND soft-releases the evicted agent's LLM client pool so the httpx connection
(sockets + held buffers) is freed promptly rather than waiting on CPython GC — AIAgent holds
reference cycles (callbacks, tool state) that delay refcount collection, so a manual release is
required to keep gateway RSS flat across many /new, /model, undo and reset operations (#29298, same
leak class as #25315).
"""
from gateway.run import _AGENT_PENDING_SENTINEL
# Prompt-stability state rides the agent-cache lifecycle: a fresh agent must re-render its
# session-context bytes (the pin) and re-see the current voice-channel state once.
state = self._peek_session_state(session_key)
if state is not None:
state.conversation.ephemeral_pin = None
state.conversation.vc_last = None
# Tests build runners with ``_agent_cache_lock = None``; evict lock-free then. With the lock
# present ``_agent_cache`` is read directly (an initialized runner always has it).
_lock = getattr(self, "_agent_cache_lock", None)
evicted = None
if _lock:
with _lock:
evicted = self._agent_cache.pop(session_key, None)
else:
_cache = getattr(self, "_agent_cache", None)
if _cache is not None:
evicted = _cache.pop(session_key, None)
agent = _first_agent(evicted)
# Never tear down an agent that's mid-turn — its client, sandbox and child subagents are in use.
if agent is None or agent is _AGENT_PENDING_SENTINEL or id(agent) in self._running_agent_ids():
return
self._spawn_release_thread(
self._release_evicted_agent_soft, (agent,), f"agent-evict-{str(session_key)[:24]}", inline_fallback=True,
)
def _spawn_release_thread(self, target, args: tuple, name: str, *, inline_fallback: bool) -> None:
"""Run a release on a daemon thread. ``inline_fallback`` runs it inline (best-effort) when no
thread can start (interpreter shutdown); otherwise a spawn failure propagates, as on main."""
try:
threading.Thread(target=target, args=args, daemon=True, name=name).start()
except Exception:
if not inline_fallback:
raise
with suppress(Exception):
target(*args)
def _commit_memory_before_soft_evict(self, agent: Any, key: str) -> None:
"""Commit the live transcript to memory providers before resource-only eviction."""
# No external memory provider (``_memory_manager`` None) — nothing to commit.
if agent is None or not hasattr(agent, "commit_memory_session") or getattr(agent, "_memory_manager", None) is None:
return
try:
messages = getattr(agent, "_session_messages", None)
agent.commit_memory_session(messages if isinstance(messages, list) else None)
logger.debug(
"Committed on_session_end extraction before soft-evicting "
"session=%s (resource eviction)", key,
)
except Exception as _e:
logger.debug("Pre-evict memory commit failed for %s: %s", key, _e)
def _commit_then_release_soft(self, agent: Any, key: str) -> None:
"""Commit end-of-session memory (if warranted), then soft-release — on the daemon eviction
thread. Order matters: commit needs the live memory manager before ``release_clients``."""
self._commit_memory_before_soft_evict(agent, key)
self._release_evicted_agent_soft(agent)
def _release_evicted_agent_soft(self, agent: Any) -> None:
"""Soft cleanup for cache-evicted agents: unlike _cleanup_agent_resources, the session may
resume, so terminal sandbox, browser daemon and bg processes outlive the AIAgent instance."""
if agent is None:
return
with suppress(Exception):
if hasattr(agent, "release_clients"):
agent.release_clients()
else:
# Older agent instance (shouldn't happen in practice) — legacy full-close path.
self._cleanup_agent_resources(agent)
# Free conversation history — tens of MB of tool output on heavy 100+-tool-call sessions.
# release_clients() preserves session tool state for resume, but the message list is rebuilt from
# persisted session JSON on the next turn, so dropping it here is safe.
if hasattr(agent, "_session_messages"):
agent._session_messages = []
# _db_flush_scan_prefix (run_agent.py, stamped on every successful flush) is a shallow copy
# sharing every message dict of the flushed transcript, so leaving it pins the multi-MB strings
# this eviction frees. Pressure-evictable agents have flushed by definition, so it's populated.
if hasattr(agent, "_db_flush_scan_prefix"):
agent._db_flush_scan_prefix = None
def _agent_cache_bounds(self):
"""Operator-configured agent-cache bounds, resolved once per process (lazily, not in
``__init__``, so ``__new__``-constructed test / slash-command runners work too)."""
from gateway.run import _load_gateway_config
bounds = getattr(self, "_agent_cache_bounds_cache", None)
if bounds is None:
from gateway.agent_cache_pressure import resolve_agent_cache_bounds
try:
bounds = resolve_agent_cache_bounds(_load_gateway_config())
except Exception as _e:
logger.debug("Agent cache bounds config read failed: %s", _e)
# Resolve from an empty config rather than bare AgentCacheBounds(): the dataclass default
# has memory_high_mb=None (pressure pass OFF) but an *absent* section means "auto" — a
# transient config read failure must not permanently disable the OOM valve.
bounds = resolve_agent_cache_bounds({})
self._agent_cache_bounds_cache = bounds
return bounds
def _agent_cache_cap(self) -> int:
"""Effective LRU cap — the configured override, else the default."""
from gateway.run import _AGENT_CACHE_MAX_SIZE
return self._agent_cache_bounds().max_size or _AGENT_CACHE_MAX_SIZE
def _agent_cache_idle_ttl(self) -> float:
"""Effective idle TTL in seconds — configured override, else default."""
from gateway.run import _AGENT_CACHE_IDLE_TTL_SECS
return self._agent_cache_bounds().idle_ttl_secs or _AGENT_CACHE_IDLE_TTL_SECS
def _sweep_agent_cache_under_pressure(self) -> int:
"""Shed cached transcripts once the gateway heap nears its budget; returns count evicted.
The LRU cap counts entries and the idle sweep counts seconds; neither knows one cached agent
pins a full ``_session_messages`` transcript (tens of MB), so RSS climbs until the cgroup
throttles. Above the anonymous-RSS budget this soft-evicts LRU agents (transcript rebuilt
from the persisted session next turn). Never touched: agents mid-turn, the most recently
used sessions, and transcripts not yet on disk.
A gateway serving many chats can hold every warm transcript for the TTL window.
Pressure eviction bounds that heap before the cgroup throttles and SIGTERM can
no longer flush inside systemd's stop timeout (#80764).
"""
from gateway.run import _AGENT_PENDING_SENTINEL
from gateway.agent_cache_pressure import (
plan_pressure_evictions, read_anon_rss_mb, transcript_persistence_caught_up
)
bounds = self._agent_cache_bounds()
_cache = getattr(self, "_agent_cache", None)
_lock = getattr(self, "_agent_cache_lock", None)
# Nothing cached — whatever is using the heap, it isn't us, and warning about it every tick
# would point at the wrong subsystem.
if not bounds.memory_high_mb or not _cache or _lock is None:
return 0
rss_mb = read_anon_rss_mb()
if rss_mb is None or rss_mb < bounds.memory_high_mb:
return 0
running_ids = self._running_agent_ids()
def _is_live(agent: Any) -> bool:
return agent is not None and agent is not _AGENT_PENDING_SENTINEL and id(agent) not in running_ids
def _is_evictable(key: str, agent: Any) -> bool:
return _is_live(agent) and transcript_persistence_caught_up(agent)
with _lock:
ordered = [(key, _first_agent(entry)) for key, entry in _cache.items()]
plan = plan_pressure_evictions(
ordered, is_evictable=_is_evictable, max_evictions=bounds.max_evictions_per_pass,
protect_recent=bounds.protect_recent,
)
for key, _ in plan:
_cache.pop(key, None)
if not plan:
_mid_turn = sum(1 for _, a in ordered if a is not None and id(a) in running_ids)
_unflushed = sum(1 for _, a in ordered if _is_live(a) and not transcript_persistence_caught_up(a))
logger.warning(
"Agent cache pressure: anon RSS %dMB over budget %dMB but no "
"evictable session (%d cached, %d mid-turn, %d blocked on "
"un-flushed persistence)%s",
rss_mb, bounds.memory_high_mb, len(ordered), _mid_turn, _unflushed,
(
" — transcripts are not reaching the session DB "
"(session persistence disabled or failing?); the memory "
"valve cannot shed sessions until they persist."
if _unflushed and not _mid_turn
else " — memory will keep climbing until those turns finish."
),
)
return 0
evicted_count = len(plan)
logger.warning(
"Agent cache pressure: anon RSS %dMB over budget %dMB — evicting %d LRU session(s): %s",
rss_mb, bounds.memory_high_mb, evicted_count, ", ".join(key for key, _ in plan),
)
try:
threading.Thread(target=self._release_pressure_batch, args=(plan,), daemon=True,
name="agent-cache-pressure").start()
except Exception:
# Thread spawn failed (interpreter shutdown): release inline, unguarded (as on main).
self._release_pressure_batch(plan)
# _release_pressure_batch drains `plan` in place (so the trim runs with no lingering agent
# refs) — len(plan) is 0 once the daemon thread finishes, hence the pre-captured count.
return evicted_count
def _release_pressure_batch(self, plan: List[tuple]) -> None:
"""Release a pressure-evicted batch sequentially on one daemon thread, then ``malloc_trim`` so
RSS actually falls. The plan is drained (``pop`` + ``del``), not iterated, so no local
reference pins evicted agents during ``gc.collect`` + trim (else the valve over-evicts)."""
while plan:
key, agent = plan.pop(0) # FIFO — evict LRU-first order preserved
try:
self._commit_then_release_soft(agent, key)
except Exception as _e:
logger.debug("Pressure release failed for %s: %s", key, _e)
del agent
with suppress(Exception):
from hermes_cli.mem_trim import trim_memory
trim_memory(force=True, reason="agent_cache_pressure")
def _enforce_agent_cache_cap(self) -> None:
"""Evict oldest cached agents past the LRU cap (requires _agent_cache_lock); cleanup on a
daemon thread. Mid-turn agents are SKIPPED, so the cache may stay over cap until the next
insert."""
_cache = getattr(self, "_agent_cache", None)
# OrderedDict.popitem(last=False) pops oldest; plain dict lacks the arg so skip enforcement
# if a test fixture swapped the cache type.
if _cache is None or not hasattr(_cache, "move_to_end"):
return
# Snapshot of agent instances mid-turn, keyed by id() so lookup is O(1) and independent of
# AIAgent.__eq__ (which MagicMock overrides in tests).
running_ids = self._running_agent_ids()
# Walk LRU → MRU; only the first (size - cap) LRU positions are candidates. An active slot is
# SKIPPED rather than evicting a newer entry — that would penalise a fresh session (no cache
# history) to protect a long-running one. Cache may stay over cap until the next insert.
cap = self._agent_cache_cap()
candidates = [(key, _tuple_agent(_cache.get(key))) for key in list(_cache.keys())[:max(0, len(_cache) - cap)]]
evict_plan = [(key, agent) for key, agent in candidates if agent is None or id(agent) not in running_ids]
for key, _ in evict_plan:
_cache.pop(key, None)
remaining_over_cap = len(_cache) - cap
if remaining_over_cap > 0:
logger.warning(
"Agent cache over cap (%d > %d); %d excess slot(s) held by "
"mid-turn agents — will re-check on next insert.",
len(_cache), cap, remaining_over_cap,
)
for key, agent in evict_plan:
logger.info("Agent cache at cap; evicting LRU session=%s (cache_size=%d)", key, len(_cache))
if agent is not None:
# Commit end-of-session memory, then soft-release, both on the daemon thread so the
# (possibly network-bound) provider call never blocks the held cache lock.
self._spawn_release_thread(self._commit_then_release_soft, (agent, key), f"agent-cache-evict-{key[:24]}", inline_fallback=False)
def _sweep_idle_cached_agents(self) -> int:
"""Evict cached agents idle past the idle TTL (lock acquired internally; cleanup on daemon
threads; mid-turn agents SKIPPED); returns the number evicted."""
_cache = getattr(self, "_agent_cache", None)
_lock = getattr(self, "_agent_cache_lock", None)
if _cache is None or _lock is None:
return 0
now = time.time()
idle_ttl = self._agent_cache_idle_ttl()
to_evict: List[tuple] = []
running_ids = self._running_agent_ids()
with _lock:
for key, entry in list(_cache.items()):
agent = _tuple_agent(entry)
if agent is None or id(agent) in running_ids:
continue # mid-turn — don't tear it down
last_activity = getattr(agent, "_last_activity_ts", None)
if last_activity is None or (now - last_activity) <= idle_ttl:
continue
to_evict.append((key, agent))
for key, _ in to_evict:
_cache.pop(key, None)
for key, agent in to_evict:
logger.info("Agent cache idle-TTL evict: session=%s (idle=%.0fs)", key, now - getattr(agent, "_last_activity_ts", now))
self._spawn_release_thread(self._commit_then_release_soft, (agent, key), f"agent-cache-idle-{key[:24]}", inline_fallback=False)
return len(to_evict)