Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.
Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
(_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
explicit operator stops (process_manage kill, /stop slash + RPC mirror,
CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
the host is going away and survivors would become PPID=1 orphans
Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
The anthropic.com-only refresh guard stopped token rotation on hosts the
resolver itself accepts as native Anthropic (*.claude.com, /anthropic
proxies), which already receive the Anthropic token at startup. Long
sessions there hit 401 and the recovery refresh was refused too.
Official hosts (anthropic.com, claude.com) rotate as before; any other
host rotates only when the current key is already an Anthropic
credential (sk-ant- or OAuth), so the #17829 case (custom key swapped
for ANTHROPIC_API_KEY / the OAuth token) stays blocked. Azure keeps its
static-key exclusion. Also corrects the _resolve_openrouter_runtime
docstring, which still claimed OPENAI_BASE_URL is never consulted.
_try_refresh_anthropic_client_credentials runs before every request and on
401 for provider 'anthropic'. A URL-bearing model alias under the built-in
'anthropic' label (#28660) points that provider at a foreign host with no
key; the pre-request refresh then swapped in ANTHROPIC_API_KEY / the OAuth
token and sent it to that host.
The third-party guard above matched "anthropic.com" as a substring of the
URL, so a host like 127.0.0.1:8080/anthropic.com still passed as native and
got the refresh. Refresh only when the endpoint's hostname is Anthropic's own
(or no endpoint is set).
Found by the C11 routing truth table (tests/e2e/core/tenancy): a /model
switch to such an alias carried the vendor key to the alias host.
With provider 'anthropic' pointed at a third-party Anthropic-compatible
endpoint, _try_refresh_anthropic_client_credentials only skipped Azure and
otherwise re-resolved native Anthropic credentials (ANTHROPIC_API_KEY, the
stored OAuth token) and rebuilt the client with them, so the endpoint got
a credential it was never given. Skip the refresh for any endpoint
build_anthropic_client already classifies as third-party.
Ported from run_agent.py onto agent/client_lifecycle.py, where the method
lives now; the separate Azure test is dropped (Azure is one such endpoint).
Salvaged from #17829.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).
Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-up to the salvaged #114470 commit.
`AIAgent.close()` hands `_close_task_resources()` the agent's session_id, but the
file tools key `FileStateRegistry` by the per-turn task_id: cron runs use
`cron:<job>:<uuid>` while their session_id is `cron_<job>_<ts>`, and delegate
children use `subagent-N-xxxx` against a fresh uuid session. `cleanup_vm(session_id)`
therefore never reached `forget_task()` for the id that owns the read stamps and
writer claims, so the purge added to `forget_task()` had no production caller for
exactly the lifecycles the issue describes. `close()` now runs
`clear_file_ops_cache()` for every id in `_process_owner_task_ids` (the set the
turn context already maintains for process ownership) before dropping the session.
Dropped from #114470: the `HERMES_FILE_STATE_WRITER_TTL` env knob and the
time-based eviction in `check_stale()`. With the lifecycle end actually releasing
the finished task's claims, a TTL only weakens the concurrent case the guard exists
for (a live sibling's hour-old write is still a real conflict), and behavioural
env vars are not a config surface. The module-level `forget_task()` wrapper and
its `__all__` entry from the salvaged commit are kept.
Tests trimmed to two invariants in the mirroring file: `forget_task()` purges the
finished task's writer claims (a live sibling still fires), and `AIAgent.close()`
releases the file state of every task id it ran even though it receives the
session_id.
Follow-up to the salvaged #111787 commit (@KoNit-K), same mechanism as #75578 (@adikpb).
What:
- Move the per-model cooldown logic out of the 2.7k-line credential_pool.py facade into
agent/credential_pool_model_cooldowns.py (mixin + module helpers), keeping only the
select/has_available/next_available_at hooks and the mark_exhausted_and_rotate branch in the facade.
- The model cooldown uses the same TTL policy as a credential-wide 429 (_exhausted_ttl: provider
reset_at wins, a sole credential keeps its 60s bench) instead of a flat 1h, so a single-credential
user is never benched LONGER for the failed model than before.
- Drop the `quota_scope == "account"` check: nothing in the tree produces that key.
- Also keep billing_unverified 429s credential-wide, matching the credential-wide branch.
- _rebind_primary_credential_pool read `rt` that was not in its scope (NameError on every
post-fallback restore); pass primary_model from the caller instead.
- resolve_anthropic_token(model=...) gates only model-aware callers; model-less diagnostics
(usage display, model discovery) keep the key as before.
- _anthropic_token_or_raise names the cooled model instead of claiming no credentials exist.
- Simplify the salvaged call sites (unconditional select(model=)/resolve_anthropic_token(model=)),
widen test stubs that lacked the new kwargs, trim the tests to two invariants on a real temp store.
- Docs: credential-pools.md documents per-model Anthropic 429 cooldowns.
Why: a generic Anthropic 429 is a per-model rate limit; benching the whole credential took every
other Claude model offline while the env/borrowed token path handed the same benched key straight
back (#111769, #61451).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: adikpb <67222969+adikpb@users.noreply.github.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up to 564aef2946 (Mantle SigV4 on /model). The same class of bug covered the
other two Bedrock wires and two more rebuild paths:
- Claude on Bedrock (anthropic_messages): startup builds an AnthropicBedrock SDK
client (SigV4 via boto3). /model, fallback-to-Bedrock and fallback restore built a
plain Anthropic client with api_key="aws-sdk" against bedrock-runtime → 401/403.
- Converse models (bedrock_converse: Nova, DeepSeek, Llama): only agent_init set
_bedrock_region / _bedrock_guardrail_config. After a rebuild the transport fell
back to us-east-1 and guardrail_config=None, so eu-/ap- users hit the wrong region
and configured Guardrails silently dropped. switch_model also built a pointless
OpenAI client against bedrock-runtime.
Introduce bedrock_adapter.bind_bedrock_runtime(agent, base_url, api_mode) as the one
place that puts an agent on a non-Mantle Bedrock wire, and call it from agent_init
(replacing the two duplicated bodies), _build_switched_client, _rebuild_primary_client
and _swap_fallback_clients. try_recover_primary_transport had a hand-copied version of
_rebuild_primary_client's ladder (and would have hit the same gap); it now calls the
shared helper. The region/guardrail parsing that lived in agent_init and
runtime_provider_backends is now bedrock_region_from_runtime_url /
bedrock_guardrail_config in the adapter.
Live probe on origin/main across switch_model, restore_primary_runtime and
_swap_fallback_clients for both wires: 0/6 correct before, 6/6 after.
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping
A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.
* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin
* docs(user-guide): free tier and signing in
New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.
* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account
Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.
* fix(model): switching off the free tier points at signing in, never hops providers
* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once
* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token
* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check
Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.
* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors
The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.
* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off
* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check
The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.
* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure
A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.
* fix(auth): decide the Nous model together with the route on every credential-pool swap
The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.
* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start
* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died
The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.
* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential
The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.
* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only
A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Independent review of the first fix found a credential-identity takeover:
_adopt_nous_key_before_expiry() called the singleton resolver
unconditionally, so an agent running on an explicitly supplied (or
pool-selected) account-A key near expiry was moved onto the logged-in
account-B key from auth.json before the A key had even failed, and the next
real SDK request went out as B. That silently changes who is billed.
The adoption now reads the `sub` claim of the key in hand and passes it as
require_account; _try_refresh_nous_client_credentials refuses any
replacement whose `sub` differs (logged at INFO, current key kept). A key
with no `sub` is not adopted proactively at all. The reactive 401 path is
unchanged, and same-account adoption (the keepalive's or a peer's fresh key)
still works.
Verified with the reviewer's own live probe (real AIAgent -> real
prepare_iteration -> real auth.json transaction -> real OpenAI SDK ->
loopback capture): explicit-account case account-A -> account-A,
adopted_fresh=False (was A -> B); same-account still adopts; 12 concurrent
near-expiry agents still 0 x 401.
Tests (2 new): a fresh key for a different account is never adopted; a key
without an account claim is left alone without touching the store.
The Nous agent key lives 3,599 s. In the 1,393-agent refactor run every
in-process agent learned about the hourly expiry from its own 401: 620
authentication_error 401s in the logged window (177 in one hour), each a
failed attempt the model never saw, and the credential pool benched the
sole credential for all of them at once. At 08:30 the storm took the
parent process down.
Two gaps. The proactive refresher (hermes_cli/nous_auth_keepalive.py) is
started only by the gateway and the web server; the CLI process, and every
subagent built inside it, never started it. And even with a fresh key in
the store, nothing adopted it before a request: a request went out with
whatever key the agent was constructed with until it 401'd.
Now _finalize_routing starts the keepalive (idempotent, process-wide,
daemon) whenever an agent resolves onto provider "nous", and
prepare_iteration calls _adopt_nous_key_before_expiry(): the agent key is
a JWT, its exp is read locally, and inside a 180 s skew the store is
re-read under the auth-store lock with force_refresh=False, so the
keepalive's (or a peer's) fresh key is adopted without a POST; when none
exists, ONE refresh runs there instead of N reactive ones after N 401s.
_try_refresh_nous_client_credentials no longer rebuilds the client when
the store returns the key already in hand.
Live A/B (local server: 401s any bearer but FRESH; store patched to hold
FRESH; 40 agents holding a JWT that expires in 30 s fire concurrently):
main 40 x 401 then recover, branch 0 x 401.
Tests (4): far from expiry the store is not touched; inside the skew the
store's fresh key is adopted with force_refresh=False; the same key back
from the store is not re-adopted; a real AIAgent routed to nous starts
the keepalive and one routed to openrouter does not.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
Adolanium review §3 (Medium/Low). BASE shape at every steer/redirect/persist/agent-cache
lock site was: lock = getattr(obj, "_x_lock", None); if lock is not None: with lock: <direct
attribute read>; else: <getattr fallback for object.__new__ test stubs>. The refactor rewrote
several of these as 'with getattr(...) or nullcontext(): getattr(slot, None)', which (a) turned a
missing slot under the lock from a loud AttributeError into silent None and (b) in the gateway
peek helper read the cache without any lock when the lock attribute was absent.
Restored BASE semantics at:
- agent/agent_runtime_helpers.py::_requeue_pending_steer
- agent/interrupt_control.py: steer, redirect, clear_interrupt, _has_pending_redirect,
_drain_pending_redirect, _drain_pending_steer (new _ic_slot helper: direct read under lock)
- agent/session_persistence.py::_persist_lock (explicit None check + BASE rationale)
- gateway/slash_commands.py::_cached_agent_for (BASE callers read ONLY under the lock; no lock -> None)
- gateway/run_agent_cache.py::_evict_cached_agent (BASE: self._agent_cache direct under lock)
- agent/client_lifecycle.py::_is_openai_client_closed: BASE body + docstring verbatim (outer
is_closed first; inner _client.is_closed only when _client exists; else False)
- agent/stream_delivery.py::_ensure_stream_writer_state: restore BASE rationale that the lock is
created unconditionally in agent_init (_STREAM_STATE) and the lazy path is stub-only
A/B (/tmp/rf/rev/ab_lock_fallbacks.py, ab_client_closed.py) is byte-identical BASE vs HEAD on
the reviewer's inputs + edge cases. tests/agent/test_lock_fallback_base_semantics.py pins it
(7 of 25 cases fail on the pre-fix tree).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.