When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:
gateway waits on all in-flight work units (#77184 don't-amputate)
-> cron agent session waits on the \hermes update\ process to exit
-> \hermes update\ waits on the gateway to exit [back to A]
The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).
Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.
Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.
Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
(the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).
Fixes#100179
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).
- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
scheduled_at, dispatched_at, lateness_seconds, and kind
(on_time / late / catch_up, classified against the ticker tolerance and
the schedule's catch-up grace window). Manual triggers and one-shots are
not stamped (no scheduled instant to be late against / retired beyond
grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
catch-up, in both the built-in ticker and external provider paths.
CLI surface only — no new tools, no policy engine.
Addresses the visibility half of #99879.
- Translate leftover English 'toolset(s)' strings in ru.ts
- Cover ru aliases/config value in languages.test.ts (parity with ar)
- List Russian among desktop UI languages in website/docs/user-guide/desktop.md
Full desktop UI translation (ru.ts, 3076 string/function leaves,
mirrors en.ts 1:1) plus locale registration in types, catalog and
language list with aliases (ru, ru-ru, ru_ru, ru-by).
Russian plurals (1 / 2-4 / 5+, 11-14 exception) via RU_PLURAL/RU_NOUN
helpers; count accepts number | string to match en.ts signatures.
Verified against current main: structural validation 3076/3076 with
placeholder parity, typecheck (renderer/electron/e2e) clean,
i18n vitest 28/28, prettier + eslint clean.
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
call's post_api_request payload (per-attempt reset in the loop)
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.
Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.
Backward compatible: existing hook subscribers ignore unknown kwargs.
The spawn-time output tail (#93608) puts child stdout into flowing mode at
spawn. Both backend spawn paths in main.ts then await claimBackendChild
(whose Windows Get-Process probe cold-starts in 2-8s) and boot-progress IPC
BEFORE waitForDashboardPortAnnouncement attaches its stdout listener. Node
streams never replay consumed chunks to late listeners, so a READY sentinel
printed during that window was lost forever — the wait hit its 90s timeout
and a healthy backend was killed (deterministic on Windows, racy on
macOS/Linux; still firing on v0.21.0 incl. concurrent multi-profile boots).
Fix (belt and suspenders, both spawn paths — primary and profile pool):
- create the port-announcement promise immediately after spawn, before any
await
- new bufferedOutput option on waitForDashboardPortAnnouncement: after
attaching its own listener, waitForDashboardPort scans the output tail's
already-buffered text for the sentinel, making listener-attach ordering
irrelevant regardless of call-site shape
The readyFile path was already ordering-safe (it polls a file, not the
stream). Approach follows stale PR #60986 by @ParaWheeler, rebased onto the
output-tail/readyFile plumbing added since.
Fixes#60323
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.
Fixes#97343
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.
Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.
Fixes#90835
contributors/emails/agent@Agents-Mac-mini.local collides on
case-insensitive filesystems (macOS/Windows checkouts) with its
lowercase sibling, breaking clean checkouts. The identity is a
local-machine artifact, not a real contributor address.
Fixes#88168Fixes#100047Fixes#99966Fixes#99821
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.
Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.
Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.
Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.
Fixes#87694
Salvaged from #87745 with expiry mitigation added.
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.
Scope now travels with the unit end to end:
- the in-process systemd loop records a scope-qualified twin of
`restarted_services` (`restarted_scoped_units`) while the bare-name list
keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
stays unqualified and is read as scope-agnostic, and an unrecognized
scope drops the skip rather than honouring it: dropping a skip can only
cost one more restart-and-verify, honouring an unreadable one can leave
a stale generation running.
Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.
Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.
Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.
Refs #92145
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.
The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.
The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.
- restart active `hermes-serve*` systemd units from the fresh child,
enumerated from systemd rather than from the misclassifying inventory,
and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
process, and never kill one -- a manual or Desktop-owned serve has no
relaunch authority;
- require every runtime family, not just the gateway leg, before a
fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
Three kills at the shared chokepoints:
1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
while the active config.yaml is corrupt (AuthError code=corrupt_config).
A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
model.provider and tier-3/4 silently adopted the PAID openrouter provider
against the user's real (unparseable) intent. New probe:
hermes_cli.config.get_active_config_parse_failure(), recorded in the
existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
a fixed file clears the block immediately. Explicit provider requests
are untouched.
2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
(nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
honored untouched (paid-lane warning retained).
3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
process per provider) when a credential is newly ingested — ingestion
itself stays allowed.
Fixes#81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
on every surface
Simplify-pass follow-ups on the salvaged alias logic:
- _to_oauth_wire_name's allow_alias kwarg was never passed False by any
call site — removed.
- The _claimed_wire_names set comprehension re-implemented the same
mcp_/mcp__/bare normalization ladder as _to_oauth_wire_name; both now
share _normalize_to_mcp_wire. Behavior identical: aliased names are
bare, so the collision probe's _MCP_TOOL_PREFIX + aliased equals
_normalize_to_mcp_wire(aliased).
Anthropic subscription OAuth (claude_code credential) misroutes Hermes
sessions carrying the session_search or memory toolset into the
extra-usage lane, surfacing as HTTP 400 "You're out of extra usage" on
a valid subscription. Live-verified via the
anthropic-ratelimit-unified-representative-claim response header
(deterministic lane oracle, no dependency on the laggy usage counter):
tool schemas are innocent, the trigger is three specific system-prompt
sentences (session_search recall + two skill_manage sentences),
required jointly — breaking any one clears the classifier.
Two independent layers:
1. OAuth wire alias (anthropic_adapter.py, transports/anthropic.py):
session_search -> chat_history_lookup, memory -> context_notes in
tool name, description, and (session_search only) system-prompt
prose, with wire-collision guarding and a normalize_response
reverse-map that keeps GH-25255 registered-tool precedence. Also
routes named tool_choice through the same normalizer, closing a gap
where a forced tool_choice would leak the raw trigger string and
stop matching tools[].
2. Prompt-preserving reword (prompt_builder.py): rewords the two
triggering SKILLS_GUIDANCE sentences while keeping the same
meaning, still naming skill_manage, and leaving the Skill Safety
Rule section untouched. Applies to all auth paths since it's a
prompt-copy change, not a wire-level transform.
Two layers rather than one because the three-sentence AND-condition
means a classifier tightening could start firing on either remaining
leg alone.
Fixes#65365
The base fake still stubbed the OLD cleanup surface (shutdown_memory_provider
/ close). Production now calls release_clients(); on fakes without it the
AttributeError is swallowed by the cleanup's except-Exception, so those
tests silently stopped exercising the cleanup path. The inner recording
fake in the renamed test keeps its close() stub deliberately — it's the
tripwire proving close() is never called on the shared session.
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):
- Gate every scope-path branch on a new _IS_LINUX constant instead of
'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
never touches systemd code — no probe subprocess, no scope argv,
byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
legacy argv byte-identical, no unit recorded) and probe-returns-False
off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
wired into the on-demand windows-venv-e2e lane): real spawn_local on
windows-latest asserting jobs run exactly as before — spawned, output
captured, exit code correct, systemd path never reached even under
faked gateway identity.
Refs #70716, #71378.
When chrome-sandbox is present but not root-owned 4755, Chromium can still
abort via setuid_sandbox_host even though the namespace sandbox works.
After the userns probe skips sudo, append --disable-setuid-sandbox so
.desktop/no-TTY launches keep the namespace sandbox without a privilege
prompt. Does not add --no-sandbox.
Fixes#51327
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).
On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.
Fixes#88032Fixes#51327Fixes#58593
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review follow-ups:
- Strip the tool_call id when populating the call-name map so it matches the
stripped lookup (and pass 1's result_call_ids). A padded id previously
skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
already dropped by the orphan pass, so it cannot reach the provider with a
stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
which only ever exercised the matching case, and add the padded-id case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.
Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.
Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The reproduction reported on #96811 is Hermes Studio's group chat, and it is
the one half of the issue that cannot be closed from inside this repository.
This states why in executable form, and pins the contract the host adoption
depends on so a later refactor cannot quietly break it.
Root cause, traced to the host. Studio reaches Hermes as a LIBRARY, not
through the gateway. Its bridge mints groupRuntimeSessionId(room, profile,
name) -- a gc_run_ prefix truncated to 96 characters plus a fresh UUID4 hex --
for every reply, writes the session row itself, and constructs AIAgent(...)
with platform / session_id / session_db and no routing identity of any kind:
no gateway_session_key, no chat_id, no user_id, no parent_session_id. So the
declared scope is unreachable, the row is a lineage root, and
resolve_prompt_cache_scope() correctly falls through to the physical id, which
moves on every reply. No stable carrier crosses the boundary. Hermes must not
recover one from the id's syntax -- that is #79017's collision class, and the
two negative controls merged via #97704 exist to keep it from trying.
What does cross the boundary is a value Studio already has:
groupBridgeSessionId(room, profile, name, sessionSeed, runtimeConfig) is
stable for one conversation in one room, already carries the room, profile,
agent name and the room-owned sessionSeed, and is already hashed and
length-bounded. Passing it as gateway_session_key is the entire adoption, and
it is a Studio-side change; this PR keeps Refs #96811 for exactly that reason.
What this suite adds is the half that IS reachable here. The existing suites
simulate Studio's id SHAPE against synthetic agent doubles; none of them runs
the host's construction path, so nothing today would fail if that path stopped
honouring a declaration. These tests build a real AIAgent in the bridge's own
order -- row written first, agent second, new physical id per reply -- and
state what the adoption buys:
- three replies with distinct physical ids hold ONE affinity scope, and the
scope carries neither the room nor the member name;
- two members of one room, and two rooms, never share a bucket;
- a new sessionSeed rotates the scope, which is Studio's own conversation
boundary and needs no reset observed by Hermes;
- equal keys under different row sources never collapse, because the row's
source -- not agent.platform -- is the identity the peer queries match on;
- a tool child and a background-review fork on the same declared key keep
their own scope, so #79161 survives this construction path too;
- and an undeclared bridge is byte-identical: per-reply scope, and a
compression rotation still walks its lineage.
Against origin/main the two declared-conversation assertions fail --
"assert 3 == 1", and the raw gc_run_ id leaking as the routing key -- which is
the reproduction; the remaining eight pass there and here, because they
describe behaviour that must not change.
Reported by @cervantesh, whose re-check of Studio main@86d0c95375 located the
missing consumer on the real path and asked for exactly this witness.
Refs #96811
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.
1. scratch/repro_96811.py is deleted. It would have landed on main as a
tracked file: scratch/ is not gitignored and has never existed on main, so
this PR was creating the directory. Nothing referenced the probe, and
TestConversationGenerationRotates / TestGenerationSurvivesPruning /
TestPeerIdentityIsSourceQualified already carry all four of its stages, so
it is dropped rather than parked under tests/.
2. Upgrade notes are written into this commit body (below) and the PR body.
There is no committed changelog to add them to: scripts/release.py
generates .release_notes.md from commit SUBJECTS at release time, and
.gitignore keeps that file out of the tree.
3. declared_conversation_scope() now reads the sessions row ONCE. The fork
verdict and the source the peer queries match on both live on that row, and
asking for them separately read it twice per resolution. The new
SessionDB.declared_scope_identity() returns the pair and keeps the marker
rules beside is_explicit_fork_child() instead of re-implementing them in the
caller. A SessionDB that does not expose the combined view keeps the
original two-call path, so nothing that predates it changes behaviour --
including the three doubles that certify the fail-closed contract, which are
untouched. TestOneIdentityReadPerResolution pins the single read, the
two-call fallback, the fail-closed degrade and the fork refusal; removing
the fold turns the first of those red.
The third read stays: the generation lives in conversation_generations, a
different table, and cannot be folded into a sessions lookup.
4. _declared_conversation_session() documents the concurrent first-turn race.
Two simultaneous first requests on one declared key can each miss the
lookup, mint a row and both bind, because each row is unkeyed at bind time
and the mismatch guard does not fire. That converges rather than crossing:
both rows carry the same key under the same source, so the lookup returns
the later one for every subsequent reply and the earlier row is an abandoned
transcript, never another conversation's identity.
The same docstring still claimed the generation was durable in
sessions.end_reason and that "nothing here needs a counter". That stopped
being true in 09004753c9, which moved the generation into
conversation_generations precisely because deriving it from prunable session
rows was ABA. Corrected, along with the same stale sentence on
TestConversationBoundariesRotate.
5. conversation_generations rows are now documented as deliberately never
collected, rather than merely uncollected. Dropping one resets that peer to
"no generation", so its next boundary writes 1 again and re-issues a gwk_
scope a retired conversation already used -- the exact ABA the table exists
to close. Worth stating because the repo already carries both patterns a
maintainer would extend: delete_session() cascades to messages, and
gateway_hygiene_state is already swept by session_key.
Upgrade notes, one-time on merge:
- One cold prompt-cache bucket per keyed conversation. Every gateway platform
declares gateway_session_key, so each keyed conversation's affinity scope
moves once from its compression-lineage root session id to the gwk_ hash.
One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
recorded as a keyed row and appears in "Active: N session(s)" where it was
invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
first from the next boundary written, so a conversation that reset before the
upgrade shares its predecessor's scope once. One warm bucket, never a crossed
identity.
Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.
Found in review by @teknium1.
Refs #96811
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.
1. _run_agent raised NameError on every opted-in declared bind. Its worker
finally evaluated `if _declared_selected:`, a local of _handle_responses /
_handle_runs that is neither a parameter nor an enclosing binding here, so
the successful declared-key paths failed at settlement after the agent run.
bind_declared_conversation already IS the gate; the inner name is gone.
2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
_run_sync bound unconditionally, so an explicit body session_id that existed
with an empty session_key was adopted by the header key even though the
header lost precedence. It now carries the same gate.
3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
delete_session() deletes the selected row and bulk prune selects ended rows,
so the aggregate can return a pair it already emitted:
(1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
a retired affinity identity. The backwards-clock shape needs no pruning at
all. The generation now lives in a conversation_generations table keyed by
(source, session_key), advanced by _bump_conversation_generation inside the
same transaction that writes each boundary -- outside prunable session
history, wall-clock-free, and increment-only. end_session() and
promote_to_session_reset() both advance it, and only when they actually
wrote a boundary, so a repeated end cannot double-count.
4. The carrier could be memoized under the wrong source. _agent_source() fell
back to agent.platform before the row landed while persistence uses
_session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
it and never re-resolves once the authoritative row appears, so under an
override both sides of a /new read the platform domain and hashed the same
scope. The pre-row path now uses the persistence resolver itself.
Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.
TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Refs #96811
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
The review asked for "real handler + DB coverage for both precedence paths".
The first pass did not deliver that: TestBindFollowsPrecedence restated the
gate expression inside the test, so it asserted a copy of the rule rather than
the rule, and would have stayed green if the handlers stopped applying it.
These drive POST /v1/responses and POST /v1/runs over real routes with a real
adapter and a real SessionDB:
- the declared key selects and records the conversation, and the row it
produces carries the key;
- three replies on one declared key land on one session id;
- an undeclared request keeps a per-request id and records nothing;
- a request carrying conversation A's previous_response_id plus a foreign
header key settles on A, records nothing, leaves A's own key intact, and the
foreign key still cannot recover A -- the end-to-end shape of the blocker;
- /v1/runs, which owns its agent lifecycle rather than routing through
_run_agent, settles on the declared conversation, and an explicit body
session_id outranks the header key without rebinding it.
The _run_agent stand-in creates the session row the way
AIAgent._ensure_db_session does and performs the bind the way _run_agent's
finally block does, so the assertions land on real rows instead of on a mock's
call args. /v1/runs is captured at _create_agent for the same reason.
The restated-gate tests are kept as the cheap unit layer beneath these.
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
Both blockers from @andrexibiza's review of 28a2d7f0ee.
1. The generation lookup was not in the same identity domain as recovery.
latest_conversation_boundary() selected on session_key alone, while
_declared_conversation_session() is qualified by (source, session_key).
X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
API conversation may legally carry the same key as a Telegram row in one
database -- a /new over there rotated this conversation's gwk_ generation
while recovery correctly refused to cross the same line, moving the affinity
identity out from under a physical identity that had not moved.
The boundary read now takes (session_key, source), and the carrier is
'source|key|generation' rather than 'key|generation' -- keying on the string
alone would also collapse two same-key conversations from different sources
onto one routing key, since this value leaves the process verbatim as
OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
from the agent's own session row, falling back to the platform the row will
be created with before it lands.
2. The declared key's stated lower precedence did not survive settlement. Both
handlers let stored_session_id / an explicit body session_id win, then called
_bind_declared_conversation() unconditionally. record_gateway_session_peer()
does SET session_key = ? across compression ancestors, so a request carrying
conversation A's chain plus header key B silently rebound A to B: A could no
longer be recovered by its own key, and B recovered A's session.
Recording is now gated on the declared key having actually selected or
minted the session, on both paths. Behind that gate the bind itself refuses
to overwrite a row already bound to a different key, so a future caller
cannot reintroduce the same defect by opting in wrongly.
test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.
Refs #96811
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.
latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.
Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.
Refs #96811
POST /v1/responses and POST /v1/runs parse and authenticate the client's
X-Hermes-Session-Key, pass it downstream for memory scoping, and then mint a
throwaway physical session id anyway whenever the client manages its own
history (no previous_response_id chain to carry one forward).
Every conversation-affinity hint Hermes sends is derived from that physical
id, so all four re-keyed on every single reply: prompt_cache_key on both
OpenAI-wire transports, the OpenRouter and Nous sticky session_id, and xAI's
x-grok-conv-id. The conversation never landed back on a warm prefix.
Fix the identity rather than the four consumers. The declared key resolves to
its live session through find_latest_gateway_session_for_peer -- the same
reset-fenced recovery every native gateway platform already uses -- and the
turn records the row it ended on through record_gateway_session_peer, which
AIAgent._ensure_db_session never did (it knows the key and writes the row
unkeyed, so the mapping the next reply needs did not exist).
Because the lookup is fenced on sessions.end_reason, the generation that must
rotate is already durable: session_reset (/new), session_switch, idle, daily,
suspended and resume_pending_expired all return None, so a new conversation
gets a new id and a cold affinity scope, and a retired generation can never be
resolved again. No counter, no new persisted field, and no new precedence rule
in the cache-scope resolver -- /branch, delegate and tool children keep the
isolation of #79161/#79017 byte for byte.
Precedence is unchanged where it already worked: an explicit body session_id
and the previous_response_id chain both still outrank the declared key, and a
request that declares nothing keeps its per-request id. Recording is opt-in
(bind_declared_conversation), so no other _run_agent caller's rows change.
Refs #96811
(cherry picked from commit e7c83dddf36784d1012bf483240ebc7f6b2ef9aa)
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.
No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.
That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.
It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.
The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.
Refs #96811
The turn-lease timeout/interrupt paths return from inside the try block
before set_affinity_scope() runs; the finally then read an unassigned
local -> UnboundLocalError. This was the cause of the 4 red
cross-process lease tests on PR #97158's CI.
(cherry picked from commit 05a3c8a4caa59831b01d13f3950e61385aadbfa5)
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).
Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.
Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.
- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
rotation AND across per-response ids). Hashed because, unlike a session id,
the key embeds platform/chat/user identifiers and leaves the process
verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
when a host declared one. The providers read the attribution id when it is
unset, so delegate trees keep sharing their parent's sticky key and every
host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
rules that keep /branch children, delegate subagents and tool children off
their parent's chat key. Background-review forks clone the live runtime, so
_persist_disabled excludes them for the same reason (#79161).
Refs #96570Fixes#96811
Address review follow-ups on the seed fix:
- Move the 'seed only from the 0 state' guard from an inline block in
build_turn_context() into
ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
the predicate with the rest of the speculative-seed lifecycle
(snapshot_preflight_display_tokens /
rollback_interrupted_preflight_display_tokens). Callers now use the
method via a getattr guard so test doubles and external context
engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
described the old >=0 guard ('treats any negative value as no real
usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
method on the compressor fixture, eliminating mirror-drift risk (the
helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
low until their next report) in the method docstring.