Bot Mode group turns on a Windows Desktop attached to a remote gateway
dialed a registry secondary per member profile. That second WebSocket
accept/closed in ~30ms and never ran session.create. When the window
primary is already that one-host-many-profiles source (sharedRemote),
request/retain/open/ensure now reuse the primary socket and scope RPCs
with profile=. Isolated SSH/pooled backends are unchanged.
The tui_gateway/run_agent writers now stamp profile_name explicitly, but
every OTHER creation path that passes no profile_name (cli.py /new,
hermes_cli/main.py --create-if-missing, foreign-session import, ACP
adapter, gateway branch/title paths, the #82616 peer self-heal INSERT,
and compression children of legacy NULL parents) still minted
profile_name = NULL rows. Rows minted NULL after the one-shot #94724
legacy-owner backfill ran stayed NULL forever: profile-keyed consumers
(desktop sidebar scope matching, @session:<profile>/<id> deep links, the
fail-closed owner ladder) treat NULL as unowned, so the sessions vanished
from the sidebar with their transcripts intact (#99222).
Fix the class at the choke point instead of chasing call sites: every
profile-tree state.db belongs to exactly one profile, so
SessionDB._insert_session_row (and the peer self-heal INSERT and the
compression-child publisher) derive the store's own profile from db_path
when the caller names none — <root>/state.db -> 'default',
<root>/profiles/<name>/state.db -> <name>. The same single-match contract
backfill_null_session_profiles and the web listing's row_profile stamp
already rely on. Explicit profile_name arguments always win; stores
outside the profile tree (tests, ad-hoc copies) keep NULL — never guess.
E2E-verified with real imports against a temp HERMES_HOME:
before (origin/main) a bare create_session on the default store persisted
NULL; after, 'default' / '<profile>' land in state.db for the default
store, a named-profile store, the peer self-heal insert, and a
compression child of a NULL parent, while explicit args and
outside-tree stores are unchanged.
Refs #99222
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).
NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.
Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.
Fixes#99222
Session mutations from the desktop sidebar silently no-op against the wrong
profile's state.db and reappear on the next refresh, unless the mutated
session happens to belong to the serving (primary) profile. Most visible in
the "All Profiles" view and on a remote-primary desktop (a registered remote
gateway as primary, every profile served from it): deleting a session flashes
away optimistically, then comes back after a profile switch.
Root cause: the mutation helpers passed the owning profile ONLY as
request.profile, which the Electron main process consumes for backend ROUTING
but which does not scope the request the backend actually receives. The
backend selects its target DB from ?profile= (DELETE) or body.profile (PATCH).
Reads (getSession, getSessionMessages) and renameSession already scope
correctly; the other mutations regressed / never did:
- deleteSession -> no ?profile= in the URL
- setSessionArchived -> no profile in the PATCH body
- setSessionPinnedRemote -> no profile in the PATCH body
- setSessionUnreadRemote -> no profile in the PATCH body
On a remote gateway whose connection has no remoteProfile alias, the main
process leaves such requests unscoped, so the backend opens its own (default)
state.db, cannot find another profile's row, and returns
{ok:true, already_absent:true} (DELETE) or no-ops (PATCH) — a fake success the
UI treats as done. deleteSession lost its ?profile= scope in the api/ module
split (it was present via PR #44138 / the pre-split hermes.ts path).
Fix: scope all four mutations the same way the working endpoints do —
deleteSession appends sessionScopeQuery(profile) to the URL; setSessionArchived
/ setSessionPinnedRemote / setSessionUnreadRemote include profile in the PATCH
body (mirroring renameSession). request.profile stays for per-profile
remote-override and global-remote routing. Single-profile / selected-profile
users are unaffected (the serving profile already matched).
Verified against a live remote-primary desktop: the DELETE now goes out as
/api/sessions/<id>?profile=<owner> and the row is actually removed from the
owning profile's state.db (confirmed server-side) instead of returning
already_absent.
Adds api/sessions.test.ts coverage: delete scopes ?profile= in the URL for
object and bare-string owners and omits it when no owner is known; archive /
pin / unread carry body.profile when owned and omit it otherwise.
Fixes#78836
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.
Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
config.get and config.set ignored the focused profile on a shared
app-global backend, so reads and persistent writes used the launch
config.yaml. Bind the existing @_profile_scoped decorator and write
_save_cfg through the request home override.
Fixes#95760
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
_env_line_defines_key() decides which .env lines the writers may rewrite or
drop. It matched on the `KEY=` prefix, but load_env() splits on the first
`=` and strips the name:
key, _, value = line.partition('=')
env_vars[key.strip()] = _parse_env_value(value)
so `OPENAI_API_KEY = sk-...` is a live assignment — the key resolves, the
provider works, and every UI shows it as set. The writers did not see it.
This is the same resurrection hole #40041 fixed for `export KEY=`, still
open for the whitespace form:
- DELETE /api/env 404s ("not found in .env") while the credential stays
active — a key the user revoked through the UI is never actually revoked
- PUT /api/env appends a SECOND line instead of replacing; a later delete
removes the appended line and the original value silently comes back
Rotate-then-delete on a spaced line therefore restores exactly the key the
user rotated away from.
Match load_env()'s parse instead of prefix-matching, so the writers accept
precisely what the reader accepts: skip blank/comment/no-'=' lines, strip an
`export ` prefix, then compare the stripped name. Commented-out lines stay
untouched and `KEY_EXTRA=`/`MY_KEY=` still do not match `KEY`.
Verified against the real dashboard endpoints on a temp HERMES_HOME: the
spaced line is now removed, rotation replaces it in place with no duplicate,
and a parity check asserts the writer matches a line iff load_env() does.
`_scrub_config_yaml_mirrors` reconciles the config.yaml copies of a credential
when it is rotated or removed through the dashboard. It walks `model`,
`auxiliary.<task>`, and `custom_providers.<name>` — but not the keyed
`providers` schema.
`providers` is not a niche section: `get_compatible_custom_providers` documents
it as "the newer keyed schema" (v12+), and it is exactly where the dashboard /
desktop write a custom endpoint's inline key —
`_write_custom_endpoint` sets `providers.<id>.api_key`. That value is a real
credential: the runtime resolver reads it (`runtime_provider` /
`hermes_cli.main` / `model_switch` all read `entry.get("api_key")` off a
`providers` entry), and an inline key outranks the env var.
So the section the scrub skips is the one the dashboard writes to, and both
callers break on it:
- save_provider_env_credential (rotation, #62269): a stale
`providers.<id>.api_key` is left at the OLD value and, being
higher-precedence than the freshly-rotated env var, shadows the rotation —
the "persistent 401 with a key the UI no longer shows" that #62269 fixed,
reintroduced for the newer schema.
- remove_provider_env_credential: its contract is to "remove a credential from
EVERY store it lives in", yet the `providers` copy survives, leaving the
secret in config.yaml after the user asked to delete it.
Walk `providers.<id>` too. The scrub stays value-matched, so an unrelated
endpoint's key is untouched. Only `api_key` is scrubbed here: in the keyed
`providers` schema `api` is the base_url alias, not a credential (unlike
model/auxiliary/custom_providers), so `_fix` takes an explicit field list and
this section passes `("api_key",)` — a provider's endpoint URL is never
rewritten even if it happened to equal the credential string.
tests/hermes_cli/test_credential_lifecycle.py: drive the real PUT/DELETE
/api/env endpoints against a `providers.<id>.api_key` mirror — rotation moves
it to the new key, delete clears it, and a `providers.<id>.api` base_url alias
is preserved. The two scrub tests fail on main (stale key survives); the
base_url guard passes on main as a control. 15 pass here; 345 pass across the
credential-lifecycle + web-server suites (the one failing honcho-merge test
fails identically on clean main).
ensureRegistryBackend() reuses the ambient primary descriptor for a
non-local, non-ssh registry primary, but the reused descriptor was
missing sharedRemote: true. The request router only appends
?profile=<profile> for sharedRemote backends, so Capabilities/Skills
requests went out unscoped and the gateway served the default profile
while a named profile was selected.
Add the flag and a regression test asserting the scoped URL.
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
comment pinning WHY it is the timeout ladder's first rung (transient
lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
test_compression_rotation_state.py / test_context_compressor.py; kept the
force-bypass test (only site pinning that behavior for split failures) and
the E2E test (now asserting the named constant).
The cooldown is recorded inside the split-failure except handler; a raw
call on a stub/partial compressor would replace the real split error with
an AttributeError. Mirror the try/except-debug convention of the adjacent
record_rejected_compaction call.
The salvaged unit tests drive _record_compression_failure_cooldown directly;
this drives the real _compress_context split-failure path (archive boom on a
real SessionDB) and asserts the cooldown recording fires with the
session_split_failed error class.
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):
1. publish_compression_child gains require_lease_refresh: the lease is
extended inside the same transaction as the expiry check (same conn, no
TOCTOU), giving a worker whose refresher thread died from transient DB
failures one final chance to keep its completed work.
2. A failed compression split now records a 60s failure cooldown, so the
next turn cannot immediately re-trigger the same compression.
Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
Attribution mapping for the #98137 salvage (author commit rides in the
same PR); must precede the authored commit so every commit passes
check-attribution.
Follow-ups on the salvaged #97797 transport:
- Blocker 2 from the #97681 exact-head review claimed near-expiry refresh
silently mints against the target's CURRENT policy. On this head the
handler DOES refuse drift (_require_unchanged_execution_policy -> 403
room_reauthorization_required), but nothing pinned the handler-level
behavior: removing the drift check still passed the entire grants suite
(the check was only unit-tested in isolation). New HTTP-level regression
test drives /v1/room-members/grants/refresh with a drifted-policy grant
and requires the 403; sabotage-verified (check removed -> test fails).
- cancel() conflict resolution: keeps our race-retry routing loop from
#99099 with this layer's peer-stop acknowledgement body inside it
(peer receipt -> settle completion -> local interrupt escalation).
- docs: NAT one-way-reachability note in bot-mode.md — Desktop is a viewer,
not a relay; put room authority on the host everyone can reach (field
finding from /bin/bash on #97681).
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.
Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.
Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
The stream consumer called prefers_fresh_final_streaming(text,
metadata=...) only, and no metadata producer stamps a platform key — so
RelayAdapter's hook always fell back to the PRIMARY descriptor's
platform (the scalar-vs-per-chat capability seam, third occurrence).
Two failure directions on multiplexed relays with
platforms.relay.extra.slack.unfurl_links/media: true (#97957):
- Slack primary fronting Telegram/Discord: every link-bearing streamed
final on the non-Slack chats finalized as a fresh send with no delete
op advertised -> the answer delivered TWICE (orphaned preview).
- Non-Slack primary fronting Slack: the hook returned False, leaving
the force-on unfurl feature dark on exactly the chats it shipped for.
Pass chat_id=self.chat_id from the consumer; the relay hook already
accepted it and resolves via _platform_by_chat + the per-platform
negotiated descriptor. Graduated TypeError fallback keeps the
single-platform hook signatures (Telegram, base class) and legacy test
doubles working unchanged.
Both regression tests verified RED against the unfixed consumer, GREEN
with the fix; single-platform relays are unaffected (#97957's own 30
tests unchanged-green).
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.
Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.
Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
The PKCE payload is a flat 'provider=...;state=...;verifier=...;next=...'
string. A raw ';' is a cookie-attribute terminator, so Python's
http.cookies emits the value in RFC 6265 quoted form with each ';'
escaped as the backslash-octal '\073'. Mainstream browsers echo that
form back verbatim and Python parsers decode it — the browser round
trip is fine. But '"' and '\' are outside the plain cookie-octet set,
and non-Python hops that re-serialize the Cookie header reject the
value and drop the cookie entirely: Go's net/http (Traefik middleware,
Authentik outposts, other gateways) refuses any cookie value
containing a backslash. The OIDC callback then 400s with "Missing
PKCE state cookie" even though the browser sent the cookie.
Field reproduction: support thread "Still unable to use Authentik for
signin with traefik" — devtools showed the browser sending the intact
quoted \073 cookie on /auth/callback while Hermes logged
missing_pkce_cookie behind a Traefik+Authentik chain.
Fix: URL-encode the whole payload in set_pkce_cookie (quote(payload,
safe='') — ';' becomes '%3B') so the wire value contains only
cookie-octets and no parser in the chain has anything to reject, and
decode through a single shared inverse, cookies.parse_pkce_payload(),
in BOTH readers: the OAuth /auth/callback and the native
password-login path (routes.login_submit), whose broker/provider
binding check would otherwise parse zero segments from the
newly-encoded value and silently disable itself.
Regression coverage: the wire-shape test pins the full cookie-octet
set (the '"'/'\' assertions are the ones a Go-parser hop fails
pre-fix), the round-trip tests drive the real /auth/login →
/auth/callback path, and the next= test pins the exact post-login
redirect byte shape. Native-flow broker assertions updated to decode
through parse_pkce_payload instead of substring-matching the raw wire
value.
Salvaged from #84065 (rebased onto current main, which gained the
SameSite=None PKCE attrs and the RFC 8252 native password flow since
the PR branched): kept main's _pkce_attrs cookie shape, extended the
fix to the login_submit reader the original PR predated, and reframed
the rationale — browsers do NOT truncate at the first ';' (there is
no literal ';' on the wire in the quoted form); the failing hop is a
strict middlebox cookie parser.
Closes#83832
Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
Render durable per-member hold state in Group Chat with canonical resume guidance and accessible, theme-safe status copy.\n\nVerified by independent pre-commit review.
Query-file DM transports do not consume stdin. Use DEVNULL for both the initial attempt and policy-gated retry so Git Bash cannot pass an invalid pseudo-handle to Windows subprocess creation.
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).
Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.
Fix, at the single status-determination choke point both the single-task
and batch paths share:
- `failed=True` on the child result now wins over a non-empty summary:
status = "failed".
- The child's classified failure_reason (rate_limit / billing /
server_error / ...) is propagated onto the batch entry so the parent
can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
through to "max_iterations" (which also wrongly set truncated=True).
Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
Follow-ups on the salvaged #97744 runner:
- tui_gateway/hosted_room_driver.py: HostedRoomRuntime.cancel() treated its
initial status read as truth, so a task transitioning queued->running (or
settling) between the read and the state call surfaced a transient
'running work requires acknowledged two-phase cancellation' /
StaleTaskError to the caller and failed groups.disband. Deterministic
repro on the PR head: test_client_event_id_cannot_squat_disband_receipt
failed 5/5 locally. cancel() now re-reads and re-routes on every
race-shaped failure (bounded retries), returns already-cancelled tasks
idempotently, and rejects truly terminal states honestly.
- methods_groups conflict resolution keeps both method sets: the replication
surface from #99047 (groups.replicate/replica_state/promote/demote) and
the runner surface from this layer (groups.stop/retry/approve).
- test_groups_replication_methods.py updated to the runner's stricter
create contract (2-6 profile-backed members, live worker service).
Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):
[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
marks_failed: an error string with the 'failed' key ABSENT (not False)
must still be status=failed + exit_reason=error. The branch order
(result.get('failed') or result.get('error')) already handles this; the
test pins the error-alone path.
[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
completed: error='' is falsy so result.get('error') falls through to the
summary-presence heuristic => status=completed. No code change; the
existing branch is correct and the new test locks it in.
[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
exit_reason consumer:
* tools/delegation_live_log.py finalize() prints exit_reason generically
and only special-cases == 'max_iterations' for a readable suffix.
* tools/process_registry.py derives truncated as
(truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
* tools/async_delegation.py passes exit_reason through generically.
The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
the delegation result's exit_reason. No exhaustive if/elif over the enum
missing an 'error' case, so nothing to add.
[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
status in {completed, interrupted, failed} and exit_reason in {completed,
max_iterations, interrupted, error}, and added a compact enum comment at
the result-entry construction. Verified the process_registry.py renderer
comment (truncated <= exit_reason == 'max_iterations') still holds — the
truncation flag is derived exactly that way, so no contradiction.
[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
explicitly set failed=False' is not adopted. No consumer produces a result
dict with an explicit failed=False and no summary while relying on
completed semantics: run_agent.py sets failed=True only on genuine failure
and omits the key on success (no failed=False producer). Also, the
proposed elif would reintroduce ambiguity (explicit failed=False + no
summary => 'completed'?) and diverge from the conservative else => 'failed'.
result.get('failed') is falsy for both explicit-False and absent, so no
distinction exists to preserve; the else is the correct default.
Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.
Closes#97654.
A provider-rejected child (e.g. HTTP 400 "<model> is not a valid model ID")
returns completed=False with failed=True + an error string as its terminal
final_response. _run_single_child keyed status on summary presence alone and
assumed completed=False meant iteration-budget exhaustion, so such a child
was reported status=completed + exit_reason=max_iterations, rendering the
false '"TRUNCATED: hit max_iterations"' banner.
Consult the structured failure fields (failed / error) before falling back to
the summary-presence heuristic, and derive exit_reason honestly: failure ->
'error', interrupted -> 'interrupted', completed -> 'completed', and only
genuine budget exhaustion (completed=False, no failure) -> 'max_iterations'.
The 'truncated' flag stays keyed on exit_reason == 'max_iterations', so it is
now correct automatically. The batch renderer needed no change (the error
field is already plumbed into the result entry for the parent).
Closes#97655
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.
- delegate_tool: result.failed now forces status 'failed' (with the
error carried on the entry); new shared format_subagent_failure_line()
renders one clean human-readable line (traceback -> exception message,
length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
terminal failure status deliver that line via _deliver_platform_notice
BEFORE all progress-queue gates; tool_progress_callback is now always
attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
Every participant gateway can now keep a durable copy of a hosted room's
ordered log and continue the room when its authority host is gone:
- gateway/hosted_room_replicas.py: replica store in root state.db.
ingest_page() persists authority-stamped groups.log pages idempotently,
refusing sequence gaps and authority-epoch regressions. promote_replica()
continues the room locally at epoch+1 with a lineage-proving
authority.claimed event; the stale owner is fenced everywhere the claim
replicates. demote_room() lets a returning stale authority fence itself
(authority.lost) upon observing a newer epoch, killing split-brain writes.
- tui_gateway/methods_groups.py: groups.replicate / groups.replica_state /
groups.promote / groups.demote RPC surface. Promotion requires
confirm=true — storage decides HOW takeover is atomic and provable, the
caller (user action now, lease/quorum driver later) decides WHEN it is
safe, matching the boundary blessed on #97681.
Validation: 20 new tests incl. a full failover round-trip (A hosts, B
replicates incrementally, A dies, B promotes with complete history, A
returns demoted and fenced); 69 total across the hosted-rooms area; E2E
with two real gateway stores and real install identities.
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.
Fix, per the documented compression.codex_app_server_auto contract:
* Session hygiene now routes codex_app_server sessions to
run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
cached agent's thread via thread/compact/start (through the existing
codex route in _compress_context) and KEEPS that agent cached; 'native'
and 'off' skip cleanly with no eviction and no local fallback. A wedged
compaction records the persistent failure cooldown; success resets the
hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
a temporary compression agent and compacts the live thread with
force=True instead (a manual compress is an explicit user decision in
every mode). No live thread -> honest "nothing to compact" reply
instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
rewriting the mirror cannot shrink the thread, so the #73715-style
local fallback (including its force=True leak into native/off) is
deliberately not adopted.
Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.
Closes#73503
Co-authored-by: webtecnica <webtecnica@gmail.com>
PR #98628 removed _build_chunk_digests, so the two lean chunk-digest
cancellation tests reintroduced by the #97512 cherry-pick target a
deleted mechanism — removed. The #96775 stall-interrupt assertions now
match the stall_interrupted marker inside the strategy/kind-stamped
durable error instead of assuming it is the prefix.