Commit Graph

95 Commits

Author SHA1 Message Date
Justin Bennington
8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Teknium
a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium
0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
686f6c61
c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
liuhao1024
a9e56c048a fix(compression): anchor codex app-server preflight on reported usage
The codex_app_server runtime bypasses the conversation loop, so the
usage-anchored context accounting captured there never ran: agent._usage_anchor
stayed None forever. Every preflight estimate therefore fell back to the rough
mirror-transcript heuristic, which is deliberately never compacted on this
runtime (_record_codex_app_server_compaction preserves the mirror), so it grows
monotonically while the real thread may be tiny or freshly compacted. With
compression.codex_app_server_auto=hermes that estimate alone tripped the
threshold and fired thread/compact/start on nearly every turn of a long-lived
session (#100381).

Mirror the main loop's post-response capture: _record_codex_app_server_usage
now snapshots the reported thread usage into agent._usage_anchor (base prompt
+ completion exactly as the provider counted them, with estimation confined to
messages appended since). A usage-less turn keeps the previous anchor, and the
post-compaction invalidation site is unchanged.
2026-09-06 13:21:17 -07:00
kshitijk4poor
3ac671dbba refactor(agent): drop the write-only agent-level codex watchdog timestamps
With a request-local watchdog state on every codex request, agent._codex_stream_last_event_ts / _last_progress_ts had no remaining reader; the state-None snapshot only runs for non-codex requests where no watchdog consults it.
2026-09-07 00:41:09 +05:30
kshitijk4poor
1ad8b5dd4a refactor(agent): trim the watchdog phase split to what the tests pin
- Explicit-vs-implicit idle timeout detection uses env_float with a sentinel
  instead of a hand-rolled parse (unset and unparseable both mean implicit).
- _on_event writes the agent-level timestamps only when there is no
  request-local watchdog state; with state present they were dual bookkeeping
  nobody read (the snapshot prefers the state).
- _abort_request keeps main's close-then-retire order: the reorder had no
  test teeth and is a separate concern from #90449.
2026-09-07 00:41:09 +05:30
PeaceMaker-best
463d4bcd15 fix(agent): defer codex idle watchdog until progress
Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
2026-09-07 00:41:09 +05:30
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
fe03731efd review-fix(suppress-audit): codex_runtime/relay_runtime/shutdown_watchdog — restore BASE exception semantics 2026-09-03 09:41:59 -07:00
Teknium
becdbd3f42 refactor(agent/codex_runtime): keep separate anchor-reset statements (source-inspection tests pin them) 2026-09-02 23:17:36 -07:00
Teknium
f6b33242c3 refactor(agent/codex_runtime): single-line remaining logger/raise call wraps 2026-09-02 23:13:32 -07:00
Teknium
ab916de51a refactor(agent/codex_runtime): single-line logger calls and signatures, suppress() approval lookup, drop redundant extra_body temp 2026-09-02 23:11:17 -07:00
Teknium
e73bd99994 refactor(agent/codex_runtime): compact docstrings and comments by hand (WHY/invariants kept) 2026-09-02 22:59:45 -07:00
Teknium
30a1acba47 refactor(agent/codex_runtime): flatten stream-assembler dispatch tables, extra_body merge, and relay stream call 2026-09-02 22:38:16 -07:00
Teknium
894abab3d9 refactor(agent/codex_runtime): collapse single-use locals and request-content probe 2026-09-02 22:28:12 -07:00
Teknium
41428e7bbf refactor(agent/codex_runtime): AST-identical blank-line strip 2026-09-02 22:24:25 -07:00
Teknium
1c56545684 refactor(agent/codex_runtime): inline single-use failure logger, suppress() for single-statement guards, table _stable_call_id prefixes, guarded post-turn hooks via _call_guarded 2026-09-02 22:22:21 -07:00
Teknium
72f1e2e2c1 refactor(agent/codex_runtime,bedrock_adapter): fold assembler state, pending-call guard, toolUse block builder (-16 LOC) 2026-09-02 19:26:47 -07:00
Teknium
db1214a406 refactor(agent/codex*,bedrock_adapter): share Mantle URL parse, billing identity, route flag unpacking; fold stream/tools preflight (-29 LOC) 2026-09-02 19:23:16 -07:00
Teknium
50dafd06b7 style(agent/codex*,bedrock_adapter): restore one blank line after nested defs 2026-09-02 19:18:13 -07:00
Teknium
97c0b89839 refactor(agent/codex_runtime): compact docstrings/comments, keep every invariant and WHY (-66 LOC) 2026-09-02 19:04:48 -07:00
Teknium
e0693c4d3e refactor(agent/codex*,bedrock_adapter): drop intra-function blank lines (AST-identical, -117 LOC) 2026-09-02 18:47:54 -07:00
Teknium
5634e32005 refactor(agent/codex_runtime): extract _output_text_of, fold accumulator init/returns, tighten item payload helpers (-40 LOC) 2026-09-02 18:39:31 -07:00
Teknium
4f5d9da6da refactor(agent/codex_runtime): fold agent-callback guard, dedupe change filtering, extract stream close; collapse signatures (-19 LOC) 2026-09-02 18:30:17 -07:00
Teknium
841d9ff4a8 refactor(agent/adapters): simplify codex responses adapter and runtime (-1080 LOC)
Split _preflight_codex_input_items into per-item-type helpers over a
_PreflightCtx; streaming assembly moves into _CodexResponseAssembler with a
per-event dispatch table; run_codex_app_server_turn loses its session/usage/
interrupt regions to _ensure_codex_session/_finish_codex_turn/
_consume_user_interrupt/_queue_token_counts (counts built lazily so stub agents
without a session DB are never touched). Responses input/normalization and
stream callback order verified byte-identical against merge-base.
2026-09-02 13:29:47 -07:00
kshitijk4poor
2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
notkisk
619ca3011e fix(agent): isolate background review snapshots 2026-09-02 13:28:41 +05:30
AgentLinker
cac9db7caf fix(codex): retired stream requests must not synthesize a completed response
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.

Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.

Fix: publish a per-request retirement token so the worker can tell it has been
retired.

- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
  `agent._active_codex_stream_request_token` before handing off to the worker
  (codex_responses only) and clears it at all four kill sites plus the worker's
  own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
  which can raise — every other call site wraps it in try/except, and a leaked
  token would let a later worker mistake itself for the owning attempt. The
  request-local `_codex_request_retired` mirror also swallows the transport
  error our own force-close causes, so the worker's local error cannot replace
  the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
  `TimeoutError` from `interrupt_check` when it no longer owns the request —
  raising rather than breaking, because a break returns the partial `final`.
  The four stream callbacks also drop post-retirement frames so an abandoned
  attempt cannot stream tokens into the live turn's bubble (the gateway caches
  AIAgent instances per session).

`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.

Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
2026-09-01 22:14:06 -07:00
Teknium
c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
Teknium
d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
Jan-Stefan Janetzky
8cc379b528 fix(memory): bind the checkpoint gate to every compaction authority
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):

- codex app-server: in "native"/"off" auto-compaction mode (native is the
  default) Hermes preflight is skipped and the codex agent compacts its
  own thread inside run_turn() — the compress_context() rejection was
  unreachable. init_agent now refuses checkpoint_required together with
  api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
  testable guard), and run_codex_app_server_turn() fails closed as
  defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
  native_compaction_context_management() now returns None while the gate
  is armed, so context_management never goes on the wire and the
  checkpoint-aware Hermes compressor stays authoritative. The suppression
  is logged once per process, not silently applied.

Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.

Refs #93986
2026-08-25 03:55:55 -07:00
kchernev
10a070bd49 fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:29:45 +05:30
Teknium
580060ffd8 fix: reuse first-observed sequence when announced items land via output_item.done
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.

Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
2026-08-23 19:02:30 -07:00
cxxCoolStar
4f3ae189a3 fix(agent): harden pending Responses tool call settlement 2026-08-23 19:02:30 -07:00
cxxCoolStar
720344cfba fix(codex): settle pending Responses tool calls when output_item.done is omitted
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
2026-08-23 19:02:30 -07:00
kshitijk4poor
f6d1d774a1 refactor(codex): name the dropped shape in the wire-guard warning
Review polish from the 3-angle pass on the final stack:

- The warning now says WHICH shape leaked (top-level, extra_body, or both).
  Relay injects top-level while request_overrides typically inject via
  extra_body, so the shape identifies the offending middleware when
  debugging.
- Fold the 'always returns a fresh mapping' assertion into the parametrized
  real-endpoint test (the caller mutates the result with stream=True, so the
  copy contract is load-bearing on no-drop paths too) and drop the
  SimpleNamespace stub test it strictly subsumes. The nested-preserve stub
  stays: the parametrized test only exercises top-level retention.
2026-08-20 03:15:14 +05:30
kshitijk4poor
26530e7df5 fix(codex): strip nested extra_body retention at the consumer Codex wire
The wire guard only removed the top-level prompt_cache_retention kwarg, but
the OpenAI SDK merges extra_body into the outgoing JSON body, so a nested
extra_body.prompt_cache_retention reaches chatgpt.com/backend-api/codex just
the same and still triggers the non-retryable HTTP 400. Both injection
vectors are real and probe-verified: the Relay overlay's 'key not in
baseline' arm admits an interceptor-added extra_body, and
request_overrides={'extra_body': {...}} lands verbatim in build_kwargs
output.

Close the gap in the same helper: strip the nested field too (copy-on-write,
never mutating the caller's mapping), drop extra_body entirely when it
empties, and log the same warning. Compatible endpoints keep nested
retention untouched.

Mutation-verified: removing the extra_body leg fails both new nested tests.

Reported by egilewski's review on #89969.
2026-08-20 03:15:14 +05:30
kshitijk4poor
ba4bc39afd test(codex): pin the retention drop to real endpoints
The salvaged compatibility test stubs `_is_codex_backend=lambda: False` on a
SimpleNamespace, so it proves the helper honors its own boolean but not that
the boolean is right for any real endpoint. A predicate change that widened
the drop onto retention-supporting hosts would keep it green.

Adds a parametrized test that builds a real AIAgent per base URL and asserts
the drop only fires for chatgpt.com/backend-api/codex, while api.meta.ai,
bedrock-mantle.*.api.aws, api.openai.com and a same-host/different-path
backend keep their supported 24h value. Also asserts prompt_cache_key
survives untouched on every endpoint, since retention and cache-key routing
are independent and the guard must not disturb caching.

Verified non-vacuous: relaxing the guard's condition to drop on every
endpoint fails 4 of the 6 cases (Meta, Bedrock, OpenAI, non-codex path).

Drive-by on the guard itself: drop the dead `None` default on the `pop` that
is already gated by an `in` check, and record why the predicate is resolved
via getattr -- run_codex_stream is driven with lightweight stand-in agents
that lack `_is_codex_backend`, so a bare call would raise AttributeError.
2026-08-20 03:15:14 +05:30
Falko
8e2949495a fix(codex): strip unsupported cache retention at wire 2026-08-20 03:15:14 +05:30
Tuck
79a41c1d32 fix: convert remaining messages.append() calls to append_message() 2026-08-15 01:04:19 -07:00
kshitij
10b2b11efa fix: widen APIConnectionError handling to finalization drain loop
The PR added APIConnectionError handling to the main request and
iteration try blocks but missed the finalization drain loop (line ~1492).
That site catches httpx transport errors to preserve an already-completed,
already-billed response when the drain iterator fails. Without the
APIConnectionError handler, an SDK-wrapped transport error during drain
would propagate uncaught and discard the completed response.

Also strengthens the test's no-payload-leak assertion to check the full
request body and URL are absent from the log message, not just the
literal string 'payload'.
2026-08-12 12:42:39 +05:30
Gille
aa3ca1c3be fix(agent): tolerate transport errors without requests 2026-08-12 12:42:39 +05:30
Gille
da1170670b fix(agent): log Codex transport failure details 2026-08-12 12:42:39 +05:30
Brooklyn Nicholson
6bb630ef78 fix(codex): split reasoning summary parts on summary_index
The native Responses stream does carry summary_index, so the part boundary is
structured data here rather than something to infer. Break on a change of
index, and leave streams that send no index (plain reasoning_text) untouched.
2026-08-06 22:02:46 -05:00
joaomarcos
7b5997f111 fix(codex): don't retry after a valid response.completed on drain error
A transport error during the post-terminal SSE drain (used only to let the
Relay transport finalize the attempt) was sharing exception handling with
the pre-terminal assembly path, so it discarded an already-completed,
already-billed response and opened a brand-new physical request. Give the
drain step its own non-fatal error handling: log a finalization warning
and still return the terminal response.

Closes #74310
2026-07-29 16:41:39 -03:00
teknium1
5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
Alex Fournier
b1a5d67e71 Merge upstream main into fix/hermes-relay-anthropic-context
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-07-28 08:10:58 -07:00