Commit Graph

20 Commits

Author SHA1 Message Date
teknium1
c350d3e44f fix(bootstrap): race IPv6/IPv4 for every sync connect from the process bootstrap, not per client
Wire the racer #114277 added once, at the seam every Hermes process already
crosses first: ``hermes_bootstrap`` (imported before anything else by
``hermes``, ``hermes-agent``, ``hermes-acp``, ``gateway.run``, ``batch_runner``,
``tui_gateway.entry`` and the slash worker). Installing it from
``init_agent`` was too late for the startup path the report is about — the TUI
gateway starts MCP discovery and the model-catalog prewarm before any AIAgent
exists, and ``hermes model`` / picker prewarm never build one.

- Move the stdlib-only racer core (``_happy_eyeballs_create_connection``,
  ``_interleave_addrinfos``) and the installer into ``hermes_bootstrap`` and
  apply it on import; ``agent.process_bootstrap`` keeps only the httpcore
  backend and imports the racer from there. The bootstrap must stay stdlib-only
  because entry points call ``harden_import_path()`` after importing it.
- Patch urllib3's own serial connect walker lazily through a one-shot
  ``sys.meta_path`` hook instead of importing urllib3 eagerly: ``hermes`` and
  the TUI gateway never load urllib3 at start, and importing it costs ~50 ms.
- Idempotence by a marker on the installed function, so a re-import of the
  bootstrap (tests, ``importlib.reload``) never wraps the racer twice.
- Drop the ``init_agent`` call site (redundant: ``process_bootstrap`` imports
  the bootstrap) and trim the five added tests to two invariants in the
  mirroring ``tests/test_hermes_bootstrap.py``; the racer unit test moves its
  monkeypatch seams to the new module.
- Docs: ``network.force_ipv4`` now describes the default racing behaviour.

Live A/B (stub resolver: blackholed 100::1 first, local IPv4 second, driving
the real clients after ``import hermes_cli.main``): catalog fetch via requests
5.05s -> 0.39s, shared keepalive httpx client 15.24s -> 0.32s, inline
httpx.Client 5.01s -> 0.26s; IPv4-only control 0.01s both sides. Refused ports
still fail instantly with the same exception types; a blackhole-only host still
raises ConnectTimeout at the configured connect timeout.

Fixes #114265
2026-09-18 09:56:28 -07:00
liuhao1024
5a88baa79b fix(agent): keep socket racer signature-complete with stock create_connection
Forward the 3.11+ keyword-only all_errors parameter instead of raising
TypeError, and resolve the timeout sentinel to socket.getdefaulttimeout()
so a process default from socket.setdefaulttimeout() survives on the
winning socket, exactly like the stock serial original. The urllib3
racer's fallback now also receives the original sentinel instead of a
resolved None. Both pinned with regression tests.
2026-09-18 09:56:28 -07:00
liuhao1024
298cfd326e fix(agent): race IPv6/IPv4 on every startup-path sync connect, not just Codex
The RFC 8305 racer in agent/process_bootstrap.py was only wired to the
Codex cloud transport (and Codex OAuth flows). Every other sync TCP
connect in the process still walked the getaddrinfo results serially:
the model catalog fetch (urllib/http.client), provider warm
(requests/urllib3), and sync LLM clients (httpcore) all stalled for the
full connect timeout per AAAA record on a network whose advertised IPv6
route is blackholed, stacking minutes onto TUI cold start (#114265).

install_happy_eyeballs_socket_connect() patches the two funnels every
one of those stacks routes through - socket.create_connection (httpcore
resolves it at call time; http.client re-reads it per connection) and
urllib3.util.connection.create_connection (requests warm) - with the
existing racer. Idempotent, installed once in init_agent next to
_install_safe_stdio; a racer crash that is not an ordinary connect
error falls back to the stock serial connect.

Fixes #114265
2026-09-18 09:56:28 -07:00
teknium1
164c5ea142 fix(agent): LLM traffic honours CIDR and wildcard NO_PROXY entries like the platform adapters do
Three answers to "is this host in NO_PROXY": process_bootstrap used the stdlib
proxy_bypass_environment (no CIDR, no `*.`), gateway/platforms/base.py had a
full matcher (should_bypass_proxy) and a second suffix-only one
(is_host_excluded_by_no_proxy, used by Slack). Live-verified: with
NO_PROXY=10.0.0.0/8 Telegram bypassed the proxy while the LLM call to a 10.x
endpoint went through it.

The full matcher moves to the leaf module agent/proxy_bypass.py (stdlib only,
importable at early boot); both base.py functions are one-line forwarders and
process_bootstrap._get_proxy_for_base_url uses it (passing host:port so
port-qualified entries match). The six-key proxy env scan is also shared.
2026-09-13 05:09:43 -07:00
Teknium
2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
c3b411dfb7 perf(agents): share one httpx transport pool across every agent's client
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.

What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.

Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).

Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
  SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
  again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
  default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
  `limits=` never reached them, so mounts ran on httpx defaults with a 5 s
  keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
  pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).

Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
2026-09-03 02:35:21 -07:00
Teknium
123d962dc7 refactor(agent): prompt_builder/plugin_llm/process_bootstrap pass-2 — inline _normalize_ref, fold snapshot read, proxy-env scan 2026-09-02 22:32:12 -07:00
Teknium
a5858a9e98 refactor(agent): process_bootstrap/markdown_tables/onboarding pass-2 — inline _pad_to_width, fold proxy bypass and pool guard, pack signatures 2026-09-02 22:21:06 -07:00
Teknium
6e20411e4a refactor(agent): prompt_builder siblings — inline snapshot write, compact __all__/locals, tidy blank-line conventions 2026-09-02 19:05:15 -07:00
Teknium
afce1b22db refactor(agent): compact prompt_builder siblings — collapse guard ladders, inline single-use helpers, trim docstrings (emitted text byte-identical) 2026-09-02 18:22:01 -07:00
Teknium
624a752e79 refactor(agent/runtime): deadline/file_safety/estop/lifecycle leaf modules — dedupe result plumbing, drop dead guard code
- deadline: _result()/_abandon() replace 7 BoundedResult constructions and 2
  cancel+callback sites; timers handled as a list; dead raise_if_timed_out removed.
- file_safety: retired classify_cross_profile_target (0 refs; get_cross_profile_warning
  stub kept for external callers), _home_and_resolved/_mirror_warning shared by
  the sandbox/container mirror guards; _find_sandbox_mirror_segments inlined.
- estop: _hermes_home/_canonical_root now the file_safety helpers;
  _reset_log_state_for_tests inlined into its only test.
- subagent_lifecycle: _validate_request driven by _UNSUPPORTED_REQUEST_FIELDS table.
- turn_liveness: dead start()/_abort_message removed, _emit_warning shared.
- process_bootstrap: _enable_happy_eyeballs reused by the client variant.
- Comment/docstring compaction across the remaining leaf modules.
2026-09-02 13:29:48 -07:00
Teknium
419232d49b fix(codex): extend Happy-Eyeballs racing to Codex OAuth/auth clients; pin async native racing
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:

- hermes_cli/auth.py Codex OAuth clients (token refresh at
  auth.openai.com/oauth/token, device-code login, token exchange, usage
  probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
  every connect eats the full timeout per AAAA before IPv4 is tried, so
  auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
  had no explicit racing wired.

Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
  installs the existing _HappyEyeballsSyncBackend on a ready-built sync
  httpx.Client's direct transports (default transport + mounts), skipping
  proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
  host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
  five Codex OAuth/probe endpoints. Best-effort: falls back to default
  serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
  RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
  no custom backend needed. Documented in build_keepalive_http_client and
  pinned by tests (contract test on the anyio signature + a live
  regression test where a blackholed 100::1 IPv6 addr hangs and local
  IPv4 wins in ~250ms instead of the serial connect timeout).

network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.

Refs #13834; follows #94388 (9cce8725).
2026-09-01 12:08:11 -07:00
Teknium
9cce872505 docs: note httpcore pin dependency in _enable_happy_eyeballs 2026-08-24 21:46:05 -07:00
Nathan Shan
d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
teknium1
51c1ba6976 fix(agent): apply pool-level keepalive to the process_bootstrap sibling builder
The salvaged #54550 converted AIAgent._build_keepalive_http_client but the
near-identical build_keepalive_http_client in agent/process_bootstrap.py
(used by auxiliary clients: compression, vision, web_extract, titles) kept
the socket_options transport and the api.githubcopilot.com bypass. Same
conversion: httpx.Limits(keepalive_expiry=20) + pool timeouts, verify
forwarded on client and no-proxy mounts, copilot hardcode removed.
2026-07-05 03:14:55 -07:00
kshitijk4poor
676236bb1d fix(agent): honor custom CA certs on aux client + harden TLS resolution
The salvaged fix wired per-provider ssl_ca_cert / ssl_verify (and
HERMES_CA_BUNDLE) into the MAIN OpenAI client. This follow-up:

- Auxiliary client parity: process_bootstrap.build_keepalive_http_client
  accepts and forwards verify; auxiliary_client._resolve_aux_verify mirrors
  the main-client TLS resolution (via load_config_readonly, the read-only
  fast path) so compression/vision/web_extract/title-gen/session_search
  honor the same per-provider CA. Without this, chat worked against a
  private-CA endpoint but every auxiliary call still failed APIConnectionError.
- switch_model now reads custom_providers from live config (load_config_readonly)
  instead of the init-time agent._custom_providers snapshot, so ssl_ca_cert /
  ssl_verify edits are honored on mid-session model switch — matching the
  context-length reload (#15779).
- Drop the dead client-level verify= where a custom httpx transport is used
  (httpx ignores it there); verify lives on the transport. Fix docstrings.
  Applies to both run_agent._build_keepalive_http_client and process_bootstrap.
- resolve_httpx_verify: add CURL_CA_BUNDLE to the env chain (consistency with
  agent/ssl_guard._CA_BUNDLE_ENV_VARS) and emit a loud logger.warning naming
  the endpoint whenever ssl_verify:false disables verification.
- get_custom_provider_tls_settings: case-insensitive base_url match (config
  dedup already lowercases; scheme/host are case-insensitive) so a mixed-case
  entry doesn't silently drop its CA. Exact match preserved — no prefix bypass.
- Demote best-effort except Exception: pass in agent_init/switch_model to
  logger.debug(exc_info=True).
- Tests for aux verify forwarding, _resolve_aux_verify, case-insensitive
  match, and prefix-bypass rejection.
2026-07-02 04:51:56 +05:30
HexLab98
073847c0f2 fix(auxiliary): use env-only proxy policy for OpenAI SDK clients (#53702)
Auxiliary clients now inject a keepalive httpx transport with explicit
HTTPS_PROXY/NO_PROXY resolution, matching the main agent. This avoids
macOS system proxy settings (which omit the ExceptionsList) breaking
vision and other auxiliary calls to internal provider endpoints.
2026-06-27 21:22:49 -07:00
teknium1
5f309ae685 refactor(run_agent): extract OpenAI proxy, safe stdio, IterationBudget
Three small extractions into focused modules:

* agent/process_bootstrap.py — \_OpenAIProxy (lazy openai.OpenAI import),
  \_SafeWriter (broken-pipe-resistant stdio wrapper), \_install_safe_stdio,
  \_get_proxy_from_env, \_get_proxy_for_base_url. All process / IO bootstrap.
* agent/iteration_budget.py — IterationBudget class (thread-safe consume/
  refund counter shared by parent agent and subagents).

run_agent re-exports every name so existing test patches like
patch('run_agent.OpenAI', ...) and 'from run_agent import IterationBudget'
keep working unchanged.  Verified the patch-rebinding contract for OpenAI
explicitly.

tests/run_agent/ + tests/agent/test_gemini_fast_fallback.py:
1347 passed, 3 skipped.
run_agent.py: 15427 -> 15261 lines (-166).
2026-05-16 17:59:32 -07:00