docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.
Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".
gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.
Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.
Review follow-up on #109610.
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.
acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.
`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.
Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.
Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.
Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.
holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.
openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.
Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
Unifying the credential pool's `_RETRY_DELAY_PATTERNS` into `RETRY_DELAY_PATTERNS`
flipped the pool's precedence: "retry after 30s; resets in 4hr" cooled the
credential for 14400 s where the pool used to take 30. A body carrying both
describes a short throttle inside a long quota window; the explicit retry-after
is the wait the provider actually asks for, so it is tried before "resets in".
Review follow-up on #109539.
`_get_proxy_for_base_url` lost its guard when it moved onto the shared matcher:
`split_host_port` read `urlsplit(...).port`, which raises ValueError for
`http://host:notaport/v1` or `:99999`, and `build_keepalive_http_client`'s
outer except then returned None -- the client silently lost the shared pool
instead of merely skipping the bypass check. The port parse now yields
`(host, None)` on ValueError only; the host still matches NO_PROXY entries.
Review follow-up on #109539.
The shared matcher took the gateway adapter's `*.` branch, which only matched
subdomains. The adapter's own `is_host_excluded_by_no_proxy` docstring promised
"leading-dot and `*.` entries match the apex domain and subdomains" (the
curl/requests convention), so `NO_PROXY=*.slack.com` silently stopped covering
`slack.com`. `*.` and `.` entries now share one apex+subdomain rule.
Review follow-up on #109539.
Three answers to "is this host in NO_PROXY": process_bootstrap used the stdlib
proxy_bypass_environment (no CIDR, no `*.`), gateway/platforms/base.py had a
full matcher (should_bypass_proxy) and a second suffix-only one
(is_host_excluded_by_no_proxy, used by Slack). Live-verified: with
NO_PROXY=10.0.0.0/8 Telegram bypassed the proxy while the LLM call to a 10.x
endpoint went through it.
The full matcher moves to the leaf module agent/proxy_bypass.py (stdlib only,
importable at early boot); both base.py functions are one-line forwarders and
process_bootstrap._get_proxy_for_base_url uses it (passing host:port so
port-qualified entries match). The six-key proxy env scan is also shared.
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.
The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.
tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.
hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.
Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
`ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
hermes_state_messages._scrub_surrogates (0 callers) deleted.
redact_for_egress's bearer sweep matched any run of token characters after
the word "Bearer", so ordinary prose ("I'm the bearer of bad news") came
back as "Bearer [redacted] bad news" on every chat and A2A reply. The
gateway and A2A sweeps this PR replaced always required 20+ chars; only
monitoring was floor-less. Restore the floor on the opaque branch and keep
the bracket branch that folds an already-masked residue to one marker.
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.
Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.
Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.
Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.
Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.
Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).
Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Both writers already go through atomic_json_write(mode=0o600) but were
missing from tests/test_private_credential_writers.py::_writers, so a
"write then chmod" regression in either would not be caught. Also removes
the `import os` left unused in agent/secret_sources/_cache.py (ruff F401).
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.
utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.
Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).
LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.
- `_status_429` now honors a structured billing code first (decisive
signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
(429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
limit" was already there; providers use both orders). "terminal billing
limit" text is deliberately NOT matched: substring rules cannot negate
the "non-terminal billing limit" wording — the structured code covers it.
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.
- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
after the read loop; surface EOF-without-[DONE] (warn + error result
when nothing was received); extract _consume_sse_line so line parsing
and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
proven red on origin/main.
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.
Two-part fix:
1. title_generator.py: Pass reasoning_config={"enabled": False} to
call_llm() so thinking is explicitly disabled for title generation.
2. chat_completions.py: In _build_gemini_thinking_config, when
reasoning is disabled (enabled=False or effort="none"), set
thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
includeThoughts: False only hides thought parts from the response
while the model still reasons internally and bills thought tokens
against maxOutputTokens. thinkingBudget: 0 truly disables thinking
so thought tokens do not consume the max_tokens budget.
Fixes#91927
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).
Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
Replace the two stubbed tests with a single test that drives the real
CodexTransport.build_kwargs (the fixture agent already carries a tool), so
the test binds the summary body actually sent, not a hand-written dict. It
covers both the first attempt and the empty-summary retry; the previous
pair asserted the same three keys twice and reused one dict across
attempts, so the second-attempt check passed trivially after the first pop.
Add the WHY comment on the pops: the transport emits tools, tool_choice and
parallel_tool_calls as one block, and strict Responses backends 400 on the
controls without tools.
Review finding: the salvaged rule assumed a lowercase-hex grammar that
AgentMail's docs do not establish (only the `am_` / `am_org_` prefix is
documented), so a non-hex key would have gone unmasked. Discriminate on
what actually separates keys from identifiers: an alphanumeric body with
no `_`/`-` and a 20-char floor. `am_example_identifier_123` still passes.
The pattern am_[A-Za-z0-9_-]{10,} was too broad and matched ordinary
identifiers like 'am_example_identifier' that happen to start with 'am_'.
Changed to am_[a-f0-9]{32,} which:
- Requires lowercase hex characters only (real AgentMail keys are hex)
- Requires 32+ characters after the prefix (reducing false positives)
FixesNousResearch/hermes-agent#10983
Review finding: `echo cd backend` injected backend/AGENTS.md, and the
rstrip(";") applied after shlex had removed quoting turned `cd 'backend;'`
into `backend`. Tokenize with punctuation_chars so operators are their own
tokens: a `cd` counts only at a segment start, `backend;ls` splits at the
operator, and a quoted `'backend;'` stays the literal name.
SubdirectoryHintTracker's generic token filter in
_extract_paths_from_command drops any token that does not contain /
or ., so relative directory names in commands like 'cd backend && ls'
were silently ignored. As a result, Hermes missed AGENTS.md /
CLAUDE.md / .cursorrules in the entered subdirectory whenever users
navigated with plain relative names -- the common case.
Add _extract_nav_command_targets, which scans the token stream for
'cd' / 'pushd' and treats the next non-flag token as a path candidate
resolved against working_dir. The generic token pass still runs so
all other shapes (absolute paths, files with extensions, etc.) keep
working. 'cd -' and bare 'cd' are intentionally skipped -- neither
points at a project subdirectory.
Three new tests cover 'cd backend && ls', 'pushd backend', and
'cd backend' appearing after an earlier chained sub-command.
Known limitation called out in review: multi-step chains like
'cd backend && cd src' still resolve each hop against working_dir
rather than simulating the shell's evolving cwd. That's a bigger
change (shell state tracking) and is out of scope for this fix;
the common single-hop case reported in the issue is now covered.
Fixes#11032
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
Two CI regressions from the previous commit, both mine:
- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
`should_fallback` to True for the MoA adapter-shape and missing-preset
verdicts. #55933 made those deliberately NOT fall back (a fallback would
silently replace the MoA route with a single model); restore
`retryable=False` only. The gate in `settle_unrecovered_error` now honours
that for real: on main these verdicts never reached the fallback branch
because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
`_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
block instead.
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).
- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
request-validation and overflow heuristics, returning format_error with
`retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
verdicts that legitimately reach that branch (policy block, TLS chain, MoA
shape/preset errors) now state `should_fallback=True` explicitly, so the gate
changes behaviour only for the new verdict. Local validation errors keep
their historical fallback.
Fixes#12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.
Co-authored-by: cuyua9 <2114364329@qq.com>
`re.sub` with the home path as a template string parsed backslashes as
escapes (re.error dropped the whole injected config block); use a callable.
The early `~/…` return also skipped `expandvars`, leaving `~/$LEAF` half
resolved. Prefix-substitute and fall through to normal expansion instead.
Review finding on #109143: `_handle_stream_error` returned True for the
stream_options rejection without checking whether `_call` had another
iteration. With HERMES_STREAM_RETRIES=0, or after the transient budget was
spent, the loop ended with neither a response nor an error set and the call
returned None instead of raising.
The compatibility retry now extends the loop by one attempt exactly once
(`_compat_retries`); the transient budget is untouched. Test pinned with
HERMES_STREAM_RETRIES=0 (red on the previous head).
Azure AI Foundry serverless (MaaS) endpoints validate the request body
strictly and reject `stream_options: {"include_usage": true}` with 422
`extra_forbidden`. Hermes sent the field on every streaming call, so the
agent was unusable against that endpoint family and the fallback chain
failed too.
When a 400/422 names `stream_options` as an extra/unsupported field and no
delta has been delivered yet, re-open the stream without the field and
remember the rejection on the agent (`_stream_options_unsupported`) so later
turns skip it up front. Streaming itself stays on — this is not the
"stream not supported" case. Usage accounting for such endpoints falls back
to the estimator, as it already does for native Gemini.
Salvage of PR #53271 by @DavidMetcalfe, reshaped onto the `_StreamingCall`
retry loop with per-agent state instead of a process-wide host set; one
end-to-end invariant test through `_interruptible_streaming_api_call`.
Fixes#9705
The _SKILL_INVALID_CHARS regex stripped all non-ASCII characters from
skill names, so skills with CJK, Cyrillic, or other Unicode names
(e.g. "小说拆条") produced an empty slug and were silently skipped.
Change the regex from [^a-z0-9-] to [\w-] so Unicode word characters
are preserved. Platform-specific sanitizers (Telegram, Discord) already
handle their own character restrictions downstream.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
`skip_context_files=True` kept AGENTS.md/CLAUDE.md out of the system prompt,
but `SubdirectoryHintTracker` was always on, so the first tool call touching
a directory with such a file spliced its full text onto the tool result.
Cron jobs without a workdir (which set skip_context_files) that relay exact
stdout then delivered `[Subdirectory context discovered: ...]` plus the file
body to Telegram/Discord.
The tracker now takes `enabled=` and agent init wires it to
`not skip_context_files`: one flag, both injection paths. Interactive
sessions and cron jobs with a workdir are unchanged.
Direction proposed in PR #9434 (@zhitiao), which gated on platform == "cron";
gating on the existing skip flag covers the same case without a platform
special-case.
Fixes#9441
Gateway sessions and batch workers append to the same default
trajectory_samples.jsonl / failed_trajectories.jsonl with a plain open("a") +
write(); concurrent writers interleaved mid-object and the file stopped
parsing (#12684). The append now holds an exclusive lock for write+flush:
flock on POSIX, a 1-byte msvcrt.locking range on Windows.
Tests: a foreign process holding the lock must block the append (red on
base); six processes appending oversized entries all land parseable.
Fixes#12684. Salvaged from #12685 by @shafdev; Windows arm and test trim ours.
`build_tool_preview()`'s generic-key fallback and the cute-message helpers
still truncated with a bare `text[:max_len - 3] + "..."`; for max_len 1-3 the
slice goes negative and returns almost the whole string (27 chars for
max_len=1). `_truncate_preview` already had the guard, so the two code paths
disagreed.
One truncation helper (`_tail_trunc`) with the guard, used everywhere; the
head-truncating `_cute_path` gets the same clamp.
Salvage of PR #48483 by @HeLLGURD (current-code fix); the earliest reports and
patches were #9464 (@LarHope), #9477 (@kagura-agent) and #9497.
Co-authored-by: LarHope <12761142+LarHope@users.noreply.github.com>
Fixes#9439
`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.
Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.
Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.
Fixes#9814
`add_provider()` flipped `_has_external` and appended the provider before
calling `get_tool_schemas()`. When schema loading raised, the broken provider
stayed registered and the single-external slot was poisoned for the rest of
the process: every later provider was rejected as "already registered".
Materialize the schema list first; state changes only after it succeeds.
Exception propagation is unchanged.
Hand-port of PR #9997 by @zhouhe-xydt onto the current add_provider() (the
reserved-core-tool filter landed in between); one invariant test.
Fixes#9948