Commit Graph

4752 Commits

Author SHA1 Message Date
teknium1
0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
teknium1
991b23ad8d refactor(sessions): one session-id minter; QQ update-prompt key from build_session_key
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.

Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".

gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.

Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
2026-09-13 05:21:02 -07:00
teknium1
b5f072e2ce fix(agent): manual /compress "nothing to compress" gate stays gateway-only
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.

Review follow-up on #109610.
2026-09-13 05:20:26 -07:00
teknium1
c0be0e0826 refactor(display): one think-tag list; ACP tool titles derive from agent/display previews
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.

acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
2026-09-13 05:20:26 -07:00
teknium1
7114da3de6 refactor(agent): manual /compress runs through one core with --preview/--aggressive on every surface
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.

`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
2026-09-13 05:20:26 -07:00
teknium1
f678ed8299 refactor(model-providers): thinking-toggle XOR effort translation lives in agent.reasoning_effort; opencode-free imports it instead of borrowing via sys.modules
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
2026-09-13 05:19:48 -07:00
teknium1
ca2ca5b9e1 fix(secret-scope): mem0, langfuse and azure credential readers stop swallowing UnscopedSecretError
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.

Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.

Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
2026-09-13 05:19:48 -07:00
teknium1
9ba5850e95 refactor(memory-plugins): background threads inherit the profile context; JSON sidecar reads and the holographic config.yaml write use core primitives
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.

Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.

holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.

openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.

Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
2026-09-13 05:19:48 -07:00
teknium1
3bef6b6a5c fix(agent): explicit "retry after N s" wins over "resets in ..." in the shared reset table
Unifying the credential pool's `_RETRY_DELAY_PATTERNS` into `RETRY_DELAY_PATTERNS`
flipped the pool's precedence: "retry after 30s; resets in 4hr" cooled the
credential for 14400 s where the pool used to take 30. A body carrying both
describes a short throttle inside a long quota window; the explicit retry-after
is the wait the provider actually asks for, so it is tried before "resets in".

Review follow-up on #109539.
2026-09-13 05:09:43 -07:00
teknium1
8121b8f438 fix(agent): a malformed port in base_url no longer drops the keepalive transport
`_get_proxy_for_base_url` lost its guard when it moved onto the shared matcher:
`split_host_port` read `urlsplit(...).port`, which raises ValueError for
`http://host:notaport/v1` or `:99999`, and `build_keepalive_http_client`'s
outer except then returned None -- the client silently lost the shared pool
instead of merely skipping the bypass check. The port parse now yields
`(host, None)` on ValueError only; the host still matches NO_PROXY entries.

Review follow-up on #109539.
2026-09-13 05:09:43 -07:00
teknium1
5b31da6b4d fix(agent): NO_PROXY *.example.com matches the apex domain again
The shared matcher took the gateway adapter's `*.` branch, which only matched
subdomains. The adapter's own `is_host_excluded_by_no_proxy` docstring promised
"leading-dot and `*.` entries match the apex domain and subdomains" (the
curl/requests convention), so `NO_PROXY=*.slack.com` silently stopped covering
`slack.com`. `*.` and `.` entries now share one apex+subdomain rule.

Review follow-up on #109539.
2026-09-13 05:09:43 -07:00
teknium1
164c5ea142 fix(agent): LLM traffic honours CIDR and wildcard NO_PROXY entries like the platform adapters do
Three answers to "is this host in NO_PROXY": process_bootstrap used the stdlib
proxy_bypass_environment (no CIDR, no `*.`), gateway/platforms/base.py had a
full matcher (should_bypass_proxy) and a second suffix-only one
(is_host_excluded_by_no_proxy, used by Slack). Live-verified: with
NO_PROXY=10.0.0.0/8 Telegram bypassed the proxy while the LLM call to a 10.x
endpoint went through it.

The full matcher moves to the leaf module agent/proxy_bypass.py (stdlib only,
importable at early boot); both base.py functions are one-line forwarders and
process_bootstrap._get_proxy_for_base_url uses it (passing host:port so
port-qualified entries match). The six-key proxy env scan is also shared.
2026-09-13 05:09:43 -07:00
teknium1
3f59b5c594 refactor(agent): /context breakdown and native-compaction retention use the canonical token estimator
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
2026-09-13 05:09:43 -07:00
teknium1
398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
teknium1
469a87f7a5 refactor(config): collapse thin _load_config copies onto the canonical readers
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.

tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
2026-09-13 05:09:06 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
f680431f2e fix(redact): bearer residue sweep needs a 20-char token floor
redact_for_egress's bearer sweep matched any run of token characters after
the word "Bearer", so ordinary prose ("I'm the bearer of bad news") came
back as "Bearer [redacted] bad news" on every chat and A2A reply. The
gateway and A2A sweeps this PR replaced always required 20+ chars; only
monitoring was floor-less. Restore the floor on the opaque branch and keep
the bracket branch that folds an already-masked residue to one marker.
2026-09-13 05:07:50 -07:00
teknium1
c849bc383a refactor(env): agent.secret_scope.load_env_file is the only .env tokenizer; six hand parsers collapse onto it
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.

Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.

Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.

Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
2026-09-13 05:07:50 -07:00
teknium1
226df89f74 refactor(redact): one secret-pattern source; a2a, gateway chat and monitoring egress scrub through redact_for_egress
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.

Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.

Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).

Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
2026-09-13 05:07:50 -07:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
602801baa1 chore(secrets): cover the spawn ledger and meet-node token in the private-writer invariant; drop unused import
Both writers already go through atomic_json_write(mode=0o600) but were
missing from tests/test_private_credential_writers.py::_writers, so a
"write then chmod" regression in either would not be caught. Also removes
the `import os` left unused in agent/secret_sources/_cache.py (ruff F401).
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1
2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Teknium
92df11f81d fix(classifier): terminal_quota_exhausted 429s classify as billing, not rate_limit
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).

LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.

- `_status_429` now honors a structured billing code first (decisive
  signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
  (429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
  limit" was already there; providers use both orders). "terminal billing
  limit" text is deliberately NOT matched: substring rules cannot negate
  the "non-terminal billing limit" wording — the structured code covers it.
2026-09-12 21:52:48 -07:00
Teknium
11d761eed9 fix(stream): flush residual SSE buffer at EOF and widen resilience to the Gemini native adapter
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.

- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
  after the read loop; surface EOF-without-[DONE] (warn + error result
  when nothing was received); extract _consume_sse_line so line parsing
  and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
  flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
  proven red on origin/main.
2026-09-12 21:18:24 -07:00
Shakti Prasad Mohapatra
29056b2335 fix(title): disable Gemini thinking tokens to prevent max_tokens starvation
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.

Two-part fix:

1. title_generator.py: Pass reasoning_config={"enabled": False} to
   call_llm() so thinking is explicitly disabled for title generation.

2. chat_completions.py: In _build_gemini_thinking_config, when
   reasoning is disabled (enabled=False or effort="none"), set
   thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
   includeThoughts: False only hides thought parts from the response
   while the model still reasons internally and bills thought tokens
   against maxOutputTokens. thinkingBudget: 0 truly disables thinking
   so thought tokens do not consume the max_tokens budget.

Fixes #91927
2026-09-12 21:10:02 -07:00
liuhao1024
5ea655771b fix(agent): disable reasoning on the title-generation pass 2026-09-12 21:10:02 -07:00
Teknium
1fec70ea48 fix(guardrails): catch repeating multi-call cycles in the stall guard
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).

Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
2026-09-12 21:07:46 -07:00
cedanoagent
387ac50d85 feat(video): add OpenRouter Hailuo 3 Max provider 2026-09-12 13:44:52 -07:00
kshitijk4poor
476a45f4f3 test(agent): one real-transport test for the codex summary tool controls
Replace the two stubbed tests with a single test that drives the real
CodexTransport.build_kwargs (the fixture agent already carries a tool), so
the test binds the summary body actually sent, not a hand-written dict. It
covers both the first attempt and the empty-summary retry; the previous
pair asserted the same three keys twice and reused one dict across
attempts, so the second-attempt check passed trivially after the first pop.

Add the WHY comment on the pops: the transport emits tools, tool_choice and
parallel_tool_calls as one block, and strict Responses backends 400 on the
controls without tools.
2026-09-12 22:40:28 +05:30
Julien Talbot
16695bfa07 fix(codex): strip tool controls from summary calls
(cherry picked from commit b058c25740b4a4bfa40ffbe2ce7bcddb9fe930dd)

Co-authored-by: Matthieu Talbot <1246794+MartyLake@users.noreply.github.com>
2026-09-12 22:40:28 +05:30
teknium1
848075c501 fix(redact): AgentMail prefix rule matches any opaque key body, not only hex
Review finding: the salvaged rule assumed a lowercase-hex grammar that
AgentMail's docs do not establish (only the `am_` / `am_org_` prefix is
documented), so a non-hex key would have gone unmasked. Discriminate on
what actually separates keys from identifiers: an alphanumeric body with
no `_`/`-` and a 20-char floor. `am_example_identifier_123` still passes.
2026-09-12 08:37:14 -07:00
nightq
cc885b8b19 fix: tighten AgentMail API key regex to avoid false positives
The pattern am_[A-Za-z0-9_-]{10,} was too broad and matched ordinary
identifiers like 'am_example_identifier' that happen to start with 'am_'.

Changed to am_[a-f0-9]{32,} which:
- Requires lowercase hex characters only (real AgentMail keys are hex)
- Requires 32+ characters after the prefix (reducing false positives)

Fixes NousResearch/hermes-agent#10983
2026-09-12 08:37:14 -07:00
teknium1
088d292d36 fix(agent): navigation targets only from a cd that starts a shell segment
Review finding: `echo cd backend` injected backend/AGENTS.md, and the
rstrip(";") applied after shlex had removed quoting turned `cd 'backend;'`
into `backend`. Tokenize with punctuation_chars so operators are their own
tokens: a `cd` counts only at a segment start, `backend;ls` splits at the
operator, and a quoted `'backend;'` stays the literal name.
2026-09-12 08:34:07 -07:00
Trevin Chow
7d61b4872a fix(agent): resolve bare cd backend targets in SubdirectoryHintTracker
SubdirectoryHintTracker's generic token filter in
_extract_paths_from_command drops any token that does not contain /
or ., so relative directory names in commands like 'cd backend && ls'
were silently ignored. As a result, Hermes missed AGENTS.md /
CLAUDE.md / .cursorrules in the entered subdirectory whenever users
navigated with plain relative names -- the common case.

Add _extract_nav_command_targets, which scans the token stream for
'cd' / 'pushd' and treats the next non-flag token as a path candidate
resolved against working_dir. The generic token pass still runs so
all other shapes (absolute paths, files with extensions, etc.) keep
working. 'cd -' and bare 'cd' are intentionally skipped -- neither
points at a project subdirectory.

Three new tests cover 'cd backend && ls', 'pushd backend', and
'cd backend' appearing after an earlier chained sub-command.

Known limitation called out in review: multi-step chains like
'cd backend && cd src' still resolve each hop against working_dir
rather than simulating the shell's evolving cwd. That's a bigger
change (shell state tracking) and is out of scope for this fix;
the common single-hop case reported in the issue is now covered.

Fixes #11032
2026-09-12 08:34:07 -07:00
Keane Yan
d1a13ae244 fix(auxiliary): normalize model on auto cache miss 2026-09-12 08:28:49 -07:00
teknium1
79007efc48 fix(agent): a no-fallback verdict wins over the local-ValueError fallback allowance
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
2026-09-12 08:27:11 -07:00
teknium1
293cb54f28 fix(agent): keep MoA verdicts fallback-free; nest the gated fallback under one if
Two CI regressions from the previous commit, both mine:

- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
  `should_fallback` to True for the MoA adapter-shape and missing-preset
  verdicts. #55933 made those deliberately NOT fall back (a fallback would
  silently replace the MoA route with a single model); restore
  `retryable=False` only. The gate in `settle_unrecovered_error` now honours
  that for real: on main these verdicts never reached the fallback branch
  because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
  `_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
  site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
  broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
  block instead.
2026-09-12 08:27:11 -07:00
teknium1
dfd4aa4a94 fix(agent): malformed tool-call-argument 400s no longer walk the fallback chain
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).

- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
  request-validation and overflow heuristics, returning format_error with
  `retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
  verdicts that legitimately reach that branch (policy block, TLS chain, MoA
  shape/preset errors) now state `should_fallback=True` explicitly, so the gate
  changes behaviour only for the new verdict. Local validation errors keep
  their historical fallback.

Fixes #12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.

Co-authored-by: cuyua9 <2114364329@qq.com>
2026-09-12 08:27:11 -07:00
teknium1
66dbb656be fix(skills): literal HOME replacement; ~/$VAR still expands the variable
`re.sub` with the home path as a template string parsed backslashes as
escapes (re.error dropped the whole injected config block); use a callable.
The early `~/…` return also skipped `expandvars`, leaving `~/$LEAF` half
resolved. Prefix-substitute and fall through to normal expansion instead.
2026-09-12 08:25:59 -07:00
yoma
86211492d1 fix(skills): expand HOME defaults against tool home 2026-09-12 08:25:59 -07:00
teknium1
10864f394d fix(agent): stream_options compatibility retry does not consume the transient budget
Review finding on #109143: `_handle_stream_error` returned True for the
stream_options rejection without checking whether `_call` had another
iteration. With HERMES_STREAM_RETRIES=0, or after the transient budget was
spent, the loop ended with neither a response nor an error set and the call
returned None instead of raising.

The compatibility retry now extends the loop by one attempt exactly once
(`_compat_retries`); the transient budget is untouched. Test pinned with
HERMES_STREAM_RETRIES=0 (red on the previous head).
2026-09-12 08:25:40 -07:00
DavidMetcalfe
6527af2286 fix(agent): retry once without stream_options when an endpoint rejects it (HTTP 400/422)
Azure AI Foundry serverless (MaaS) endpoints validate the request body
strictly and reject `stream_options: {"include_usage": true}` with 422
`extra_forbidden`. Hermes sent the field on every streaming call, so the
agent was unusable against that endpoint family and the fallback chain
failed too.

When a 400/422 names `stream_options` as an extra/unsupported field and no
delta has been delivered yet, re-open the stream without the field and
remember the rejection on the agent (`_stream_options_unsupported`) so later
turns skip it up front. Streaming itself stays on — this is not the
"stream not supported" case. Usage accounting for such endpoints falls back
to the estimator, as it already does for native Gemini.

Salvage of PR #53271 by @DavidMetcalfe, reshaped onto the `_StreamingCall`
retry loop with per-agent state instead of a process-wide host set; one
end-to-end invariant test through `_interruptible_streaming_api_call`.

Fixes #9705
2026-09-12 08:25:40 -07:00
enigma
036b3ff5ba fix(skills): allow non-ASCII characters in skill command names (#12351)
The _SKILL_INVALID_CHARS regex stripped all non-ASCII characters from
skill names, so skills with CJK, Cyrillic, or other Unicode names
(e.g. "小说拆条") produced an empty slug and were silently skipped.

Change the regex from [^a-z0-9-] to [\w-] so Unicode word characters
are preserved. Platform-specific sanitizers (Telegram, Discord) already
handle their own character restrictions downstream.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-12 08:25:36 -07:00
teknium1
5148b76677 fix(agent): skip_context_files also disables subdirectory hint injection
`skip_context_files=True` kept AGENTS.md/CLAUDE.md out of the system prompt,
but `SubdirectoryHintTracker` was always on, so the first tool call touching
a directory with such a file spliced its full text onto the tool result.
Cron jobs without a workdir (which set skip_context_files) that relay exact
stdout then delivered `[Subdirectory context discovered: ...]` plus the file
body to Telegram/Discord.

The tracker now takes `enabled=` and agent init wires it to
`not skip_context_files`: one flag, both injection paths. Interactive
sessions and cron jobs with a workdir are unchanged.

Direction proposed in PR #9434 (@zhitiao), which gated on platform == "cron";
gating on the existing skip flag covers the same case without a platform
special-case.

Fixes #9441
2026-09-12 08:25:25 -07:00
teknium1
330004e597 fix(agent): keep trajectory serialization inside the best-effort handler
A non-JSON value used to log 'Failed to save trajectory' and return; moving
json.dumps above the try made it raise to the caller.
2026-09-12 08:25:14 -07:00
shafdev
bd70d7387a fix(agent): lock save_trajectory() appends so concurrent processes cannot corrupt the JSONL
Gateway sessions and batch workers append to the same default
trajectory_samples.jsonl / failed_trajectories.jsonl with a plain open("a") +
write(); concurrent writers interleaved mid-object and the file stopped
parsing (#12684). The append now holds an exclusive lock for write+flush:
flock on POSIX, a 1-byte msvcrt.locking range on Windows.

Tests: a foreign process holding the lock must block the append (red on
base); six processes appending oversized entries all land parseable.

Fixes #12684. Salvaged from #12685 by @shafdev; Windows arm and test trim ours.
2026-09-12 08:25:14 -07:00
HeLLGURD
9d09598111 fix(display): tool previews never exceed a tiny max_len (1-3)
`build_tool_preview()`'s generic-key fallback and the cute-message helpers
still truncated with a bare `text[:max_len - 3] + "..."`; for max_len 1-3 the
slice goes negative and returns almost the whole string (27 chars for
max_len=1). `_truncate_preview` already had the guard, so the two code paths
disagreed.

One truncation helper (`_tail_trunc`) with the guard, used everywhere; the
head-truncating `_cute_path` gets the same clamp.

Salvage of PR #48483 by @HeLLGURD (current-code fix); the earliest reports and
patches were #9464 (@LarHope), #9477 (@kagura-agent) and #9497.

Co-authored-by: LarHope <12761142+LarHope@users.noreply.github.com>
Fixes #9439
2026-09-12 08:23:42 -07:00
MonkeyLeeT
289e0d506f fix(insights): count tool usage per session before merging the two sources
`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.

Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.

Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.

Fixes #9814
2026-09-12 08:23:13 -07:00
zhouhe-xydt
5084237f8e fix(memory): load provider schemas before mutating MemoryManager state
`add_provider()` flipped `_has_external` and appended the provider before
calling `get_tool_schemas()`. When schema loading raised, the broken provider
stayed registered and the single-external slot was poisoned for the rest of
the process: every later provider was rejected as "already registered".

Materialize the schema list first; state changes only after it succeeds.
Exception propagation is unchanged.

Hand-port of PR #9997 by @zhouhe-xydt onto the current add_provider() (the
reserved-core-tool filter landed in between); one invariant test.

Fixes #9948
2026-09-12 08:22:57 -07:00