Commit Graph

4537 Commits

Author SHA1 Message Date
teknium1
735831f776 fix(mcp): reconcile chore lives in run_profile_reconcile; prune lazy + mid-connect servers
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.

reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.

test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
2026-09-13 06:29:12 -07:00
Teknium
4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
teknium1
3a179fe524 refactor(gateway): every Ogg/Opus voice transcode routes through base.transcode_to_ogg_opus
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).

`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.

Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
2026-09-13 05:32:38 -07:00
teknium1
3e066dfedd fix(computer_use): approval goes through the shared gate; no callback now fails closed
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.

_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.

Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
  headless -q, plain library use): destructive actions are now REFUSED
  with a BLOCKED error and never reach the backend. Previously they
  silently ran. cron honors approvals.cron_mode, unattended platforms
  approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
  approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
  command_allowlist entry (`cua:click:background`) and is scoped to that
  action+mode — the old blanket "always_approve unlocks everything for
  the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
  "BLOCKED: Action timed out ..."); the error JSON keeps `action`.

Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
2026-09-13 05:21:02 -07:00
teknium1
db5827e3d2 fix(skills_hub): clawhub 429 keeps a zero Retry-After as zero
The rebase resolution used 'parsed or 5', which turned a legitimate 0 s
(negative headers clamp to 0) into the 5 s default; the pre-refactor code
kept it. Test None explicitly.
2026-09-13 05:09:43 -07:00
teknium1
1512edfb85 refactor(tools): terminal, execute_code, MCP and the bounded collector truncate through one head/tail helper
Four copies of the 40/60 head/tail algorithm with a near-identical notice
(terminal_tool_result, mcp_tool_content, code_execution_tool,
environments/base_output) collapse into tools/tool_output_truncate.py, so the
ratio and the `... [<LABEL> TRUNCATED - N <unit> omitted out of T total] ...`
marker are defined once. execute_code keeps byte mode + spill path and only
shares the notice/split. Visible change: the terminal notice now uses
thousands separators like the other three (`9,000 chars` not `9000 chars`).

kanban_specify._truncate: comment claimed escape stripping the body never did;
comment now says what the plain clamp is for.
2026-09-13 05:09:43 -07:00
teknium1
3f59b5c594 refactor(agent): /context breakdown and native-compaction retention use the canonical token estimator
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
2026-09-13 05:09:43 -07:00
teknium1
398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
teknium1
469a87f7a5 refactor(config): collapse thin _load_config copies onto the canonical readers
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.

tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
2026-09-13 05:09:06 -07:00
teknium1
93d0dba281 fix(profiles): HERMES_HOME-blind path sites resolve through hermes_constants
worktree_gc archived untracked files under ~/.hermes regardless of the active
profile or HERMES_HOME (and the Windows LOCALAPPDATA default); dashboard_procs'
remote-lock dir, both bot_mode `_default_home` copies, methods_bot_relay's
`_relay_root` (a hand copy of get_default_hermes_root's profiles/ strip) and
load_hermes_dotenv's default home all re-derived env-or-`~/.hermes` by hand and
so diverged from the platform default. Each now calls the canonical getter
with the intent it already had: get_hermes_home() where the active profile
matters (archive), get_process_hermes_home() where the process asset must stay
visible under a routed-profile override (locks, bot mode, startup .env),
get_default_hermes_root() for install-wide relay state.
2026-09-13 05:09:06 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
dd1baee0e4 refactor(secrets): drop scope-aware env shims; runtime_provider and the voice/xai tools read the canonical getters
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.

tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.

Behavior change: none.
2026-09-13 05:07:50 -07:00
teknium1
c849bc383a refactor(env): agent.secret_scope.load_env_file is the only .env tokenizer; six hand parsers collapse onto it
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.

Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.

Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.

Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
2026-09-13 05:07:50 -07:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1
2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Teknium
82199439c7 feat(image_gen): add Meta Muse Image ($0.01/img) to the FAL catalog
meta/muse-image/text-to-image + paired meta/muse-image/edit, the FAL
listing of Meta's Muse Image model (launched on the Meta Model API in
Aug 2026 at $0.01/image).

- aspect_ratio size family (16:9 / 1:1 / 9:16 from the vendor's
  21:9..9:21 enum); always sent on t2i for deterministic framing,
  deliberately omitted on edits so Muse follows the input image.
- No seed in the vendor schema (Grok Imagine 2.0 precedent) - the
  supports whitelist filters it.
- Edit takes 1-10 reference image_urls (max_reference_images=10).

Schema verified against FAL's OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403), same as prior catalog
additions.
2026-09-12 22:18:06 -07:00
Teknium
23af232837 fix(tools): reject malformed tool parameter schemas at registration
Port from earendil-works/pi#9300 fix (acaa253cc): a plugin registering a
tool whose schema["parameters"] is not a dict (a list, string, etc.)
previously registered fine and the malformed schema was serialized into
every provider request, 400-ing turns far from the offending plugin.
Live probe on main confirmed the bad schema flows into _fn_def() and the
OpenAI wire unchanged.

Fail at registry.register() with the tool name in the error instead. The
plugin loader already catches registration exceptions and marks the
plugin errored, so a broken plugin degrades gracefully rather than
breaking every session. Schemas that omit "parameters" stay valid
(no-argument tools); MCP tools are unaffected (their schemas pass
through _normalize_mcp_input_schema first, which always returns a dict).
2026-09-12 22:11:56 -07:00
Teknium
7f61ae589f Port from openai/codex#41436: answer blocking terminal queries in background PTY sessions
Programs run via terminal(background=true, pty=true) can block forever when
they probe their terminal — device-status (ESC[5n), window-size (ESC[18t),
cursor-position (ESC[6n), or DEC private-mode (ESC[?N$p) queries — because
nothing on the PTY master side answers, and the raw query bytes leak into
captured output.

- tools/pty_query_responder.py: incremental byte scanner that strips the
  handled queries from PTY output (chunk splits included) and produces
  bounded replies; everything else passes through untouched.
- tools/process_registry.py: wire the responder into _pty_reader_loop
  (POSIX only — ConPTY answers its own queries); flush partial escape
  tails at end-of-stream.
- tests mirror the codex fixtures plus a live-PTY E2E where a subprocess
  blocks on ESC[6n until answered.
2026-09-12 22:09:19 -07:00
Indigo Karasu
41380ccef9 fix(process): list refreshes no longer leave exited children running
Carve the direct-child list reconciliation from Indigo Karasu's earliest
PR #60506 (2a96ae2cbf806ccdc3e9b584911774f32622f421), corroborated by
fangliquanflq's narrow #81385 (50ffd243d92627e4a03a3ee8427ad0f5e090c3ab).
Run the existing helper after task/session filtering and reuse the
idempotent owner-stamped completion path. Do not import cross-session
disclosure, bare-PID healing, forget RPCs, or reader rewrites.

A real-child regression fails before this change and passes after it:
the direct child exits while its descendant keeps writing to stdout;
listing reports exit without consuming the result or waiting for EOF.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-12 22:06:03 -07:00
shuwen.wu
9db604ce38 fix(tools): reload permanent allowlist by replacement 2026-09-12 22:03:19 -07:00
durden
98d54ee7b3 docs(approval): say that save_permanent_allowlist can only add, and that a revocation waits for the next write
Review follow-up. Two of the three items taken as written; the third declined
with a reason.

1. Taken. The reconcile semantics mean `patterns` may only ADD -- an entry left
   out of it is not removed, because the on-disk list wins for anything this
   process did not approve itself. Every caller in the tree is additive today,
   so nothing breaks, but the signature does not say so. Stated in the
   docstring, and pinned by
   `test_a_caller_that_passes_a_smaller_set_does_not_remove` so a future
   `allowlist remove` finds out here instead of in production.

   NOT taken: the `reconcile: bool = True` opt-out. There is no caller that
   wants it, and AGENTS.md:98-101 names exactly this -- "Speculative
   infrastructure. Hooks, callbacks, or extension points with no concrete
   consumer." The removal path is editing config.yaml, which the docstring now
   says.

2. Taken. website/docs/user-guide/security.md, next to the existing
   `hermes config edit` tip, which is where an operator reads about removing a
   pattern: the list is read at startup, a pattern removed while a session is
   running stays approved in that session until the next write or a restart,
   and if it was removed for safety reasons, restart.

3. Taken. `test_save_failure_is_logged_not_raised` asserted non-raising but
   never asserted the log its name promises. Now asserts
   "Could not save allowlist" via caplog.

    scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
    === Summary: 1 files, 9 tests passed, 0 failed (100% complete) in 0.4s
2026-09-12 22:03:19 -07:00
durden
9b06d3d081 fix(approval): saving the allowlist deletes entries the operator added by hand and resurrects ones they revoked
`load_permanent_allowlist()` runs exactly once, at module import
(tools/approval.py, the call at the bottom of the module), and
`load_permanent()` only unions into `_permanent_approved` (:2866-2869) --
nothing ever removes. `save_permanent_allowlist()` then wrote that in-memory
set straight back over `config["command_allowlist"]`, at eight call sites.

`command_allowlist` is a file the operator edits, and deleting a line from it
is the documented way to withdraw a standing approval. Any hand edit made
while a Hermes process is live was undone by that process's next `[a]lways`,
in both directions at once.

Reproduced on this tree with a temp HERMES_HOME:

    BEFORE (tools/approval.py at fcbd107)
      on disk before this process starts : ['git status', 'ls *']
      operator edits config.yaml by hand : ['ls *', 'npm test']
                                            (revoked 'git status', added 'npm test')
      after ONE [a]lways                 : ['docker *', 'git status', 'ls *']
      is_approved still honours revoked? : True

    AFTER
      after ONE [a]lways                 : ['docker *', 'ls *', 'npm test']
      is_approved still honours revoked? : False

`npm test` was silently deleted from the operator's own config file, and
`git status` -- a standing approval they had just withdrawn -- was written
back and kept auto-approving. Neither prints anything.

The same shape loses writes between two live Hermes processes: whichever
saves second overwrites the other's entry.

The fix reconciles at write time. The file is re-read and the result is what
is on disk now, plus what this process approved since its own baseline, where
the baseline is what `command_allowlist` held the last time this process
synchronised with the file. That difference is what separates "the operator
granted this here" from "this was on disk at import and may since have been
revoked". Revoked entries are also dropped from `_permanent_approved` so
`is_approved()` stops honouring them for the rest of the process.

It does NOT make a revocation take effect the instant the file changes --
nothing re-reads the file on the approval hot path, and adding a stat there is
a separate change with its own cost. It makes the next write stop undoing the
operator's edit.

`_lock` is `threading.Lock` and not reentrant; all eight call sites were
checked and none holds it across the call, so the added critical section
cannot deadlock. The failure path still logs and returns rather than raising,
as before.

Searched open and merged PRs and issues for `command_allowlist revoke`,
`permanent allowlist reload`, `approval allowlist clobber` and
`save_permanent_allowlist` -- nothing covers this.

Tests: tests/tools/test_permanent_allowlist_reconcile.py, 8 cases -- both
halves of the bug, the two-process race, idempotence, the unedited round trip,
and the existing contract that a config write failure is logged rather than
raised.

    scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
    === Summary: 1 files, 8 tests passed, 0 failed (100% complete) in 0.5s

No regression across the 29 test files in tests/ that touch the allowlist or
the approval module: 25 failed before and after, byte-identical failure set
(pre-existing missing-dependency failures in my local venv).
2026-09-12 22:03:19 -07:00
Teknium
e21a6fb159 fix(tools): make the empty/whitespace old_string rejection actionable
Port from cline/cline#13970: models that send patch calls with an empty
old_string got back 'old_string cannot be empty' — an error that names the
problem but not the recovery, so the next call was byte-identical and the
run burned turns until loop detection killed it (upstream repro: Kimi K3
looping on old_text: null).

The rejection now states the recovery: set old_string to the exact text the
replacement should replace, read the file first if unsure, use write_file
for new files/full rewrites, and do not re-send the call unchanged. The
whitespace-only rejection gets the same treatment. No behavior change for
valid calls.
2026-09-12 21:55:07 -07:00
Eugeniusz Gilewski
fef98ff00f fix(skills): bound streamed ClawHub ZIP downloads (#57571)
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.

Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.

Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
2026-09-12 21:50:30 -07:00
liuhao1024
0e13fa98ec fix(web): cap web_extract provider dispatch with a wall-clock timeout (salvage #57180)
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.

Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).

Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."

Fixes #57155

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-12 21:30:38 -07:00
Teknium
cbd4492f1f fix(tools): steer/redirect releases a blocking process_manage wait (port of MoonshotAI/kimi-code#3697)
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).

wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.

Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.

Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
2026-09-12 21:25:00 -07:00
entropy-0x
f79cb77224 fix(tools): reject V4A Add File onto an existing path
A V4A `*** Add File:` operation is meant to create a new file. The apply
path called `write_file` unconditionally and `_validate_operations` had no
pre-check for ADD, so an Add targeting a path that already existed
overwrote the file with only the patch's `+` lines, returned success, and
emitted a `--- /dev/null` diff that hid what was lost. Models frequently
confuse Add with Update, so this destroyed existing file contents with no
error. The MOVE path already guards its destination against clobbering;
ADD now follows the same rule.

Makes a V4A `Add File` operation fail when its target already exists,
instead of silently overwriting the existing file. Validation now rejects
the operation before any write happens, so the two-phase
validate-then-apply contract ("no files were modified" on a validation
failure) holds for ADD as it already does for UPDATE/MOVE/DELETE. A
matching re-check in the apply phase closes the validate-to-apply race.

N/A

- [x] 🐛 Bug fix (non-breaking change that fixes an issue)

- `tools/patch_parser.py`: add an ADD branch in `_validate_operations`
  that errors when `read_file_raw` finds an existing file, and a
  defensive existence re-check in `_apply_add` before `write_file`,
  mirroring the existing MOVE destination guard.
- `tests/tools/test_patch_parser.py`: add a test asserting an Add onto an
  existing path fails validation and leaves the original bytes unwritten;
  add `read_file_raw` to three ADD-path LSP fakes so they match the real
  `file_ops` interface now exercised on ADD.

1. Build a V4A patch with `*** Add File: <path>` where `<path>` already
   exists on disk.
2. Apply it via `apply_v4a_operations`. Before this change the file is
   overwritten with the patch's `+` lines and the result is success;
   after, the result is a validation failure and the file is untouched.
3. Run `scripts/run_tests.sh tests/tools/test_patch_parser.py` —
   `TestApplyOperations::test_add_onto_existing_file_fails_and_preserves_contents`
   covers the regression.

- [x] I've read the [Contributing Guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md)
- [x] My commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) (`fix(scope):`, `feat(scope):`, etc.)
- [x] I searched for [existing PRs](https://github.com/NousResearch/hermes-agent/pulls) to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix/feature (no unrelated commits)
- [x] I've run `pytest tests/ -q` and all tests pass
- [x] I've added tests for my changes (required for bug fixes, strongly encouraged for features)
- [x] I've tested on my platform: macOS 15 (Darwin 25.5.0)

- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) per the [compatibility guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md#cross-platform-compatibility) — or N/A
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
2026-09-12 21:13:40 -07:00
Teknium
24692ee790 feat(approval): flag cloud metadata-endpoint (IMDS) credential fetches for approval
On a cloud VM the instance-metadata service hands live IAM/service-account
credentials to any local process with no auth, so a fetch against it is
credential exfiltration unless the operator expects it — yet
detect_dangerous_command() auto-approved `curl` against the link-local
metadata IP, metadata.google.internal, and the Alibaba endpoint. Add one
DANGEROUS_PATTERNS entry covering 169.254.169.254 (AWS/Azure/GCP/OpenStack),
its AWS IPv6 form fd00:ec2::254, metadata.google.internal, and Alibaba's
100.100.100.200. The host literals have no other use, so their appearance in
a command is the signal regardless of HTTP client; lookarounds keep other
169.254.x.x link-local addresses and longer host/dotted strings out.

This prompts for approval (legit uses exist on real cloud VMs); it is NOT a
hardline block. Deterministic containment-escape detection at the approval
layer, same class as the existing credential-path detectors.
2026-09-12 21:12:18 -07:00
Teknium
a565e2d493 fix(mcp): widen the SSE fallback trigger and harden its guards (salvage #53764)
Relocated onto the decomposed module layout and hardened:

- Trigger covers the rejection CLASS, not just literal 400: SSE-only
  servers' load balancers answer the chunked Streamable HTTP initialize
  POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
  rejections as an opaque -32603 'Server returned an error response'
  (error class per #104363 by @RohithPariki). Timeouts and 5xx never
  trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
  clears _ready before re-entering the transport, so the original guard
  also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
  to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
  and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
  boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
  connects + latches; no fallback on reconnect/timeout/5xx; both-fail
  error is actionable.

The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.

Fixes #53676
2026-09-12 21:05:29 -07:00
aieng-abdullah
b2465f1608 fix(mcp): auto-fallback to SSE transport when Streamable HTTP returns 400
SSE-only MCP servers (e.g. WigAI for Bitwig Studio) reject the
Streamable HTTP initialize request with 400 Bad Request, causing
permanent failure with 0 active tools. The only workaround was
manually setting transport: sse in config.

When Streamable HTTP returns 400 during initial connect, log a
warning and retry with SSE before reporting failure. Reconnects
are excluded so a genuine 400 on an established transport is not
silently masked.

Extracted inline SSE code into _run_sse() helper shared by the
explicit config path and the new fallback path.

Fixes #53676
2026-09-12 21:05:29 -07:00
Søren L. Hansen
0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
Teknium
a49a9d79b3 Port from lobehub/lobehub#19329: surface environment recreation in terminal tool results
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.

Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.

Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
2026-09-12 20:55:04 -07:00
Max Freedom Pollard
10c34dd7e2 fix(tools): stop read_file rendering a phantom empty line for newline-terminated files
_add_line_numbers split on '\n', so a file ending in a newline (the normal,
well-formed case) produced a trailing empty element that got its own line
number. read_file therefore showed a phantom '<N+1>|' line that is not in the
file, on every terminal backend and every OS, matching neither cat -n nor the
reported total_lines. Drop the single terminating newline before splitting.

Fixes #49451
2026-09-12 20:52:49 -07:00
nikkoxgonzales
71063b1dbe fix(tools): count final unterminated line in read_file total_lines
wc -l counts newlines, not lines, so a file without a trailing newline
reported one fewer total_lines than the content it returned. The read
paths already probe the last byte (file_ends_with_newline) to strip cut's
phantom newline; use that same signal at the shared assembler choke point
so total_lines, truncation, and the past-EOF guard agree on every path
(compound, sequential, native).

Fixes #3907. Supersedes #3908: single adjustment instead of a per-path
helper, covering the native and sequential paths as well.
2026-09-12 20:52:49 -07:00
joaomarcos
3966e5de94 security(state): harden async_delegation's direct state.db writer
tools/async_delegation.py:_connect() opens the same state.db as
SessionDB via a bare sqlite3.connect(), bypassing the owner-only
(0600) hardening added for SessionDB. Apply the same
_create_owner_only / _secure_wal_files policy here, reusing
hermes_state's helpers (managed/container skip included).

Addresses teknium1's review on #59716.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 20:43:42 -07:00
Hermes fleet-fix
c445987559 fix(memory): secure built-in memory lock files 2026-09-12 20:43:42 -07:00
teknium1
c6f87deb2c feat(video): OpenRouter backend covers every model on the live video catalog
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):

- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
  no duration are outside the unified video_generate surface and are dropped)
  with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
  advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
  live limits (nearest by value/height/ratio) and drops generate_audio/seed
  for models that lack them (the API 400s otherwise); reference images ride
  in input_references; local file inputs are refused (OpenRouter fetches
  URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
  never to a provider-supplied unsigned_urls host (kept from #103267)

Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.

Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
2026-09-12 13:44:52 -07:00
Xipong
3d7f773bb4 fix(kanban): honor explicit platform tool opt-ins across configuration surfaces 2026-09-12 12:32:55 -07:00
kshitijk4poor
1c671beab2 refactor(process-registry): word the degrade warning as a host-level notice
The warning fires once per process but was prefixed with the first
job's unit suffix, reading as a per-job notice for a host-level
condition. Drop the suffix; the message now says what applies to every
later dispatch.
2026-09-12 23:17:12 +05:30
kshitijk4poor
f003e449be refactor(cron): trim the scope-degrade dispatch to its invariants
Follow-up to the cherry-picked #102431 fix, addressing the review findings:

- The two real-helper scheduler tests ran the Linux-only helper unmarked
  and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
  scripts/check-windows-footguns.py --all (lint lane red). The remedy text
  is now built once in the helper and passed into the warning, so the
  "scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
  once (helper level); default config still Popens externally and
  `require_restart_safe_scope: true` raises (scheduler level, real helper).
  Dropped the stubbed duplicate, the standalone config-raise test and the
  in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
  with the same `except Exception` guard as the sibling
  `failure_nudge_threshold` read, so a config error no longer escapes
  the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
  instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
  `cron.require_restart_safe_scope` documented in the cron user guide.
2026-09-12 23:17:12 +05:30
Paul Robertson
560b6d2e81 fix(cron): degrade gracefully when systemd user scopes are unavailable
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
2026-09-12 23:17:12 +05:30
teknium1
d62716c704 fix(mcp): breaker opened by tool errors says "rejected", not "unreachable"
Application errors (isError payloads) keep counting as breaker strikes:
that is #10447's point (a server answering errors made the model hammer
it 8x in 10s) and #109180 just reasserted it. What #11113 actually hit is
the open-breaker MESSAGE: after three rejected fetches the model was told
the server was "unreachable" and went to the user instead of fixing its
URL. Track whether the streak was all application errors and word the
pause accordingly; one transport strike restores the unreachable text.
2026-09-12 09:05:56 -07:00
LehaoLin
c59c1a98e3 fix(browser): suppress KeyboardInterrupt in atexit cleanup thread stop
A second Ctrl+C while the atexit hook joins the browser janitor thread
surfaced as "Exception ignored in atexit callback: _stop_browser_cleanup_thread"
with a KeyboardInterrupt traceback. The janitor is a daemon thread and the
interpreter is already exiting, so nothing is lost by swallowing the
interrupt — the terminal tool's sibling _stop_cleanup_thread already does.

Salvaged from #10765 (function has since moved to browser_tool_lifecycle.py).

Fixes #10764
Co-authored-by: LehaoLin <lehaolin98@outlook.com>
2026-09-12 08:43:28 -07:00
teknium1
ce0b10cb21 fix(approval): anchor uninstall rules at command position, allow option operands
The package-manager uninstall patterns were bare \b-anchored, so quoted
prose (`git commit -m "document npm uninstall usage"`) prompted while the
real `npm --prefix DIR uninstall x` slipped past because the option group
did not allow an operand. Use the file's _CMDPOS anchor like every other
command-name rule and let each global option take one operand.

Found by independent review before merge.
2026-09-12 08:42:43 -07:00
konsisumer
85ce25687e fix(approval): require confirmation for package uninstalls
`npm uninstall -g`, pnpm/yarn remove, `pip uninstall` and `brew uninstall`
remove software outside the project yet matched no dangerous-command
pattern, so the agent ran them without asking (#10199). Add one
"package manager uninstall" rule per manager; installs and updates stay
unprompted.

Hand-ported from PR #64175 (the patterns moved from tools/approval.py to
tools/approval_detection.py after it was opened).
2026-09-12 08:42:43 -07:00
easyvibecoding
4e1b3daa86 fix(memory): warn when MEMORY.md / USER.md exceed their char limit on load
The cap only fires on add/replace, so an externally written over-budget
file rode silently in the system prompt while every later add was refused
with no visible cause. Warn at load; entries stay loaded (never truncate
a user's memories).

Salvage of #10886 (original hunk targeted memory_tool.py before the
store split); authored by @easyvibecoding.

Refs #10877
2026-09-12 08:30:52 -07:00
teknium1
02b398bda5 fix(mcp): recovered application errors keep the breaker strike
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (3ff18ffe14,
#10447: a server answering errors made the model hammer it 8x in 10s).
Route the recovered result through _record_call_outcome so the caller
sees the tool's answer and the counter still moves the right way.
2026-09-12 08:23:28 -07:00
Dennis Hermsmeier
5dca47a651 fix(mcp): preserve application results after transport recovery 2026-09-12 08:23:28 -07:00
teknium1
c7efbbdab2 chore(deps): mirror slack-sdk 3.44.1 pin in tools/lazy_deps.py
pyproject extras and the lazy installer must agree on the slack-sdk pin;
the salvaged bump only touched pyproject.toml + uv.lock, so the
`platform.slack` lazy-install spec would still have pulled 3.43.0.
2026-09-12 08:12:38 -07:00