Commit Graph

15997 Commits

Author SHA1 Message Date
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
teknium1
2ef929dfe0 test(config): pin load_env's dotenv inline-comment semantics
load_env now tokenizes through agent.secret_scope.load_env_file, which
strips an unquoted ` # ...` tail (dotenv standard; the behaviour the
profile secret scope already had). Pin both halves of the rule so the
value-semantics change is a stated contract, not an accident: unquoted
`abc #123` -> `abc`, quoted `"abc #123"` -> `abc #123`. Hermes' own
writer (_quote_env_value) quotes any value containing `#`, so saved
secrets round-trip.
2026-09-13 05:07:50 -07:00
teknium1
9401cc1643 test(redact): a2a superset test asserts every credential class; drop dead mask branch
The a2a invariant test skipped any synthesized token the canonical
redactor itself let through, so a boundary/shape regression would have
passed silently. Every class scrubs today, so the escape hatch goes and
the soft ">= 50" count becomes the exact registry size.

The gateway body test still asserted a "[REDACTED]" fallback marker that
no longer exists; redact_for_egress masks via _mask_token ("***"), so
assert that alone. honcho oauth.py's `re` import became unused when its
private pattern list moved to the registry.
2026-09-13 05:07:50 -07:00
teknium1
f680431f2e fix(redact): bearer residue sweep needs a 20-char token floor
redact_for_egress's bearer sweep matched any run of token characters after
the word "Bearer", so ordinary prose ("I'm the bearer of bad news") came
back as "Bearer [redacted] bad news" on every chat and A2A reply. The
gateway and A2A sweeps this PR replaced always required 20+ chars; only
monitoring was floor-less. Restore the floor on the opaque branch and keep
the bracket branch that folds an already-masked residue to one marker.
2026-09-13 05:07:50 -07:00
teknium1
dd1baee0e4 refactor(secrets): drop scope-aware env shims; runtime_provider and the voice/xai tools read the canonical getters
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.

tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.

Behavior change: none.
2026-09-13 05:07:50 -07:00
teknium1
c849bc383a refactor(env): agent.secret_scope.load_env_file is the only .env tokenizer; six hand parsers collapse onto it
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.

Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.

Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.

Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
2026-09-13 05:07:50 -07:00
teknium1
226df89f74 refactor(redact): one secret-pattern source; a2a, gateway chat and monitoring egress scrub through redact_for_egress
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.

Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.

Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).

Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
2026-09-13 05:07:50 -07:00
teknium1
3a75078c3d test(channel_directory): fault the canonical JSON serializer, not json.dump
atomic_json_write now serializes with json.dumps before touching the file (the
surrogate-escape fix), so a json.dump stub never fired and the disk-full test
silently passed the write. The invariant is unchanged (a failed write keeps the
previous cache); the fault is injected at utils._dump_json.
2026-09-13 05:07:11 -07:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
602801baa1 chore(secrets): cover the spawn ledger and meet-node token in the private-writer invariant; drop unused import
Both writers already go through atomic_json_write(mode=0o600) but were
missing from tests/test_private_credential_writers.py::_writers, so a
"write then chmod" regression in either would not be caught. Also removes
the `import os` left unused in agent/secret_sources/_cache.py (ruff F401).
2026-09-13 05:07:11 -07:00
teknium1
b2688440a9 fix(docker): rebootstrap re-seed temp is randomly named so a stale temp never blocks recovery
reseed_if_terminal created its temp as <auth>.rebootstrap.<pid>.tmp with
O_CREAT|O_EXCL. Boot-hook PIDs inside a container are near-deterministic,
so a run SIGKILL'd between create and replace leaves a same-named file and
every later boot hits FileExistsError - which main() swallows as
"error (ignored)", leaving the terminal-session recovery path dead until
someone deletes the temp by hand.

tempfile.mkstemp in the auth dir gives a random name at 0600 (stdlib only,
matching the script's no-hermes-imports rule); the fsync + os.replace +
unlink-on-failure semantics are unchanged.
2026-09-13 05:07:11 -07:00
teknium1
714013c493 fix(utils): new non-secret atomic writes follow the process umask again
Every hand-rolled writer this PR folded into utils._atomic_write created a
NEW file with write_text()/open("w"), i.e. at 0o666 masked by the umask
(0644 under 022). The canonical helper publishes through mkstemp, whose
temp is 0600, and with no explicit mode and no existing target to copy bits
from it left that 0600 in place - so debug, model_catalog, profiles,
breadcrumbs, worktree_ops, web_result_cache, plugin_compat, write_approval,
rich_sent_store, active_sessions and the google_meet state files were
silently tightened to owner-only, the volume-mount hazard
_restore_file_metadata's own docstring warns about. Undeclared in the PR.

Fix at the canonical: when mode is None and the target does not exist,
apply default_new_file_mode() (0o666 masked by the umask, read via the
umask two-call trick with a transient 0o077 so a racing thread can only get
a tighter file). The helper is hermes_cli/backup._default_new_file_mode
moved into utils and reused. Secret writers (mode=0o600) are 0600 before,
during and after as before; an existing target keeps its bits; on non-POSIX
the helper returns None so nothing is chmod'd.
2026-09-13 05:07:11 -07:00
teknium1
4157494d2f fix(utils): atomic_json_write escapes lone surrogates instead of raising UnicodeEncodeError
The canonical writer defaults to ensure_ascii=False, but ~10 of the sites
repointed onto it (terminal breadcrumbs, shell-hook allowlist, active
sessions, debug pending, model-catalog cache, the credential writers)
previously used json's ensure_ascii=True default. A surrogate-escaped str
(os.fsdecode of a non-UTF-8 cwd/argv) that json used to persist as \udcff
now made the utf-8 text handle raise UnicodeEncodeError - a ValueError that
the callers' `except OSError` never catches, so breadcrumbs silently stopped
writing and the other sites leaked a new error type.

Fix at the canonical: serialize to a str first (so nothing lands in the
temp file on failure), and on UnicodeEncodeError retry the dump with
ensure_ascii=True. That escape round-trips - json.loads returns the same
str with the lone surrogate - whereas encoding with surrogateescape emits a
raw 0xFF byte the reader's utf-8 decode rejects. The happy path is
unchanged: normal content keeps its raw UTF-8 bytes on disk.
2026-09-13 05:07:11 -07:00
teknium1
30657d197d refactor(update): psutil Android installer extracts through the shared safe-tar guard
hermes_cli/psutil_android carried its own tar path-traversal / link-member
guard (a 0.87 copy of archive_safe.safe_extract_targz). One guard for every
tar.gz we extract; the installer keeps raising PsutilAndroidInstallError so
its callers' except clauses are unchanged.

Behavior change: the psutil path now also rejects Windows-absolute and
backslash-smuggled member names (archive_safe.normalize_archive_parts), and
chmod failures on extracted files are suppressed identically.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1
2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Konstantin Khlopkov
ccd360e94f fix(state): tighten existing db files by chmod(2), not an open/fchmod/close cycle
POSIX fcntl locks are owned per (process, inode): closing any descriptor
for state.db releases every lock the process holds on that inode,
including the locks of an already-open SQLite connection. The
owner-only hardening cycle opened the live database and its -wal/-shm
read-only, fchmod'ed, and closed, so any process that already held a
connection (gateway, desktop hermes serve, dashboard share one) dropped
its live locks on every SessionDB init. A sibling process then took the
shared-memory DMS exclusively at its own close, checkpointed, and
unlinked the sidecars while long-lived holders kept the deleted inodes
open, tripping the deleted-WAL generation guard.

chmod(2) on the path never opens the file, so it cannot disturb locks.
The descriptor path remains only for first-time main-db creation, where
no locks can exist yet.
2026-09-13 04:29:06 -07:00
Teknium
d3202bbc8d feat(tui_gateway): fire agent_loop_stopped on session.interrupt too
Widens the new hook to the sibling interrupt surface: the TUI/desktop
session.interrupt path stops a live turn exactly like the gateway's
/stop, so plugins holding per-turn external resources get the same
signal there (platform='tui'). Gated on a genuinely running turn;
dispatch failures are swallowed so a plugin can never break the
interrupt. Docs updated to describe both surfaces.

Inspired by ChatGPT Work / Codex CLI 0.150.0 'Interrupt' hooks
(hooks that run when an active top-level turn is interrupted).
2026-09-12 22:19:44 -07:00
Franci Penov
f361971eed feat(gateway): fire agent_loop_stopped plugin hook on interrupt
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.

_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.

Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.

Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
2026-09-12 22:19:44 -07:00
Teknium
82199439c7 feat(image_gen): add Meta Muse Image ($0.01/img) to the FAL catalog
meta/muse-image/text-to-image + paired meta/muse-image/edit, the FAL
listing of Meta's Muse Image model (launched on the Meta Model API in
Aug 2026 at $0.01/image).

- aspect_ratio size family (16:9 / 1:1 / 9:16 from the vendor's
  21:9..9:21 enum); always sent on t2i for deterministic framing,
  deliberately omitted on edits so Muse follows the input image.
- No seed in the vendor schema (Grok Imagine 2.0 precedent) - the
  supports whitelist filters it.
- Edit takes 1-10 reference image_urls (max_reference_images=10).

Schema verified against FAL's OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403), same as prior catalog
additions.
2026-09-12 22:18:06 -07:00
Teknium
a5522f69c0 feat(slack): route status/title through the Agent Sessions API (slack-sdk 3.44.0)
Slack deprecates the Assistant messaging experience (assistant_view) in
February 2027: assistant.threads.setStatus/setTitle are replaced by
agents.sessions.setStatus/rename. slack-sdk 3.44.0 (Aug 27 2026) ships
the typed methods with drop-in-compatible signatures.

- adapter: capability probe on the AsyncWebClient CLASS (never instance —
  mock auto-attributes lie), cached; status set/clear + thread title route
  through agents.sessions.* when available, legacy otherwise
- pins: slack-sdk 3.43.0 -> 3.44.0 (pyproject messaging+slack extras,
  lazy_deps, uv.lock)
- tests: autouse fixture pins the probe to legacy under the mocked SDK;
  5 new tests cover both routing paths for typing, clear, and title
- docs: slack.md scope table + status-line notes mention both methods
2026-09-12 22:15:11 -07:00
teknium1
ae38c903b1 test(cron): unreachable-retry ladder test uses an interval schedule, not a wall-clock cron
"0 3 * * *" fires at a fixed UTC minute; when CI runs in the half hour before
it (observed 02:30:56Z), the 30-minute rung lands after the natural occurrence
and plan_retry correctly yields to the schedule, clearing the state the test
asserts on. "every 24h" always has its natural fire a full day out, so every
rung is strictly earlier regardless of when the test runs.
2026-09-12 22:13:33 -07:00
Teknium
ec58e08a35 feat(cron): automatic bounded re-runs when a fire never reached the model
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."

A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.

Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.

- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
  notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
2026-09-12 22:13:33 -07:00
Teknium
23af232837 fix(tools): reject malformed tool parameter schemas at registration
Port from earendil-works/pi#9300 fix (acaa253cc): a plugin registering a
tool whose schema["parameters"] is not a dict (a list, string, etc.)
previously registered fine and the malformed schema was serialized into
every provider request, 400-ing turns far from the offending plugin.
Live probe on main confirmed the bad schema flows into _fn_def() and the
OpenAI wire unchanged.

Fail at registry.register() with the tool name in the error instead. The
plugin loader already catches registration exceptions and marks the
plugin errored, so a broken plugin degrades gracefully rather than
breaking every session. Schemas that omit "parameters" stay valid
(no-argument tools); MCP tools are unaffected (their schemas pass
through _normalize_mcp_input_schema first, which always returns a dict).
2026-09-12 22:11:56 -07:00
Teknium
7f61ae589f Port from openai/codex#41436: answer blocking terminal queries in background PTY sessions
Programs run via terminal(background=true, pty=true) can block forever when
they probe their terminal — device-status (ESC[5n), window-size (ESC[18t),
cursor-position (ESC[6n), or DEC private-mode (ESC[?N$p) queries — because
nothing on the PTY master side answers, and the raw query bytes leak into
captured output.

- tools/pty_query_responder.py: incremental byte scanner that strips the
  handled queries from PTY output (chunk splits included) and produces
  bounded replies; everything else passes through untouched.
- tools/process_registry.py: wire the responder into _pty_reader_loop
  (POSIX only — ConPTY answers its own queries); flush partial escape
  tails at end-of-stream.
- tests mirror the codex fixtures plus a live-PTY E2E where a subprocess
  blocks on ESC[6n until answered.
2026-09-12 22:09:19 -07:00
Teknium
0c2e66ea8b test(auth): OpenRouter PKCE invariants, fake-authority A/B harness, docs, contributor map
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
2026-09-12 22:07:41 -07:00
Indigo Karasu
41380ccef9 fix(process): list refreshes no longer leave exited children running
Carve the direct-child list reconciliation from Indigo Karasu's earliest
PR #60506 (2a96ae2cbf806ccdc3e9b584911774f32622f421), corroborated by
fangliquanflq's narrow #81385 (50ffd243d92627e4a03a3ee8427ad0f5e090c3ab).
Run the existing helper after task/session filtering and reuse the
idempotent owner-stamped completion path. Do not import cross-session
disclosure, bare-PID healing, forget RPCs, or reader rewrites.

A real-child regression fails before this change and passes after it:
the direct child exits while its descendant keeps writing to stdout;
listing reports exit without consuming the result or waiting for EOF.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-12 22:06:03 -07:00
teknium1
8dc4b6f565 test(approval): reconcile fixture uses the per-home baseline map
The multiplex-scoped rebase replaced the flat _permanent_baseline set with
_permanent_baseline_by_home (keyed by profile home, "" = unscoped); the
fixture must reset and seed that map.
2026-09-12 22:03:19 -07:00
Teknium
2aba244ff5 test: trim reconcile suite to the invariant cases 2026-09-12 22:03:19 -07:00
shuwen.wu
9db604ce38 fix(tools): reload permanent allowlist by replacement 2026-09-12 22:03:19 -07:00
durden
98d54ee7b3 docs(approval): say that save_permanent_allowlist can only add, and that a revocation waits for the next write
Review follow-up. Two of the three items taken as written; the third declined
with a reason.

1. Taken. The reconcile semantics mean `patterns` may only ADD -- an entry left
   out of it is not removed, because the on-disk list wins for anything this
   process did not approve itself. Every caller in the tree is additive today,
   so nothing breaks, but the signature does not say so. Stated in the
   docstring, and pinned by
   `test_a_caller_that_passes_a_smaller_set_does_not_remove` so a future
   `allowlist remove` finds out here instead of in production.

   NOT taken: the `reconcile: bool = True` opt-out. There is no caller that
   wants it, and AGENTS.md:98-101 names exactly this -- "Speculative
   infrastructure. Hooks, callbacks, or extension points with no concrete
   consumer." The removal path is editing config.yaml, which the docstring now
   says.

2. Taken. website/docs/user-guide/security.md, next to the existing
   `hermes config edit` tip, which is where an operator reads about removing a
   pattern: the list is read at startup, a pattern removed while a session is
   running stays approved in that session until the next write or a restart,
   and if it was removed for safety reasons, restart.

3. Taken. `test_save_failure_is_logged_not_raised` asserted non-raising but
   never asserted the log its name promises. Now asserts
   "Could not save allowlist" via caplog.

    scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
    === Summary: 1 files, 9 tests passed, 0 failed (100% complete) in 0.4s
2026-09-12 22:03:19 -07:00
durden
9b06d3d081 fix(approval): saving the allowlist deletes entries the operator added by hand and resurrects ones they revoked
`load_permanent_allowlist()` runs exactly once, at module import
(tools/approval.py, the call at the bottom of the module), and
`load_permanent()` only unions into `_permanent_approved` (:2866-2869) --
nothing ever removes. `save_permanent_allowlist()` then wrote that in-memory
set straight back over `config["command_allowlist"]`, at eight call sites.

`command_allowlist` is a file the operator edits, and deleting a line from it
is the documented way to withdraw a standing approval. Any hand edit made
while a Hermes process is live was undone by that process's next `[a]lways`,
in both directions at once.

Reproduced on this tree with a temp HERMES_HOME:

    BEFORE (tools/approval.py at fcbd107)
      on disk before this process starts : ['git status', 'ls *']
      operator edits config.yaml by hand : ['ls *', 'npm test']
                                            (revoked 'git status', added 'npm test')
      after ONE [a]lways                 : ['docker *', 'git status', 'ls *']
      is_approved still honours revoked? : True

    AFTER
      after ONE [a]lways                 : ['docker *', 'ls *', 'npm test']
      is_approved still honours revoked? : False

`npm test` was silently deleted from the operator's own config file, and
`git status` -- a standing approval they had just withdrawn -- was written
back and kept auto-approving. Neither prints anything.

The same shape loses writes between two live Hermes processes: whichever
saves second overwrites the other's entry.

The fix reconciles at write time. The file is re-read and the result is what
is on disk now, plus what this process approved since its own baseline, where
the baseline is what `command_allowlist` held the last time this process
synchronised with the file. That difference is what separates "the operator
granted this here" from "this was on disk at import and may since have been
revoked". Revoked entries are also dropped from `_permanent_approved` so
`is_approved()` stops honouring them for the rest of the process.

It does NOT make a revocation take effect the instant the file changes --
nothing re-reads the file on the approval hot path, and adding a stat there is
a separate change with its own cost. It makes the next write stop undoing the
operator's edit.

`_lock` is `threading.Lock` and not reentrant; all eight call sites were
checked and none holds it across the call, so the added critical section
cannot deadlock. The failure path still logs and returns rather than raising,
as before.

Searched open and merged PRs and issues for `command_allowlist revoke`,
`permanent allowlist reload`, `approval allowlist clobber` and
`save_permanent_allowlist` -- nothing covers this.

Tests: tests/tools/test_permanent_allowlist_reconcile.py, 8 cases -- both
halves of the bug, the two-process race, idempotence, the unedited round trip,
and the existing contract that a config write failure is logged rather than
raised.

    scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
    === Summary: 1 files, 8 tests passed, 0 failed (100% complete) in 0.5s

No regression across the 29 test files in tests/ that touch the allowlist or
the approval module: 25 failed before and after, byte-identical failure set
(pre-existing missing-dependency failures in my local venv).
2026-09-12 22:03:19 -07:00
teknium1
08a2e7dbcc refactor(cli): vim mode is a config key only — drop the /vim slash command
Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
2026-09-12 22:00:02 -07:00
Teknium
4f03480787 test(cli): add vim to the tracked inline-handled command list 2026-09-12 22:00:02 -07:00
Sam Foreman
968fdc4746 fix(cli): apply /vim salvage to decomposed layout, trim tests, docs
- Relocate _handle_vim_command and _vim_mode_label onto
  cli_status_bar_mixin.py (the mixin owning status-bar commands and
  rendering) — the original targeted pre-decomposition cli.py; slash
  dispatch picks the handler up by naming convention
- Wire the vim label into the current status-bar fragment builder
- Trim tests to 3 invariant tests per the salvage bar
- Document /vim in reference/slash-commands.md

Live E2E (tmux + PTY, temp HERMES_HOME): /vim toggles on, Esc/i flip
NORMAL/INSERT in the status bar, /vim off restores emacs bindings,
display.vim_mode persisted to config.yaml.
2026-09-12 22:00:02 -07:00
Sam Foreman
1571752a70 test(cli): cover /vim toggle, persistence, and status label
Adds 11 tests for the vim mode surface:

- /vim status reports without persisting
- bare /vim toggles; on|off set explicitly; each persists to
  display.vim_mode
- editing_mode is applied to a running Application in both directions
- a missing Application is not fatal (toggle before the TUI starts)
- invalid arguments leave state untouched and print usage
- _vim_mode_label() maps vi input modes to NORMAL/INSERT/REPLACE and
  stays empty when disabled or before the app exists
- the config default is off and /vim is registered with subcommands
2026-09-12 22:00:02 -07:00
Teknium
63584da036 feat(video): add MiniMax H3 Max Turbo family (fal post-train, 480P-1080P)
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.

Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
2026-09-12 21:58:25 -07:00
Teknium
e21a6fb159 fix(tools): make the empty/whitespace old_string rejection actionable
Port from cline/cline#13970: models that send patch calls with an empty
old_string got back 'old_string cannot be empty' — an error that names the
problem but not the recovery, so the next call was byte-identical and the
run burned turns until loop detection killed it (upstream repro: Kimi K3
looping on old_text: null).

The rejection now states the recovery: set old_string to the exact text the
replacement should replace, read the file first if unsure, use write_file
for new files/full rewrites, and do not re-send the call unchanged. The
whitespace-only rejection gets the same treatment. No behavior change for
valid calls.
2026-09-12 21:55:07 -07:00
Teknium
92df11f81d fix(classifier): terminal_quota_exhausted 429s classify as billing, not rate_limit
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).

LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.

- `_status_429` now honors a structured billing code first (decisive
  signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
  (429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
  limit" was already there; providers use both orders). "terminal billing
  limit" text is deliberately NOT matched: substring rules cannot negate
  the "non-terminal billing limit" wording — the structured code covers it.
2026-09-12 21:52:48 -07:00
Teknium
833594cda6 test: opt the ClawHub owner-HTTP fixture into private URLs for the SSRF-guarded stream
The streamed download path now routes through the SSRF-safe client, which
(correctly) refuses the test fixture's 127.0.0.1 registry. Set
HERMES_ALLOW_PRIVATE_URLS for the fixture's lifetime and reset the module
cache on both sides so the guard still fail-closes everywhere else.
2026-09-12 21:50:30 -07:00
Eugeniusz Gilewski
fef98ff00f fix(skills): bound streamed ClawHub ZIP downloads (#57571)
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.

Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.

Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
2026-09-12 21:50:30 -07:00
Teknium
7b037f0efa feat(cli): opt-in git_branch status-bar field (⎇ current branch)
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.

Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.

Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
2026-09-12 21:47:53 -07:00
Teknium
053f8b1b17 fix(kanban): make archive-time worker termination race-safe and audited
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:

- Snapshot status/pid/claim INSIDE the archive txn so the kill only
  happens when this caller wins the archive transition; a losing
  concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
  skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
  hold the SQLite write lock. Safe because archived is terminal — no
  dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
  event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
  -> no signal, no event), live E2E: worker survived archive on main,
  terminated (<0.3s, clean SIGTERM) with the fix.

Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
2026-09-12 21:45:15 -07:00
Teknium
77f0c83ec3 fix(sessions): honor CLAUDE_CONFIG_DIR and CODEX_HOME in foreign session discovery
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.

_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.

Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
2026-09-12 21:39:10 -07:00
liuhao1024
0e13fa98ec fix(web): cap web_extract provider dispatch with a wall-clock timeout (salvage #57180)
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.

Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).

Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."

Fixes #57155

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-12 21:30:38 -07:00
salch-cred
ad0398eed8 fix(kanban): preserve sticky block on tasks created with initial_status=blocked (#107398) 2026-09-12 21:27:21 -07:00
Teknium
cbd4492f1f fix(tools): steer/redirect releases a blocking process_manage wait (port of MoonshotAI/kimi-code#3697)
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).

wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.

Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.

Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
2026-09-12 21:25:00 -07:00
Teknium
50f17b1300 fix(google-workspace): support multi-tab Google Docs (port cloudflare/cloudflare-os#450)
Google Docs can hold a tree of tabs, each with its own body and its own
1-based character index space; the legacy top-level `body` only carries
the first tab. `docs get` was silently dropping every other tab's
content, and `docs append` computed its insert index from the first
tab's body and sent the write with no tabId — so on a tabbed doc the
append could land at a wrong offset in the wrong tab.

- Requests now pass includeTabsContent=true (both gws and SDK paths).
- `docs get` returns a `tabs` array (preorder flattening of the
  tabs/childTabs tree, nested tabs included); single-tab docs keep the
  `body` field so existing callers work, and `--tab <tabId>` reads one
  tab. Legacy no-tabs responses are unchanged.
- `docs append` targets exactly one tab: the insert location carries
  the tabId, the end index is computed inside that tab's own body, an
  unknown `--tab` errors instead of falling back to the first tab, and
  a multi-tab doc without `--tab` errors with the tab list rather than
  guessing. Tabs are never merged — index spaces are independent.

Adapted from cloudflare/cloudflare-os#450 (gatekeeper-google), which
fixed the same provider behavior: reads must traverse Document.tabs and
every write Location must carry the immutable tabId.

Two pre-existing bare read_text/write_text in the touched test file
gained encoding="utf-8" (windows-footgun sweep rule).
2026-09-12 21:20:42 -07:00
Teknium
11d761eed9 fix(stream): flush residual SSE buffer at EOF and widen resilience to the Gemini native adapter
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.

- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
  after the read loop; surface EOF-without-[DONE] (warn + error result
  when nothing was received); extract _consume_sse_line so line parsing
  and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
  flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
  proven red on origin/main.
2026-09-12 21:18:24 -07:00