Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).
LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.
- `_status_429` now honors a structured billing code first (decisive
signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
(429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
limit" was already there; providers use both orders). "terminal billing
limit" text is deliberately NOT matched: substring rules cannot negate
the "non-terminal billing limit" wording — the structured code covers it.
The streamed download path now routes through the SSRF-safe client, which
(correctly) refuses the test fixture's 127.0.0.1 registry. Set
HERMES_ALLOW_PRIVATE_URLS for the fixture's lifetime and reset the module
cache on both sides so the guard still fail-closes everywhere else.
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.
Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.
Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.
Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.
Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:
- Snapshot status/pid/claim INSIDE the archive txn so the kill only
happens when this caller wins the archive transition; a losing
concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
hold the SQLite write lock. Safe because archived is terminal — no
dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
-> no signal, no event), live E2E: worker survived archive on main,
terminated (<0.3s, clean SIGTERM) with the fix.
Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.
_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.
Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.
Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).
Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."
Fixes#57155
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).
wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.
Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.
Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
Google Docs can hold a tree of tabs, each with its own body and its own
1-based character index space; the legacy top-level `body` only carries
the first tab. `docs get` was silently dropping every other tab's
content, and `docs append` computed its insert index from the first
tab's body and sent the write with no tabId — so on a tabbed doc the
append could land at a wrong offset in the wrong tab.
- Requests now pass includeTabsContent=true (both gws and SDK paths).
- `docs get` returns a `tabs` array (preorder flattening of the
tabs/childTabs tree, nested tabs included); single-tab docs keep the
`body` field so existing callers work, and `--tab <tabId>` reads one
tab. Legacy no-tabs responses are unchanged.
- `docs append` targets exactly one tab: the insert location carries
the tabId, the end index is computed inside that tab's own body, an
unknown `--tab` errors instead of falling back to the first tab, and
a multi-tab doc without `--tab` errors with the tab list rather than
guessing. Tabs are never merged — index spaces are independent.
Adapted from cloudflare/cloudflare-os#450 (gatekeeper-google), which
fixed the same provider behavior: reads must traverse Document.tabs and
every write Location must carry the immutable tabId.
Two pre-existing bare read_text/write_text in the touched test file
gained encoding="utf-8" (windows-footgun sweep rule).
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.
- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
after the read loop; surface EOF-without-[DONE] (warn + error result
when nothing was received); extract _consume_sse_line so line parsing
and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
proven red on origin/main.
Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.
1. `[DONE]` marker didn't break the outer chunk loop
---------------------------------------------------
The `break` on `[DONE]` only exited the inner line-parse `while`,
leaving the outer `async for chunk in resp.content.iter_any():` to
keep reading. If the upstream held the connection open after
`[DONE]` (buggy proxy, crashed server, network hang), the client
waited up to sock_read=1800 seconds (30 min) for the next chunk.
Fix: set a `done` flag when `[DONE]` is seen and check it at the
top of the outer loop.
2. No TCP connect timeout
-----------------------
`ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
the default `None` (no timeout). An unreachable proxy host (DNS
fail, firewall, remote down) would hang on TCP connect for the OS
default (minutes) before surfacing an error to the user.
Fix: add `sock_connect=30` so connect failures surface within 30s.
3. SSE JSON parse exception handling was too narrow
-------------------------------------------------
The inner parse caught only `json.JSONDecodeError`. A response like
`{"choices": [null]}` parsed successfully, then
`choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
object has no attribute 'get'`. That bubbled up to the outer
`except Exception`, aborting the entire stream — any further chunks
were lost, and the user saw the accumulated partial response
without knowing why.
Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
chunk now skips, the stream keeps parsing.
New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:
- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
chunks and verifies both good chunks are captured, bad ones skipped
A V4A `*** Add File:` operation is meant to create a new file. The apply
path called `write_file` unconditionally and `_validate_operations` had no
pre-check for ADD, so an Add targeting a path that already existed
overwrote the file with only the patch's `+` lines, returned success, and
emitted a `--- /dev/null` diff that hid what was lost. Models frequently
confuse Add with Update, so this destroyed existing file contents with no
error. The MOVE path already guards its destination against clobbering;
ADD now follows the same rule.
Makes a V4A `Add File` operation fail when its target already exists,
instead of silently overwriting the existing file. Validation now rejects
the operation before any write happens, so the two-phase
validate-then-apply contract ("no files were modified" on a validation
failure) holds for ADD as it already does for UPDATE/MOVE/DELETE. A
matching re-check in the apply phase closes the validate-to-apply race.
N/A
- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- `tools/patch_parser.py`: add an ADD branch in `_validate_operations`
that errors when `read_file_raw` finds an existing file, and a
defensive existence re-check in `_apply_add` before `write_file`,
mirroring the existing MOVE destination guard.
- `tests/tools/test_patch_parser.py`: add a test asserting an Add onto an
existing path fails validation and leaves the original bytes unwritten;
add `read_file_raw` to three ADD-path LSP fakes so they match the real
`file_ops` interface now exercised on ADD.
1. Build a V4A patch with `*** Add File: <path>` where `<path>` already
exists on disk.
2. Apply it via `apply_v4a_operations`. Before this change the file is
overwritten with the patch's `+` lines and the result is success;
after, the result is a validation failure and the file is untouched.
3. Run `scripts/run_tests.sh tests/tools/test_patch_parser.py` —
`TestApplyOperations::test_add_onto_existing_file_fails_and_preserves_contents`
covers the regression.
- [x] I've read the [Contributing Guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md)
- [x] My commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) (`fix(scope):`, `feat(scope):`, etc.)
- [x] I searched for [existing PRs](https://github.com/NousResearch/hermes-agent/pulls) to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix/feature (no unrelated commits)
- [x] I've run `pytest tests/ -q` and all tests pass
- [x] I've added tests for my changes (required for bug fixes, strongly encouraged for features)
- [x] I've tested on my platform: macOS 15 (Darwin 25.5.0)
- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) per the [compatibility guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md#cross-platform-compatibility) — or N/A
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
On a cloud VM the instance-metadata service hands live IAM/service-account
credentials to any local process with no auth, so a fetch against it is
credential exfiltration unless the operator expects it — yet
detect_dangerous_command() auto-approved `curl` against the link-local
metadata IP, metadata.google.internal, and the Alibaba endpoint. Add one
DANGEROUS_PATTERNS entry covering 169.254.169.254 (AWS/Azure/GCP/OpenStack),
its AWS IPv6 form fd00:ec2::254, metadata.google.internal, and Alibaba's
100.100.100.200. The host literals have no other use, so their appearance in
a command is the signal regardless of HTTP client; lookarounds keep other
169.254.x.x link-local addresses and longer host/dotted strings out.
This prompts for approval (legit uses exist on real cloud VMs); it is NOT a
hardline block. Deterministic containment-escape detection at the approval
layer, same class as the existing credential-path detectors.
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.
Two-part fix:
1. title_generator.py: Pass reasoning_config={"enabled": False} to
call_llm() so thinking is explicitly disabled for title generation.
2. chat_completions.py: In _build_gemini_thinking_config, when
reasoning is disabled (enabled=False or effort="none"), set
thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
includeThoughts: False only hides thought parts from the response
while the model still reasons internally and bills thought tokens
against maxOutputTokens. thinkingBudget: 0 truly disables thinking
so thought tokens do not consume the max_tokens budget.
Fixes#91927
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).
Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
Relocated onto the decomposed module layout and hardened:
- Trigger covers the rejection CLASS, not just literal 400: SSE-only
servers' load balancers answer the chunked Streamable HTTP initialize
POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
rejections as an opaque -32603 'Server returned an error response'
(error class per #104363 by @RohithPariki). Timeouts and 5xx never
trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
clears _ready before re-entering the transport, so the original guard
also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
connects + latches; no fallback on reconnect/timeout/5xx; both-fail
error is actionable.
The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.
Fixes#53676
SSE-only MCP servers (e.g. WigAI for Bitwig Studio) reject the
Streamable HTTP initialize request with 400 Bad Request, causing
permanent failure with 0 active tools. The only workaround was
manually setting transport: sse in config.
When Streamable HTTP returns 400 during initial connect, log a
warning and retry with SSE before reporting failure. Reconnects
are excluded so a genuine 400 on an established transport is not
silently masked.
Extracted inline SSE code into _run_sse() helper shared by the
explicit config path and the new fallback path.
Fixes#53676
Port from openclaw/openclaw#140531: users paste the application ID from the
Developer Portal's General Information page instead of the bot token (Bot
page); the gateway then fails at runtime with an opaque 401. A real bot token
is dot-separated base64 and never purely numeric, so the setup wizard now
rejects an all-digit answer with pointed guidance and re-prompts once. A
second consecutive numeric answer is kept (user override), and non-numeric
tokens are saved exactly as before.
Adapted to hermes: the guard lives in the Discord plugin's interactive_setup
(the active setup path per the #9983 review — legacy _setup_standard_platform
is not the dispatch target), mirroring the existing Telegram token-shape
validation in hermes_cli/setup_platforms.py.
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).
Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.
Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.
Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.
Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
Two long-standing, mutually-masking defects in read_file's line accounting,
present on all three read paths (compound shell probe, sequential probes,
native):
1. total_lines came from `wc -l`, which counts newline bytes: a file whose
final line has no trailing newline was undercounted by one. For an
N*limit+1-line file read in pages, the last page was never offered
(truncated=False) and the past-EOF guard refused `offset=total` reads of
the real final line.
2. _add_line_numbers() split on '\n' without dropping the single terminating
newline, so every well-formed newline-terminated file rendered a phantom
`N+1|` empty gutter line that does not exist in the file (models routinely
tried to patch/reference it). Exactly one terminator is dropped, so a
genuinely selected trailing blank line keeps its number (`cat -n`
semantics).
Both fixes land at the shared choke point (_assemble_read_result /
_add_line_numbers) so the compound, sequential, and native paths agree.
Existing tests that froze the buggy rendering as expected output are updated
to the corrected contract.
Salvages community PRs #106888 (@nikkoxgonzales) and #49453
(@MaxFreedomPollard); same bug class independently reported/fixed in
#3907/#3908, #3927, #20814, #22945, #42929, #55696, #91306.
Cross-validated against anomalyco/opencode#47420's read-page serialization
fix (their trailing-blank-line class; hermes' page path is already
blank-line-safe once the terminator handling is right — verified live).
_add_line_numbers split on '\n', so a file ending in a newline (the normal,
well-formed case) produced a trailing empty element that got its own line
number. read_file therefore showed a phantom '<N+1>|' line that is not in the
file, on every terminal backend and every OS, matching neither cat -n nor the
reported total_lines. Drop the single terminating newline before splitting.
Fixes#49451
wc -l counts newlines, not lines, so a file without a trailing newline
reported one fewer total_lines than the content it returned. The read
paths already probe the last byte (file_ends_with_newline) to strip cut's
phantom newline; use that same signal at the shared assembler choke point
so total_lines, truncation, and the past-EOF guard agree on every path
(compound, sequential, native).
Fixes#3907. Supersedes #3908: single adjustment instead of a per-path
helper, covering the native and sequential paths as well.
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.
Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.
Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).
Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."
Fixes#97460
Co-authored-by: 686f6c61 <github@00b.tech>
tools/async_delegation.py:_connect() opens the same state.db as
SessionDB via a bare sqlite3.connect(), bypassing the owner-only
(0600) hardening added for SessionDB. Apply the same
_create_owner_only / _secure_wal_files policy here, reusing
hermes_state's helpers (managed/container skip included).
Addresses teknium1's review on #59716.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The deepseek plugin is bundled and always loads, so the try/except around
the delegation import was defense-in-depth for a path that cannot fail.
Tests reduced to the two contracts that matter: /reasoning none reaches
the wire as thinking.disabled, and DeepSeek ids produce exactly the native
DeepSeek profile's output while non-DeepSeek families stay a no-op.
The bundled-plugin loader pops half-registered modules when a plugin
fails to load, so the lazy 'from plugins.model_providers.deepseek import
deepseek' could raise ImportError on every DeepSeek-routed CommandCode
turn — turning the soft 'thinking uncontrollable' bug into a hard
crash. Catch ImportError, log, and return the pre-fix no-op (review
feedback on #95241).
CommandCode fronts DeepSeek with vendor-prefixed ids
(deepseek/deepseek-v4-flash). DeepSeek V4+ defaults to thinking mode
when the thinking field is omitted, so /reasoning none changed the
Hermes session state but not the actual request -- the turn sat in
reflecting.../brainstorming... for minutes (#95232). Strip the vendor
prefix for DeepSeek-family ids and delegate to the native DeepSeek
profile's build_api_kwargs_extras (extra_body.thinking +
reasoning_effort mapping); other CommandCode model families keep the
base no-op behavior. The prior no-op tests codified the bug and are
rewritten to pin the new contract.
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.
`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).
Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.
`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.
The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).
Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.
Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:
- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
model_switch applied it. Version bumps, dated snapshots and qualifiers
all sit inside 0.9 of a sibling, so a newer release the listing lacked
was swapped for the older one under the user's label. The matcher now
does exact membership -> suggestion text only; the id goes to the wire
verbatim and a genuine typo is refused with the listed siblings named.
Every branch that carried the correction (live listing, static catalog,
curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
loses it in one place.
- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
bare key (the slug Desktop picker rows carry) validated as a built-in
and hit the hard-rejecting live-listing branch; the same endpoint as
`custom:<key>` soft-accepted. Both spellings now validate as the user's
custom endpoint.
- apps/desktop: `manualPickRemoved` (composer reseed) and
`reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
sticky pick to the profile default / the row's first model whenever the
provider row did not list it. Rows are hints (discovered, curated,
capped); the gateway's switch result is the only authority on a pick.
Both helpers are removed; the pick stays put.
Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):
- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
no duration are outside the unified video_generate surface and are dropped)
with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
live limits (nearest by value/height/ratio) and drops generate_audio/seed
for models that lack them (the API 400s otherwise); reference images ride
in input_references; local file inputs are refused (OpenRouter fetches
URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
never to a provider-supplied unsigned_urls host (kept from #103267)
Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.
Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
Salvage follow-up to Xipong's #107736. Kept the core: `kanban` is a
configurable, default-off toolset whose check_fn answers the schema
build's own selection (ContextVar) instead of the legacy top-level
`toolsets` key, so `platform_toolsets.<platform>: [.., kanban]` — what
`hermes tools enable kanban --platform X` writes — actually reaches the
gateway agent's tool schema.
Dropped the `tui_gateway/server.py` change: turning an explicitly empty
CLI selection from "all" into "nothing" is a separate behaviour flip
already tracked by #107452, not part of this bug. The two TUI loader
tests that asserted `kanban` is auto-recovered onto a saved `[memory]`
list now assert the opposite: a configurable opt-in is never recovered.
Follow-up to the cherry-picked #102431 fix, addressing the review findings:
- The two real-helper scheduler tests ran the Linux-only helper unmarked
and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
scripts/check-windows-footguns.py --all (lint lane red). The remedy text
is now built once in the helper and passed into the warning, so the
"scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
once (helper level); default config still Popens externally and
`require_restart_safe_scope: true` raises (scheduler level, real helper).
Dropped the stubbed duplicate, the standalone config-raise test and the
in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
with the same `except Exception` guard as the sibling
`failure_nudge_threshold` read, so a config error no longer escapes
the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
`cron.require_restart_safe_scope` documented in the cron user guide.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).
Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.
The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.
Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.
Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
Replace the POSIX-only jobs-flock contention test (skipped off-POSIX,
~120 LOC of monkeypatched flock plumbing) with a single invariant test
that fails on pre-fix code in <1s: hold the per-job fire fence from a
worker thread, assert the heartbeat still returns True on the calling
thread, and that a takeover is still detected (False). The docstring on
heartbeat_fire_claim now records WHY it is not under the fence, so the
next refactor does not put it back.
Co-authored-by: Oliver Heckmann <46627487+oheckmann74@users.noreply.github.com>
Co-authored-by: salch-cred <141555468+salch-cred@users.noreply.github.com>
heartbeat_fire_claim only CAS-refreshes claim.at via _with_job; wrapping
_under_fire_fence across save_jobs let a blocked .jobs.lock pin the fence
and cause mark_job_run to fail closed on completed jobs.
Replace the two stubbed tests with a single test that drives the real
CodexTransport.build_kwargs (the fixture agent already carries a tool), so
the test binds the summary body actually sent, not a hand-written dict. It
covers both the first attempt and the empty-summary retry; the previous
pair asserted the same three keys twice and reused one dict across
attempts, so the second-attempt check passed trivially after the first pop.
Add the WHY comment on the pops: the transport emits tools, tool_choice and
parallel_tool_calls as one block, and strict Responses backends 400 on the
controls without tools.
Iteration-limit summaries retry once when the first attempt comes back empty.
Both attempts share one builder closure, so assert the retry body is as
tool-free as the first: forces an empty summary, then checks every captured
codex request for tools / tool_choice / parallel_tool_calls.
Addresses the retry-coverage request on #32777.