Commit Graph

34080 Commits

Author SHA1 Message Date
Sam Foreman
2a5373da2c feat(cli): add display.vim_mode config and register /vim command
Introduces the opt-in surface for vi editing in the input composer:

- display.vim_mode config default (False, so existing users are
  unaffected and prompt_toolkit keeps its standard emacs bindings)
- /vim command registered with on|off|status subcommands, matching
  the established /battery and /timestamps pattern

Part of #4254.
2026-09-12 22:00:02 -07:00
Teknium
63584da036 feat(video): add MiniMax H3 Max Turbo family (fal post-train, 480P-1080P)
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.

Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
2026-09-12 21:58:25 -07:00
Teknium
e21a6fb159 fix(tools): make the empty/whitespace old_string rejection actionable
Port from cline/cline#13970: models that send patch calls with an empty
old_string got back 'old_string cannot be empty' — an error that names the
problem but not the recovery, so the next call was byte-identical and the
run burned turns until loop detection killed it (upstream repro: Kimi K3
looping on old_text: null).

The rejection now states the recovery: set old_string to the exact text the
replacement should replace, read the file first if unsure, use write_file
for new files/full rewrites, and do not re-send the call unchanged. The
whitespace-only rejection gets the same treatment. No behavior change for
valid calls.
2026-09-12 21:55:07 -07:00
Teknium
92df11f81d fix(classifier): terminal_quota_exhausted 429s classify as billing, not rate_limit
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).

LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.

- `_status_429` now honors a structured billing code first (decisive
  signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
  (429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
  limit" was already there; providers use both orders). "terminal billing
  limit" text is deliberately NOT matched: substring rules cannot negate
  the "non-terminal billing limit" wording — the structured code covers it.
2026-09-12 21:52:48 -07:00
Teknium
833594cda6 test: opt the ClawHub owner-HTTP fixture into private URLs for the SSRF-guarded stream
The streamed download path now routes through the SSRF-safe client, which
(correctly) refuses the test fixture's 127.0.0.1 registry. Set
HERMES_ALLOW_PRIVATE_URLS for the fixture's lifetime and reset the module
cache on both sides so the guard still fail-closes everywhere else.
2026-09-12 21:50:30 -07:00
Eugeniusz Gilewski
fef98ff00f fix(skills): bound streamed ClawHub ZIP downloads (#57571)
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.

Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.

Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
2026-09-12 21:50:30 -07:00
Teknium
7b037f0efa feat(cli): opt-in git_branch status-bar field (⎇ current branch)
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.

Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.

Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
2026-09-12 21:47:53 -07:00
Teknium
bef5b7c2c3 chore: map asaxinw@gmail.com -> PINKIIILQWQ (salvage #42858) 2026-09-12 21:45:15 -07:00
Teknium
053f8b1b17 fix(kanban): make archive-time worker termination race-safe and audited
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:

- Snapshot status/pid/claim INSIDE the archive txn so the kill only
  happens when this caller wins the archive transition; a losing
  concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
  skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
  hold the SQLite write lock. Safe because archived is terminal — no
  dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
  event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
  -> no signal, no event), live E2E: worker survived archive on main,
  terminated (<0.3s, clean SIGTERM) with the fix.

Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
2026-09-12 21:45:15 -07:00
PINKIIILQWQ
907f409dc4 fix(kanban): archive_task now kills the worker process
archive_task() was a pure DB operation — it cleared worker_pid,
claim_lock, and status from the tasks row but never sent SIGTERM
to the actual OS process. A running worker stayed alive until it
next called kanban_complete/kanban_block and discovered it was
archived, burning API quota and compute resources.

Fix: snapshot pid+claim_lock before the write_txn clears them,
then call _terminate_reclaimed_worker() — same function reclaim_task
uses — which sends SIGTERM, waits 5s, then SIGKILL if still alive.
The termination metadata is included in the 'archived' event so
operators can see what happened.

Order matches reclaim_task: terminate first, then DB update.
For non-running / non-local tasks, _terminate_reclaimed_worker
returns immediately as a no-op.

Closes #33774 reprise: the scratch-workspace side was fixed in
fc8afd500, but the orphaned-process side was never addressed.
2026-09-12 21:45:15 -07:00
Teknium
47c029927f feat(skills): add dream-loop — concept-art visual fidelity build loop (port of achimala/dream-loop, MIT)
Port of https://github.com/achimala/dream-loop (MIT, 400+ stars in 48h).
An autonomous loop for building visually impressive 3D scenes/games:
generate photorealistic concept art, build (three.js/WebGL/Blender),
screenshot the live build, judge screenshot-vs-concept on a 5-tier score
ladder, iterate to convergence with explicit exit criteria.

Prose-only port rebound to Hermes-native tools: image_generate for
concept art, vision_analyze for judging (side-by-side composite
workaround documented), browser_exec capture_screenshot for live builds,
delegate_task for parallel asset work. Upstream ladder, failure modes,
time-budget and exit rules preserved. optional-skills/ placement.
2026-09-12 21:42:37 -07:00
Teknium
77f0c83ec3 fix(sessions): honor CLAUDE_CONFIG_DIR and CODEX_HOME in foreign session discovery
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.

_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.

Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
2026-09-12 21:39:10 -07:00
Teknium
8c4b0b751e docs: explain session hygiene — why /new still matters for memory and cost
A community deep-dive found that running one never-ending gateway session for weeks means memory injection, session_search, and pre-reset distillation almost never fire, while token cost grows. Sessions doc stated conversations never expire but never explained why users should still create boundaries. Adds a Session hygiene subsection with the mechanism and a practical rule.
2026-09-12 21:36:54 -07:00
Teknium
d5774ad880 fix(tests): pay heavy view imports at collection, not the first test's budget
Three CI-load flakes from the same class — a fixed per-test timeout billed
for one-time module-transform/env-init cost:

- apps/desktop messaging/index.test.tsx: `await import('./index')` ran inside
  renderMessaging(), so the FIRST test paid the whole MessagingView transform.
  On loaded runners that alone blew the 15s testTimeout and cascade-failed all
  subsequent tests in the file (unmounted DOM). Red on main runs 34599517793,
  34600757569, 34601269252 (green file takes 15.7s on a green main run —
  already over the first test's budget when billed there). Import moved to
  module scope, where vitest bills it to collection.
- apps/desktop skills/index.test.tsx: same pattern, 9 call sites; the file ran
  18.6s on a green main run. Deduplicated to one module-scope import (the
  existing 60s describe-timeout stays for the legitimately slow tests).
- web SessionsPage.test.tsx: the web vitest project still ran on vitest's 5s
  default while its per-row routing test legitimately takes 3.6-4.6s on GREEN
  runs; run 34600757569 tipped it to 5079ms. Gave web/vitest.config.ts the
  same 15s testTimeout the desktop project already carries, with the same
  rationale comment.

Validation: both desktop files 5x consecutive green + green pinned to 1 CPU
core (worst-case contention); SessionsPage 3x green; full desktop ui project
(801 files / 7622 tests) green; tsc + eslint clean on touched files.
2026-09-12 21:34:17 -07:00
liuhao1024
0e13fa98ec fix(web): cap web_extract provider dispatch with a wall-clock timeout (salvage #57180)
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.

Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).

Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."

Fixes #57155

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-12 21:30:38 -07:00
teknium1
f41b2c615d feat(skills): system-atlas — explorable isometric architecture atlases (port of inkboard/system-atlas, MIT)
Ports the system-atlas agent skill: one data.mjs file renders both an
interactive isometric HTML map (progressive-disclosure chapters, moving
data packets, question tracking by Q-ID) and a generated text twin
(SYSTEM.md) for the repo.

Why: architecture discussions that outgrow a single diagram — the skill
encodes hard-earned rules (max 3 structures per chapter, shapes+labels,
docs-policy ask before committing). Upstream 410 stars since Aug 20,
VoltAgent-listed; complements architecture-diagram (static SVG) with an
explorable, stateful artifact.

- optional-skills/creative/system-atlas/: SKILL.md (89 lines),
  references/ (design language, process lessons), assets/ (build.mjs,
  template.html, data.example.mjs) vendored verbatim; LICENSE.txt (MIT,
  Harshyt Goel)
- Live smoke: node --check both .mjs OK; node build.mjs produced
  atlas.html (41.5 KB) + SYSTEM.md from the example data, zero deps
- docs: own catalog row + generated page + sidebar entry only
2026-09-12 21:28:58 -07:00
salch-cred
ad0398eed8 fix(kanban): preserve sticky block on tasks created with initial_status=blocked (#107398) 2026-09-12 21:27:21 -07:00
Teknium
cbd4492f1f fix(tools): steer/redirect releases a blocking process_manage wait (port of MoonshotAI/kimi-code#3697)
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).

wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.

Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.

Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
2026-09-12 21:25:00 -07:00
Teknium
50f17b1300 fix(google-workspace): support multi-tab Google Docs (port cloudflare/cloudflare-os#450)
Google Docs can hold a tree of tabs, each with its own body and its own
1-based character index space; the legacy top-level `body` only carries
the first tab. `docs get` was silently dropping every other tab's
content, and `docs append` computed its insert index from the first
tab's body and sent the write with no tabId — so on a tabbed doc the
append could land at a wrong offset in the wrong tab.

- Requests now pass includeTabsContent=true (both gws and SDK paths).
- `docs get` returns a `tabs` array (preorder flattening of the
  tabs/childTabs tree, nested tabs included); single-tab docs keep the
  `body` field so existing callers work, and `--tab <tabId>` reads one
  tab. Legacy no-tabs responses are unchanged.
- `docs append` targets exactly one tab: the insert location carries
  the tabId, the end index is computed inside that tab's own body, an
  unknown `--tab` errors instead of falling back to the first tab, and
  a multi-tab doc without `--tab` errors with the tab list rather than
  guessing. Tabs are never merged — index spaces are independent.

Adapted from cloudflare/cloudflare-os#450 (gatekeeper-google), which
fixed the same provider behavior: reads must traverse Document.tabs and
every write Location must carry the immutable tabId.

Two pre-existing bare read_text/write_text in the touched test file
gained encoding="utf-8" (windows-footgun sweep rule).
2026-09-12 21:20:42 -07:00
Teknium
11d761eed9 fix(stream): flush residual SSE buffer at EOF and widen resilience to the Gemini native adapter
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.

- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
  after the read loop; surface EOF-without-[DONE] (warn + error result
  when nothing was received); extract _consume_sse_line so line parsing
  and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
  flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
  proven red on origin/main.
2026-09-12 21:18:24 -07:00
0xbyt4
c2b9a517e0 fix(gateway): harden proxy mode SSE streaming — 3 resilience bugs
Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.

1. `[DONE]` marker didn't break the outer chunk loop
   ---------------------------------------------------
   The `break` on `[DONE]` only exited the inner line-parse `while`,
   leaving the outer `async for chunk in resp.content.iter_any():` to
   keep reading. If the upstream held the connection open after
   `[DONE]` (buggy proxy, crashed server, network hang), the client
   waited up to sock_read=1800 seconds (30 min) for the next chunk.

   Fix: set a `done` flag when `[DONE]` is seen and check it at the
   top of the outer loop.

2. No TCP connect timeout
   -----------------------
   `ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
   the default `None` (no timeout). An unreachable proxy host (DNS
   fail, firewall, remote down) would hang on TCP connect for the OS
   default (minutes) before surfacing an error to the user.

   Fix: add `sock_connect=30` so connect failures surface within 30s.

3. SSE JSON parse exception handling was too narrow
   -------------------------------------------------
   The inner parse caught only `json.JSONDecodeError`. A response like
   `{"choices": [null]}` parsed successfully, then
   `choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
   object has no attribute 'get'`. That bubbled up to the outer
   `except Exception`, aborting the entire stream — any further chunks
   were lost, and the user saw the accumulated partial response
   without knowing why.

   Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
   dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
   to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
   chunk now skips, the stream keeps parsing.

New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:

- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
  chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
  kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
  chunks and verifies both good chunks are captured, bad ones skipped
2026-09-12 21:18:24 -07:00
Teknium
d46f79233e fix: sort imports in markdown-text.tsx per perfectionist rule 2026-09-12 21:17:02 -07:00
Teknium
e1c05ffa32 feat(desktop): persist video playback speed across transcript videos
Port from block/buzz#7336: a playback rate picked in any transcript
video's native controls persists as a device-level preference, so a
viewer who watches at 2x doesn't re-select it for every clip. New
players (and other open windows, via the persistentAtom storage sync)
start at the saved rate; out-of-range or malformed stored values fall
back to 1x, and returning to 1x removes the stored key.

Adapted from Buzz's hand-rolled localStorage module + custom player to
our persistentAtom store and the single <video> render site in
markdown-text.tsx (MediaAttachment), wrapped as TranscriptVideo.
2026-09-12 21:17:02 -07:00
entropy-0x
f79cb77224 fix(tools): reject V4A Add File onto an existing path
A V4A `*** Add File:` operation is meant to create a new file. The apply
path called `write_file` unconditionally and `_validate_operations` had no
pre-check for ADD, so an Add targeting a path that already existed
overwrote the file with only the patch's `+` lines, returned success, and
emitted a `--- /dev/null` diff that hid what was lost. Models frequently
confuse Add with Update, so this destroyed existing file contents with no
error. The MOVE path already guards its destination against clobbering;
ADD now follows the same rule.

Makes a V4A `Add File` operation fail when its target already exists,
instead of silently overwriting the existing file. Validation now rejects
the operation before any write happens, so the two-phase
validate-then-apply contract ("no files were modified" on a validation
failure) holds for ADD as it already does for UPDATE/MOVE/DELETE. A
matching re-check in the apply phase closes the validate-to-apply race.

N/A

- [x] 🐛 Bug fix (non-breaking change that fixes an issue)

- `tools/patch_parser.py`: add an ADD branch in `_validate_operations`
  that errors when `read_file_raw` finds an existing file, and a
  defensive existence re-check in `_apply_add` before `write_file`,
  mirroring the existing MOVE destination guard.
- `tests/tools/test_patch_parser.py`: add a test asserting an Add onto an
  existing path fails validation and leaves the original bytes unwritten;
  add `read_file_raw` to three ADD-path LSP fakes so they match the real
  `file_ops` interface now exercised on ADD.

1. Build a V4A patch with `*** Add File: <path>` where `<path>` already
   exists on disk.
2. Apply it via `apply_v4a_operations`. Before this change the file is
   overwritten with the patch's `+` lines and the result is success;
   after, the result is a validation failure and the file is untouched.
3. Run `scripts/run_tests.sh tests/tools/test_patch_parser.py` —
   `TestApplyOperations::test_add_onto_existing_file_fails_and_preserves_contents`
   covers the regression.

- [x] I've read the [Contributing Guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md)
- [x] My commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) (`fix(scope):`, `feat(scope):`, etc.)
- [x] I searched for [existing PRs](https://github.com/NousResearch/hermes-agent/pulls) to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix/feature (no unrelated commits)
- [x] I've run `pytest tests/ -q` and all tests pass
- [x] I've added tests for my changes (required for bug fixes, strongly encouraged for features)
- [x] I've tested on my platform: macOS 15 (Darwin 25.5.0)

- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) per the [compatibility guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md#cross-platform-compatibility) — or N/A
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
2026-09-12 21:13:40 -07:00
Teknium
24692ee790 feat(approval): flag cloud metadata-endpoint (IMDS) credential fetches for approval
On a cloud VM the instance-metadata service hands live IAM/service-account
credentials to any local process with no auth, so a fetch against it is
credential exfiltration unless the operator expects it — yet
detect_dangerous_command() auto-approved `curl` against the link-local
metadata IP, metadata.google.internal, and the Alibaba endpoint. Add one
DANGEROUS_PATTERNS entry covering 169.254.169.254 (AWS/Azure/GCP/OpenStack),
its AWS IPv6 form fd00:ec2::254, metadata.google.internal, and Alibaba's
100.100.100.200. The host literals have no other use, so their appearance in
a command is the signal regardless of HTTP client; lookarounds keep other
169.254.x.x link-local addresses and longer host/dotted strings out.

This prompts for approval (legit uses exist on real cloud VMs); it is NOT a
hardline block. Deterministic containment-escape detection at the approval
layer, same class as the existing credential-path detectors.
2026-09-12 21:12:18 -07:00
Teknium
6dd091a89c test: fold Gemini thinking-budget tests into invariant form 2026-09-12 21:10:02 -07:00
Shakti Prasad Mohapatra
29056b2335 fix(title): disable Gemini thinking tokens to prevent max_tokens starvation
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.

Two-part fix:

1. title_generator.py: Pass reasoning_config={"enabled": False} to
   call_llm() so thinking is explicitly disabled for title generation.

2. chat_completions.py: In _build_gemini_thinking_config, when
   reasoning is disabled (enabled=False or effort="none"), set
   thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
   includeThoughts: False only hides thought parts from the response
   while the model still reasons internally and bills thought tokens
   against maxOutputTokens. thinkingBudget: 0 truly disables thinking
   so thought tokens do not consume the max_tokens budget.

Fixes #91927
2026-09-12 21:10:02 -07:00
liuhao1024
5ea655771b fix(agent): disable reasoning on the title-generation pass 2026-09-12 21:10:02 -07:00
Teknium
1fec70ea48 fix(guardrails): catch repeating multi-call cycles in the stall guard
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).

Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
2026-09-12 21:07:46 -07:00
Teknium
a565e2d493 fix(mcp): widen the SSE fallback trigger and harden its guards (salvage #53764)
Relocated onto the decomposed module layout and hardened:

- Trigger covers the rejection CLASS, not just literal 400: SSE-only
  servers' load balancers answer the chunked Streamable HTTP initialize
  POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
  rejections as an opaque -32603 'Server returned an error response'
  (error class per #104363 by @RohithPariki). Timeouts and 5xx never
  trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
  clears _ready before re-entering the transport, so the original guard
  also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
  to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
  and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
  boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
  connects + latches; no fallback on reconnect/timeout/5xx; both-fail
  error is actionable.

The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.

Fixes #53676
2026-09-12 21:05:29 -07:00
aieng-abdullah
b2465f1608 fix(mcp): auto-fallback to SSE transport when Streamable HTTP returns 400
SSE-only MCP servers (e.g. WigAI for Bitwig Studio) reject the
Streamable HTTP initialize request with 400 Bad Request, causing
permanent failure with 0 active tools. The only workaround was
manually setting transport: sse in config.

When Streamable HTTP returns 400 during initial connect, log a
warning and retry with SSE before reporting failure. Reconnects
are excluded so a genuine 400 on an established transport is not
silently masked.

Extracted inline SSE code into _run_sse() helper shared by the
explicit config path and the new fallback path.

Fixes #53676
2026-09-12 21:05:29 -07:00
Teknium
e151d0b345 fix(discord): reject a numeric application ID pasted as the bot token during setup
Port from openclaw/openclaw#140531: users paste the application ID from the
Developer Portal's General Information page instead of the bot token (Bot
page); the gateway then fails at runtime with an opaque 401. A real bot token
is dot-separated base64 and never purely numeric, so the setup wizard now
rejects an all-digit answer with pointed guidance and re-prompts once. A
second consecutive numeric answer is kept (user override), and non-numeric
tokens are saved exactly as before.

Adapted to hermes: the guard lives in the Discord plugin's interactive_setup
(the active setup path per the #9983 review — legacy _setup_standard_platform
is not the dispatch target), mirroring the existing Telegram token-shape
validation in hermes_cli/setup_platforms.py.
2026-09-12 21:03:12 -07:00
Teknium
0f3199bd65 fix(dashboard): OAuth start routes resolve pollers late so test mocks intercept the spawned thread
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).

Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
2026-09-12 21:00:55 -07:00
Teknium
b35c0284a1 feat(skills): add mono-color — one/two-ink editorial print images (port of yanliudesign/mono-color-skill, MIT)
Port of https://github.com/yanliudesign/mono-color-skill (MIT, 2.9k
stars in 3 weeks). Generates original one-ink or controlled two-ink
editorial print images (risograph/duotone/monochrome poster aesthetic)
driven by machine-readable design-system catalogs: substrate/ink
palettes, composition geometry, typography roles, rhythm, controlled
imperfections — all vendored verbatim as the source of truth.

Upstream 34 KB SKILL.md restructured into a hub (143-line core + three
references). Image generation rebound to image_generate; hardcoded
'~/Desktop/Claude skills' path replaced with a neutral output dir.
Upstream examples/ are all-rights-reserved (ASSET-LICENSE.md) and are
NOT vendored — MIT text + catalogs only, with a NOTICE. optional-skills/
placement.
2026-09-12 20:58:38 -07:00
Teknium
1f855ec2a0 chore: map contributor email for #80648 salvage 2026-09-12 20:57:21 -07:00
Teknium
101b986fdf docs(cron): document resnap alongside the drift guard
The drift-guard section only offered pin-or-disable; resnap is the middle
path (adopt the new default, stay unpinned). Same PR as the salvaged
feature per docs-in-same-PR policy.
2026-09-12 20:57:21 -07:00
Søren L. Hansen
0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
Teknium
a49a9d79b3 Port from lobehub/lobehub#19329: surface environment recreation in terminal tool results
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.

Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.

Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
2026-09-12 20:55:04 -07:00
Teknium
6f62aff7f2 fix(tools): read_file line accounting — count unterminated final lines, drop the phantom trailing line
Two long-standing, mutually-masking defects in read_file's line accounting,
present on all three read paths (compound shell probe, sequential probes,
native):

1. total_lines came from `wc -l`, which counts newline bytes: a file whose
   final line has no trailing newline was undercounted by one. For an
   N*limit+1-line file read in pages, the last page was never offered
   (truncated=False) and the past-EOF guard refused `offset=total` reads of
   the real final line.

2. _add_line_numbers() split on '\n' without dropping the single terminating
   newline, so every well-formed newline-terminated file rendered a phantom
   `N+1|` empty gutter line that does not exist in the file (models routinely
   tried to patch/reference it). Exactly one terminator is dropped, so a
   genuinely selected trailing blank line keeps its number (`cat -n`
   semantics).

Both fixes land at the shared choke point (_assemble_read_result /
_add_line_numbers) so the compound, sequential, and native paths agree.
Existing tests that froze the buggy rendering as expected output are updated
to the corrected contract.

Salvages community PRs #106888 (@nikkoxgonzales) and #49453
(@MaxFreedomPollard); same bug class independently reported/fixed in
 #3907/#3908, #3927, #20814, #22945, #42929, #55696, #91306.
Cross-validated against anomalyco/opencode#47420's read-page serialization
fix (their trailing-blank-line class; hermes' page path is already
blank-line-safe once the terminator handling is right — verified live).
2026-09-12 20:52:49 -07:00
Max Freedom Pollard
10c34dd7e2 fix(tools): stop read_file rendering a phantom empty line for newline-terminated files
_add_line_numbers split on '\n', so a file ending in a newline (the normal,
well-formed case) produced a trailing empty element that got its own line
number. read_file therefore showed a phantom '<N+1>|' line that is not in the
file, on every terminal backend and every OS, matching neither cat -n nor the
reported total_lines. Drop the single terminating newline before splitting.

Fixes #49451
2026-09-12 20:52:49 -07:00
nikkoxgonzales
71063b1dbe fix(tools): count final unterminated line in read_file total_lines
wc -l counts newlines, not lines, so a file without a trailing newline
reported one fewer total_lines than the content it returned. The read
paths already probe the last byte (file_ends_with_newline) to strip cut's
phantom newline; use that same signal at the shared assembler choke point
so total_lines, truncation, and the past-EOF guard agree on every path
(compound, sequential, native).

Fixes #3907. Supersedes #3908: single adjustment instead of a per-path
helper, covering the native and sequential paths as well.
2026-09-12 20:52:49 -07:00
Teknium
eae250534f docs(memory): explain that memory needs session boundaries on gateways
A community write-up showed a real usage pattern: running Hermes on a messaging gateway for weeks without ever issuing /new. The memory docs never said that the recall loop (MEMORY.md/USER.md snapshot + session_search) only fires at session boundaries, and that gateway chats are intentionally one continuous session across restarts. Add a section spelling out why boundaries matter and recommending /new at natural break points, cross-linked to the session-continuity docs.
2026-09-12 20:50:34 -07:00
686f6c61
2bc9ed9f23 fix(mcp): fully redact credential headers in MCP probe errors and test display (salvage #97466)
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.

Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.

Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).

Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."

Fixes #97460

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-12 20:49:18 -07:00
teknium1
d7a56474fc feat(skills): pr-lens — animated architecture/data-flow diagrams for PRs (port of coldteadotai/pr-lens, 1.1k-star MIT)
Ports the pr-lens agent skill: represent a diff or subsystem as one
graph.json document and render it as animated SVG diagrams via the MIT
npx CLI (@coldtea/pr-lens-cli), with optional opt-in publishing to a
shareable canvas link.

Why: PR review and architecture explanation keep producing hand-drawn
Mermaid; this gives validated, animated, drill-down diagrams with a
deterministic document format. Upstream created Aug 20, 1.1k stars in
3 weeks, GitHub App + Action + CLI + skill.

- optional-skills/software-development/pr-lens/: SKILL.md (145 lines),
  references/ (config, graph document format, valid example) vendored
  near-verbatim, LICENSE.txt (MIT, Coldtea AI)
- gh --attach caveat handled: installed gh 2.97 lacks the flag; skill
  documents honest fallbacks (gist, canvas link, local path)
- Live smoke: validate + render of the vendored example graph passed
  (4 SVGs + manifest produced)
- docs: own catalog row + generated page + sidebar entry only
2026-09-12 20:47:00 -07:00
teknium1
d21913b09c fix(state): let sqlite raise the canonical error for a directory db_path
The owner-only pre-create helper ran before sqlite3.connect() and turned
a directory-as-state.db misconfiguration into IsADirectoryError instead
of the sqlite OperationalError the open path (and its lock-patience
classifier) expects. A directory leaks no row data, so skip it and let
sqlite fail canonically. Also map the salvage carry-commit author email
for the attribution gate.
2026-09-12 20:43:42 -07:00
joaomarcos
3966e5de94 security(state): harden async_delegation's direct state.db writer
tools/async_delegation.py:_connect() opens the same state.db as
SessionDB via a bare sqlite3.connect(), bypassing the owner-only
(0600) hardening added for SessionDB. Apply the same
_create_owner_only / _secure_wal_files policy here, reusing
hermes_state's helpers (managed/container skip included).

Addresses teknium1's review on #59716.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 20:43:42 -07:00
Hermes fleet-fix
1916cb249d security: make state databases and snapshots owner-only 2026-09-12 20:43:42 -07:00
Hermes fleet-fix
c445987559 fix(memory): secure built-in memory lock files 2026-09-12 20:43:42 -07:00
teknium1
baf1200e0a refactor(commandcode): drop the deepseek ImportError guard, trim tests to two invariants
The deepseek plugin is bundled and always loads, so the try/except around
the delegation import was defense-in-depth for a path that cannot fail.
Tests reduced to the two contracts that matter: /reasoning none reaches
the wire as thinking.disabled, and DeepSeek ids produce exactly the native
DeepSeek profile's output while non-DeepSeek families stay a no-op.
2026-09-12 20:41:32 -07:00
liuhao1024
fbb3a244b7 fix(commandcode): degrade to no-op when the deepseek plugin shim is missing
The bundled-plugin loader pops half-registered modules when a plugin
fails to load, so the lazy 'from plugins.model_providers.deepseek import
deepseek' could raise ImportError on every DeepSeek-routed CommandCode
turn — turning the soft 'thinking uncontrollable' bug into a hard
crash. Catch ImportError, log, and return the pre-fix no-op (review
feedback on #95241).
2026-09-12 20:41:32 -07:00