A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
The composer showed "<model> · Med" in one truncating pill and the only way
to change the effort was to open the model menu, find the active model's
row, and hover it for the per-row options submenu. Users read the pill as
"this model is medium only" and never found the submenu.
- New `ReasoningPill` next to the model pill: shows the active model's live
effort (session value, else the profile default) and opens the same
Thinking / Fast / Effort rows the catalog submenu offers, for the active
model only. Hidden when the catalog reports `reasoning: false`; stays
while capabilities are unknown so it never flickers during the fetch.
Folds away with the model pill in the compact composer stages.
- `useModelMenuController` (shell sibling) now owns the session write /
preset / optimistic-store / rollback logic that lived inside
`ModelMenuPanel`; the model menu and the new `ReasoningMenuPanel` share it
so an edit from either surface is one code path. Tiles get their own
pill bound to their SessionView, primary or tile — never the globals.
- `ModelOptionsContent` (the submenu body) is exported container-free so
the pill's top-level menu renders it without a Radix Sub wrapper.
- The model pill drops the effort suffix (`formatModelPillLabel`: name +
Fast); `formatModelStatusLabel` had no other caller and is removed.
- `currentModelCapabilities()` in lib/model-options resolves the active
pick's caps through `catalogProviderMatches` (aliases, custom slugs).
Live (headless Electron + worktree `hermes serve`, CDP): before — one
pill "Deepseek V4 Flash · Low", no effort control; after — "Deepseek V4
Flash" + "Low" pill; pick High → `config.get reasoning` on the live
session returns high; a `reasoning:false` cap unmounts the pill; the
catalog row submenu still writes through and the pill mirrors it.
Credit: the dedicated-pill direction was proposed independently in
composer selector on current main with the shared-controller shape.
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.
- mcp-research-data/: 224K of July tool-search bench result rows. The
harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
gitignored dir; the rows were committed by hand once and the headline
numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
WebResearchEnv that no longer exists; the yaml paths point at a
configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
for a resolved bug, two RFCs whose work shipped, an unimplemented
profile-builder proposal, the kanban dialog mock HTML and the kanban v1
spec PDF. profile-routing.md duplicated the profile_routes section of
website/docs/user-guide/multi-profile-gateways.md.
Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
Widens the new hook to the sibling interrupt surface: the TUI/desktop
session.interrupt path stops a live turn exactly like the gateway's
/stop, so plugins holding per-turn external resources get the same
signal there (platform='tui'). Gated on a genuinely running turn;
dispatch failures are swallowed so a plugin can never break the
interrupt. Docs updated to describe both surfaces.
Inspired by ChatGPT Work / Codex CLI 0.150.0 'Interrupt' hooks
(hooks that run when an active top-level turn is interrupted).
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.
_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.
Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.
Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
The Matrix docs described six agent-exposed matrix_* tools and three MATRIX_TOOLS_ALLOW_* env gates that were never implemented — the tool names and gates appear nowhere in code. Docs now describe actual behavior: no Matrix-specific agent tools; reactions/redactions are internal to approval prompts and pickers; MATRIX_ALLOWED_ROOMS scopes responses. Fixes#100535.
Slack deprecates the Assistant messaging experience (assistant_view) in
February 2027: assistant.threads.setStatus/setTitle are replaced by
agents.sessions.setStatus/rename. slack-sdk 3.44.0 (Aug 27 2026) ships
the typed methods with drop-in-compatible signatures.
- adapter: capability probe on the AsyncWebClient CLASS (never instance —
mock auto-attributes lie), cached; status set/clear + thread title route
through agents.sessions.* when available, legacy otherwise
- pins: slack-sdk 3.43.0 -> 3.44.0 (pyproject messaging+slack extras,
lazy_deps, uv.lock)
- tests: autouse fixture pins the probe to legacy under the mocked SDK;
5 new tests cover both routing paths for typing, clear, and title
- docs: slack.md scope table + status-line notes mention both methods
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."
A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.
Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.
- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
Review follow-up. Two of the three items taken as written; the third declined
with a reason.
1. Taken. The reconcile semantics mean `patterns` may only ADD -- an entry left
out of it is not removed, because the on-disk list wins for anything this
process did not approve itself. Every caller in the tree is additive today,
so nothing breaks, but the signature does not say so. Stated in the
docstring, and pinned by
`test_a_caller_that_passes_a_smaller_set_does_not_remove` so a future
`allowlist remove` finds out here instead of in production.
NOT taken: the `reconcile: bool = True` opt-out. There is no caller that
wants it, and AGENTS.md:98-101 names exactly this -- "Speculative
infrastructure. Hooks, callbacks, or extension points with no concrete
consumer." The removal path is editing config.yaml, which the docstring now
says.
2. Taken. website/docs/user-guide/security.md, next to the existing
`hermes config edit` tip, which is where an operator reads about removing a
pattern: the list is read at startup, a pattern removed while a session is
running stays approved in that session until the next write or a restart,
and if it was removed for safety reasons, restart.
3. Taken. `test_save_failure_is_logged_not_raised` asserted non-raising but
never asserted the log its name promises. Now asserts
"Could not save allowlist" via caplog.
scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
=== Summary: 1 files, 9 tests passed, 0 failed (100% complete) in 0.4s
Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.
Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.
Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
Port of https://github.com/achimala/dream-loop (MIT, 400+ stars in 48h).
An autonomous loop for building visually impressive 3D scenes/games:
generate photorealistic concept art, build (three.js/WebGL/Blender),
screenshot the live build, judge screenshot-vs-concept on a 5-tier score
ladder, iterate to convergence with explicit exit criteria.
Prose-only port rebound to Hermes-native tools: image_generate for
concept art, vision_analyze for judging (side-by-side composite
workaround documented), browser_exec capture_screenshot for live builds,
delegate_task for parallel asset work. Upstream ladder, failure modes,
time-budget and exit rules preserved. optional-skills/ placement.
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.
_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.
Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
A community deep-dive found that running one never-ending gateway session for weeks means memory injection, session_search, and pre-reset distillation almost never fire, while token cost grows. Sessions doc stated conversations never expire but never explained why users should still create boundaries. Adds a Session hygiene subsection with the mechanism and a practical rule.
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.
Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).
Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."
Fixes#57155
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Ports the system-atlas agent skill: one data.mjs file renders both an
interactive isometric HTML map (progressive-disclosure chapters, moving
data packets, question tracking by Q-ID) and a generated text twin
(SYSTEM.md) for the repo.
Why: architecture discussions that outgrow a single diagram — the skill
encodes hard-earned rules (max 3 structures per chapter, shapes+labels,
docs-policy ask before committing). Upstream 410 stars since Aug 20,
VoltAgent-listed; complements architecture-diagram (static SVG) with an
explorable, stateful artifact.
- optional-skills/creative/system-atlas/: SKILL.md (89 lines),
references/ (design language, process lessons), assets/ (build.mjs,
template.html, data.example.mjs) vendored verbatim; LICENSE.txt (MIT,
Harshyt Goel)
- Live smoke: node --check both .mjs OK; node build.mjs produced
atlas.html (41.5 KB) + SYSTEM.md from the example data, zero deps
- docs: own catalog row + generated page + sidebar entry only
Port of https://github.com/yanliudesign/mono-color-skill (MIT, 2.9k
stars in 3 weeks). Generates original one-ink or controlled two-ink
editorial print images (risograph/duotone/monochrome poster aesthetic)
driven by machine-readable design-system catalogs: substrate/ink
palettes, composition geometry, typography roles, rhythm, controlled
imperfections — all vendored verbatim as the source of truth.
Upstream 34 KB SKILL.md restructured into a hub (143-line core + three
references). Image generation rebound to image_generate; hardcoded
'~/Desktop/Claude skills' path replaced with a neutral output dir.
Upstream examples/ are all-rights-reserved (ASSET-LICENSE.md) and are
NOT vendored — MIT text + catalogs only, with a NOTICE. optional-skills/
placement.
The drift-guard section only offered pin-or-disable; resnap is the middle
path (adopt the new default, stay unpinned). Same PR as the salvaged
feature per docs-in-same-PR policy.
A community write-up showed a real usage pattern: running Hermes on a messaging gateway for weeks without ever issuing /new. The memory docs never said that the recall loop (MEMORY.md/USER.md snapshot + session_search) only fires at session boundaries, and that gateway chats are intentionally one continuous session across restarts. Add a section spelling out why boundaries matter and recommending /new at natural break points, cross-linked to the session-continuity docs.
Ports the pr-lens agent skill: represent a diff or subsystem as one
graph.json document and render it as animated SVG diagrams via the MIT
npx CLI (@coldtea/pr-lens-cli), with optional opt-in publishing to a
shareable canvas link.
Why: PR review and architecture explanation keep producing hand-drawn
Mermaid; this gives validated, animated, drill-down diagrams with a
deterministic document format. Upstream created Aug 20, 1.1k stars in
3 weeks, GitHub App + Action + CLI + skill.
- optional-skills/software-development/pr-lens/: SKILL.md (145 lines),
references/ (config, graph document format, valid example) vendored
near-verbatim, LICENSE.txt (MIT, Coldtea AI)
- gh --attach caveat handled: installed gh 2.97 lacks the flag; skill
documents honest fallbacks (gist, canvas link, local path)
- Live smoke: validate + render of the vendored example graph passed
(4 SVGs + manifest produced)
- docs: own catalog row + generated page + sidebar entry only
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.
`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).
Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.
`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.
The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).
Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.
Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
Follow-up to the cherry-picked #102431 fix, addressing the review findings:
- The two real-helper scheduler tests ran the Linux-only helper unmarked
and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
scripts/check-windows-footguns.py --all (lint lane red). The remedy text
is now built once in the helper and passed into the warning, so the
"scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
once (helper level); default config still Popens externally and
`require_restart_safe_scope: true` raises (scheduler level, real helper).
Dropped the stubbed duplicate, the standalone config-raise test and the
in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
with the same `except Exception` guard as the sibling
`failure_nudge_threshold` read, so a config error no longer escapes
the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
`cron.require_restart_safe_scope` documented in the cron user guide.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
The profiles guide promised "fresh sessions and memory" while
_CLONE_SUBDIR_FILES deliberately copies memories/MEMORY.md and USER.md
(curated identity, same tier as SOUL.md). State the real behaviour and how
to get a blank memory, so #10376's first half stops surprising people.
Refs #10376
Drop the categorical 'missing any step always means 200340' and
'200672/673 necessarily is an adapter bug' claims; the Interactive Card
toggle's role is unverified. Bring the Chinese guide in line (it still
pointed at the Event tab). Review follow-up.
Six issues (#10251, #10073, #13924, #38305, #25886, #8246) and four PRs
proposed json.dumps()-ing CallBackCard.data. The lark_oapi SDK types
`data` as Any and marshals the whole response with JSON.marshal(), so a
dict already reaches Feishu as the nested object the card-callback doc
shows in its response example; a string would nest a JSON *string* there
instead. Feishu documents 200340 as "the application has not configured
the card callback address or the configured request address is invalid"
— the click is rejected before delivery (nothing reaches Hermes), which
matches the reports (no gateway log, all four buttons identical).
The docs told users to subscribe to `card.action.trigger` under Event
Subscriptions; it lives under the separate Callback Configuration tab and
needs its own long-connection/URL mode and a published app version.
Spell out the four steps and map the sibling codes (200342/200343 =
unreachable URL; 200672/200673 = our payload).
Refs #10251
Every start (and every `ensure_hermes_home` from a sibling CLI invocation)
chmod'd the data directory to 0700, wiping group/other bits and the ACL
mask on a bind mount shared with other containers (hermes-webui, Nix
desktop + dashboard). _secure_file already skipped containers for this
reason; _secure_dir did not.
In a container the directory mode is now left to the operator unless
HERMES_HOME_MODE is set explicitly, which is still applied. cron/jobs.py
had its own 0700/0600 copies that bypassed the managed/container rules;
they now delegate to the shared helpers so cron/output stops re-locking
the mount as well.
Fixes#10757
`platforms.webhook.port: 9100` (and `routes`, `secret`, api_server `key`/
`cors_origins`, any adapter setting) was silently dropped unless nested
under `extra:` — PlatformConfig.from_dict only read a fixed set of typed
fields. Two partial bridges (a per-platform port/host/secret table in the
loader and an api_server-only block) covered a few keys and had to be
extended for every new one.
from_dict now promotes every non-typed top-level key into `extra`, with an
explicit `extra:` value winning on a clash and typed fields never leaking
into `extra` on a to_dict/from_dict roundtrip. Both hand-written bridges
are removed.
Same direction as PRs #10208/#10211/#10453 (rainow's #10206 diagnosis) and
#20506; those targeted the pre-loader layout.
Fixes#10206
The salvaged paragraph described one mechanism that is actually two.
Per developer.chrome.com/blog/remote-debugging-port, Chrome 136+ simply
stops respecting --remote-debugging-port/--remote-debugging-pipe against
the default data dir — no dialog, nothing listens. The "Allow remote
debugging?" prompt the contributor saw is the separate Chrome 144+ opt-in
approval flow (chrome://inspect/#remote-debugging), which asks per
incoming connection, not per launch. Say so, cite both sources, and keep
the fix (non-default --user-data-dir, or use_real_profile) unchanged.
The `/browser connect` tip said Chrome 136+ "silently refuse[s] to open the
remote debugging port" on the default user-data-dir and that "there is no error
message". Chrome does surface it: on Chrome 152/macOS, launching the default
profile with --remote-debugging-port puts up an "Allow remote debugging?" dialog,
and the port opens once you press Allow. Port 9222 refusing connections is the
symptom while that consent is pending, not a permanent silent block.
Two details this cost time to rediscover, now written down:
- the consent is asked per launch, so it reappears on every browser restart
- the remote-debugging toggle in chrome://settings does not suppress it; that
toggle only makes the feature available
Also cross-references browser.use_real_profile for the case the tip leaves
unanswered — wanting existing logins in the agent's browser. A dedicated
user-data-dir avoids the dialog but starts signed out; the real-profile snapshot
gets both, since the snapshot copy is itself a non-default user-data-dir.
Verified locally: default profile + --remote-debugging-port=9222 shows the
dialog and leaves 9222 closed, while a --user-data-dir launch answers
/json/version in under a second with no prompt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
~480 scientific research skills become searchable/installable through the Skills Hub
with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and
synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0).
A new optional tap-level `bucket` key stamps extra["category"] on every skill from a
tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub
category; a sidecar grouping still wins when present. Both repos stay at community
trust (not in TRUSTED_REPOS) so the guard scans every install.
Re-grafted from #60559 onto the post-split tools/skills_hub_github.py.
Snyk lands as a standalone Agent Plugins v1 package
(NousResearch/hermes-plugin-snyk, pinned 2a41a07f) instead of an
optional-mcps entry: the package carries the pinned `npx -y snyk@1.1306.0
mcp` stdio launch with CLI analytics disabled AND the workflow skill that
tells the agent when to use which scanner, so a single `hermes plugins
install snyk` gives both the tools and the playbook. Third-party product
integrations ship outside the core tree per the contribution rubric.
Supersedes the optional-mcps manifest from #73860 (same pin, same
telemetry posture, tool-pruning rationale moved into the skill).
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
The twozero TouchDesigner integration now lives in one installable unit:
plugin-catalog/touchdesigner.yaml points at NousResearch/hermes-plugin-touchdesigner
(portable Agent Plugins v1: mcp.json registers the twozero Streamable HTTP hub, skills/
carries touchdesigner-mcp). `hermes plugins install touchdesigner` + `hermes plugins
enable td` replaces the optional skill whose setup.sh hand-wrote an mcp_servers block.
The manifest name is `td` because Hermes names portable MCP tools
mcp__agent_plugin_<name>_<hash>__<server>__<tool>; twozero's longest tool under a
`touchdesigner` namespace is 71 chars, past the 64-char provider function-name cap.
optional-skills/creative/touchdesigner-mcp and its generated docs (bundled, optional,
zh-Hans) are removed; catalog tables, sidebar and kanban-video-orchestrator references
are updated to point at the plugin. Supersedes #68607 (MCP-catalog-only approach).
Replace the 'a secondary must not enable a port-binding platform' rule with the
shared-listener contract and a per-platform URL table (Twilio, LINE, Teams,
BlueBubbles, Microsoft Graph, WhatsApp Cloud, WeCom callback, Feishu webhook), plus
the status/dashboard surfaces that print the URL.
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.