Commit Graph

7413 Commits

Author SHA1 Message Date
Sam Foreman
2a5373da2c feat(cli): add display.vim_mode config and register /vim command
Introduces the opt-in surface for vi editing in the input composer:

- display.vim_mode config default (False, so existing users are
  unaffected and prompt_toolkit keeps its standard emacs bindings)
- /vim command registered with on|off|status subcommands, matching
  the established /battery and /timestamps pattern

Part of #4254.
2026-09-12 22:00:02 -07:00
Teknium
7b037f0efa feat(cli): opt-in git_branch status-bar field (⎇ current branch)
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.

Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.

Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
2026-09-12 21:47:53 -07:00
Teknium
053f8b1b17 fix(kanban): make archive-time worker termination race-safe and audited
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:

- Snapshot status/pid/claim INSIDE the archive txn so the kill only
  happens when this caller wins the archive transition; a losing
  concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
  skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
  hold the SQLite write lock. Safe because archived is terminal — no
  dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
  event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
  -> no signal, no event), live E2E: worker survived archive on main,
  terminated (<0.3s, clean SIGTERM) with the fix.

Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
2026-09-12 21:45:15 -07:00
PINKIIILQWQ
907f409dc4 fix(kanban): archive_task now kills the worker process
archive_task() was a pure DB operation — it cleared worker_pid,
claim_lock, and status from the tasks row but never sent SIGTERM
to the actual OS process. A running worker stayed alive until it
next called kanban_complete/kanban_block and discovered it was
archived, burning API quota and compute resources.

Fix: snapshot pid+claim_lock before the write_txn clears them,
then call _terminate_reclaimed_worker() — same function reclaim_task
uses — which sends SIGTERM, waits 5s, then SIGKILL if still alive.
The termination metadata is included in the 'archived' event so
operators can see what happened.

Order matches reclaim_task: terminate first, then DB update.
For non-running / non-local tasks, _terminate_reclaimed_worker
returns immediately as a no-op.

Closes #33774 reprise: the scratch-workspace side was fixed in
fc8afd500, but the orphaned-process side was never addressed.
2026-09-12 21:45:15 -07:00
Teknium
77f0c83ec3 fix(sessions): honor CLAUDE_CONFIG_DIR and CODEX_HOME in foreign session discovery
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.

_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.

Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
2026-09-12 21:39:10 -07:00
salch-cred
ad0398eed8 fix(kanban): preserve sticky block on tasks created with initial_status=blocked (#107398) 2026-09-12 21:27:21 -07:00
Teknium
0f3199bd65 fix(dashboard): OAuth start routes resolve pollers late so test mocks intercept the spawned thread
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).

Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
2026-09-12 21:00:55 -07:00
Søren L. Hansen
0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
686f6c61
2bc9ed9f23 fix(mcp): fully redact credential headers in MCP probe errors and test display (salvage #97466)
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.

Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.

Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).

Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."

Fixes #97460

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-12 20:49:18 -07:00
Hermes fleet-fix
1916cb249d security: make state databases and snapshots owner-only 2026-09-12 20:43:42 -07:00
teknium1
205645ee42 fix(gateway): register gateway.multiplex_profiles; explicit migrate --multiplex flips it with no standalone secondary
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.

`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
2026-09-12 18:35:21 -07:00
teknium1
acbecf588a fix(profiles): --clone leaves messaging channels behind; --clone-channels opts in
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).

Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.

`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.

The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
2026-09-12 18:35:21 -07:00
teknium1
de2d6a1b93 fix(config): a fresh process recovers the last good config.yaml instead of running on defaults
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).

Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.

Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
2026-09-12 16:17:04 -07:00
teknium1
d595e636c8 fix(model): a selected model id is never rewritten to a catalog neighbour
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:

- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
  any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
  model_switch applied it. Version bumps, dated snapshots and qualifiers
  all sit inside 0.9 of a sibling, so a newer release the listing lacked
  was swapped for the older one under the user's label. The matcher now
  does exact membership -> suggestion text only; the id goes to the wire
  verbatim and a genuine typo is refused with the listed siblings named.
  Every branch that carried the correction (live listing, static catalog,
  curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
  loses it in one place.

- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
  bare key (the slug Desktop picker rows carry) validated as a built-in
  and hit the hard-rejecting live-listing branch; the same endpoint as
  `custom:<key>` soft-accepted. Both spellings now validate as the user's
  custom endpoint.

- apps/desktop: `manualPickRemoved` (composer reseed) and
  `reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
  sticky pick to the profile default / the row's first model whenever the
  provider row did not list it. Rows are hints (discovered, curated,
  capped); the gateway's switch result is the only authority on a pick.
  Both helpers are removed; the pick stays put.

Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
2026-09-12 14:05:36 -07:00
teknium1
6a66a5d481 fix(desktop,dashboard): served profile's api_server/webhook read connected with their /p/<profile>/ URL; shared-gateway restart asks first
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.

- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
  webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
  (`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
  adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
  the real one, not a config guess. `hermes status` lists those URLs beside the other
  shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
  when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
  (statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
  shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
  Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
  start/stop on a served profile renders as an inline notice instead of a raw error toast.
2026-09-12 12:52:19 -07:00
Xipong
3d7f773bb4 fix(kanban): honor explicit platform tool opt-ins across configuration surfaces 2026-09-12 12:32:55 -07:00
kshitijk4poor
f003e449be refactor(cron): trim the scope-degrade dispatch to its invariants
Follow-up to the cherry-picked #102431 fix, addressing the review findings:

- The two real-helper scheduler tests ran the Linux-only helper unmarked
  and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
  scripts/check-windows-footguns.py --all (lint lane red). The remedy text
  is now built once in the helper and passed into the warning, so the
  "scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
  once (helper level); default config still Popens externally and
  `require_restart_safe_scope: true` raises (scheduler level, real helper).
  Dropped the stubbed duplicate, the standalone config-raise test and the
  in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
  with the same `except Exception` guard as the sibling
  `failure_nudge_threshold` read, so a config error no longer escapes
  the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
  instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
  `cron.require_restart_safe_scope` documented in the cron user guide.
2026-09-12 23:17:12 +05:30
Paul Robertson
560b6d2e81 fix(cron): degrade gracefully when systemd user scopes are unavailable
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
2026-09-12 23:17:12 +05:30
teknium1
2d121aa322 fix(gateway): hot-serve reaches pooled Desktop backends; deleted profiles leave no stale runtime entries
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
  ?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
  process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
  recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
  them as `stopped`.
2026-09-12 08:49:16 -07:00
teknium1
d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
teknium1
da451afb46 fix(update): probe configured-feature deps in the target venv, not the updater
The check called the registry check_fn inside the updater's own process,
whose import caches predate the install just performed (and which may be
the outer Python entirely), so a healthy freshly installed SDK produced a
false "will fail to load" warning. Run the same registry check in the
target interpreter via the existing _venv_probe path used by the core
dependency verifier.

Found by independent review before merge.
2026-09-12 08:47:21 -07:00
teknium1
2ef9c55a92 fix(update): name configured platforms whose extras failed to install
When `.[all]` fails and the per-extra fallback also fails for e.g.
`feishu`, the update printed only "Skipped optional extras that still
failed" and finished green. The running gateway kept its already-imported
modules, so the loss surfaced hours later as "No adapter available for
feishu" on the next restart (#10651).

After the fallback, check every enabled+configured platform through its
registry `check_fn` (and MCP when `mcp_servers` is set) and print which
configured feature will fail to load, with its install hint. Unconfigured
extras stay a quiet skipped line.

Reworks PR #10733 (LeonSGP43) against the registry instead of a
hand-written platform->module table so plugin platforms are covered.

Fixes #10651
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
2026-09-12 08:47:21 -07:00
teknium1
199544e054 fix(openrouter): canonical config URL keeps the pool; mirror key follows the selected endpoint
Two regressions in the mirror support: (1) the canonical
https://openrouter.ai/api/v1 that `hermes setup` persists under
provider: openrouter was treated as a custom endpoint, dropping the
auth.json credential pool and returning an empty API key; (2) an
unrelated CUSTOM_BASE_URL (which outranks the config mirror) still
received OPENROUTER_API_KEY because key selection tested mirror
eligibility, not the endpoint actually selected. A config URL is a
mirror only when its host is not openrouter.ai, and the mirror key
branch fires only when base_url is the config URL.

Found by independent review before merge.
2026-09-12 08:47:04 -07:00
JackJin
63f1016bea fix(cli): honor config base_url mirror for explicit openrouter provider
When config.yaml sets `model.provider: openrouter` together with a
`model.base_url` mirror/proxy, an explicit `--provider openrouter`
request ignored the mirror and sent traffic to the public OpenRouter
endpoint: the config base_url was only trusted for auto/custom, the
credential pool was still consulted (so a pooled key won over the
mirror), and even when the mirror URL was used its host failed the
openrouter.ai match so OPENROUTER_API_KEY was not selected for it.

Trust the config base_url for the explicit openrouter case, treat that
mirror as an OpenRouter context for key selection, and bypass the pool
like the other custom-endpoint cases already do.

Fixes #10622
2026-09-12 08:47:04 -07:00
teknium1
b78abb4710 fix(config): stop forcing 0700 on HERMES_HOME inside containers
Every start (and every `ensure_hermes_home` from a sibling CLI invocation)
chmod'd the data directory to 0700, wiping group/other bits and the ACL
mask on a bind mount shared with other containers (hermes-webui, Nix
desktop + dashboard). _secure_file already skipped containers for this
reason; _secure_dir did not.

In a container the directory mode is now left to the operator unless
HERMES_HOME_MODE is set explicitly, which is still applied. cron/jobs.py
had its own 0700/0600 copies that bypassed the managed/container rules;
they now delegate to the shared helpers so cron/output stops re-locking
the mount as well.

Fixes #10757
2026-09-12 08:43:43 -07:00
teknium1
7aa8249876 fix(cli): clear the preparing-line dedupe set on every stream reset
A tool batch that is cancelled or errors before any tool.started event
left the tool name in _tool_gen_announced, muting the next invocation's
'preparing <tool>…' line. Reset it in _reset_stream_state alongside the
other per-invocation state. Found by independent review before merge.
2026-09-12 08:43:12 -07:00
teknium1
cc269866aa fix(cli): announce "preparing <tool>…" once per tool per batch
The tool-gen callback fires once per tool CALL, so a model issuing three
parallel terminal calls printed three identical status lines before the
spinner took over (#10478). Coalesce repeats of the same tool name until a
tool actually starts (tool.started), which marks the next generation batch.

The line itself stays: it is the only feedback while a large argument
payload (e.g. a 45 KB write_file) streams, before the spinner exists.
PR #10598 removed it entirely; this keeps the first line and drops only
the duplicates.

Live A/B (in-process HTTP mock streaming 3 parallel terminal calls into a
real HermesCLI with display.streaming on): before 3 lines, after 1.

Fixes #10478
2026-09-12 08:43:12 -07:00
周鹤0668001310
92a8398087 fix(config): treat explicit false values in HERMES_MANAGED as unmanaged
Previously, setting HERMES_MANAGED=false (or 0, no, off) would be
interpreted as a literal managed-system name, causing is_managed()
to incorrectly return True and block update/config commands.

- Add _MANAGED_FALSE_VALUES tuple for canonical false strings
- Check false values before true values in get_managed_system()
- Add parametrized regression tests for all false variants

Fixes #12864
2026-09-12 08:41:00 -07:00
teknium1
1de9963ce7 fix(cli): backslash continuation runs before the paste-timing guard
Review finding: with multiline shortcuts on, a `\` + Enter arriving inside
a paste kept the literal backslash while a typed one removed it. The
continuation branch inserts a newline anyway, so it runs first and the
outcome no longer depends on timing.
2026-09-12 08:35:16 -07:00
teknium1
40f123ceab fix(cli): Enter arriving mid-paste is a newline, not a per-line steer
Without bracketed paste (tmux strips it; IME/voice dictation emits keys
one at a time) each pasted newline reaches _tui_handle_enter as its own
Enter event, so a multi-line paste while the agent ran became N separate
steering messages. An Enter within 50 ms of the last buffer change is
text still arriving; insert the newline and let the buffer accumulate.
Humans type >50 ms apart, so a real submit is unaffected.

Same mechanism as #10997 (@iRonin) and #19498 (@Ralojeil1), which
targeted the removed classic-CLI handle_enter.

Refs #10994
2026-09-12 08:35:16 -07:00
teknium1
847369e6c6 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:34:43 -07:00
teknium1
b70e0f4603 fix(mcp): CLI readers no longer reopen include: [] as "all tools enabled"
The runtime already registers nothing for an explicit empty include list
(fb1ec36a4b), but every CLI reader still coerced `[]` to "no filter":
`hermes mcp list` printed "all", `hermes mcp configure` and `hermes tools`
pre-checked every tool (so confirming the picker silently re-enabled all of
them), and a catalog reinstall pre-checked the manifest defaults over the
user's zero-tool choice.

`_tool_filters` now returns the list whenever the key holds a list; only an
absent/non-list key is None. The pickers and list output branch on `is not
None`, matching `tools/mcp_tool_registration.py`.

Fixes #12865. Builds on #13096 (@dingn42) and #52874 (@Bartok9).
2026-09-12 08:34:43 -07:00
vominh1919
d0093a11f0 fix(sessions): /branch and API fork end the parent only after the child exists
Both paths ended the source session as "branched" before create_session
ran, so a failed create left the user on a session already marked ended
with no branch behind it. Create the child first; the parent is ended
only once the branch is real.

Salvage of #11048 (targeted the pre-split cli.py handler; ported to
hermes_cli/cli_commands_mixin.py and the api_server fork sibling);
authored by @vominh1919.

Refs #11030
2026-09-12 08:33:08 -07:00
teknium1
9d0d886c43 fix(identity): spawn-ledger write survives surrogate-escaped argv
`_append_entry` moved to `utils.atomic_json_write`, which dumps with
`ensure_ascii=False` through a utf-8 text handle. An argv token holding
surrogate-escaped bytes (a non-UTF-8 project path via os.fsdecode) makes
json.dump raise UnicodeEncodeError — a ValueError, so the `except OSError`
does not catch it and callers silently lose their registration.

Expose `ensure_ascii` on `atomic_json_write` (default unchanged) and pass
True at the ledger call site, restoring the previous json.dumps behaviour
while keeping mode=0o600. One round-trip test, red on base.

Follow-up to #109156.
2026-09-12 08:27:53 -07:00
teknium1
12fe7684e1 fix(plugins): async-await helper thread runs under the caller's ContextVars
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
2026-09-12 08:26:48 -07:00
teknium1
24444e52eb fix(plugins): await async hook callbacks instead of collecting bare coroutines
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.

Fixes #12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.

Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
2026-09-12 08:26:48 -07:00
teknium1
71c72e078a fix(cli): one unverified-accept message for custom endpoints without /models; tests trimmed
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
2026-09-12 08:26:23 -07:00
Victor Nogueira
a3671787b0 fix(cli): clarify unverified custom model warning 2026-09-12 08:26:23 -07:00
Victor Nogueira
eda1e6cb15 fix(cli): accept unverified custom models
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
2026-09-12 08:26:23 -07:00
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1
535fd88c70 fix(backup): keep durable cache/ artifacts (images, citation ledger) in full backups
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
2026-09-12 08:25:49 -07:00
mrwanstudio
947e027f61 fix(backup): skip non-regular filesystem entries 2026-09-12 08:25:49 -07:00
Brad Estes
3b3f354933 fix(backup): exclude profile caches from full backups 2026-09-12 08:25:49 -07:00
teknium1
7817af2a16 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:25:36 -07:00
teknium1
2b685f08fa fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:24:52 -07:00
Drexuxux
45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1
d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao
c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1
44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1
5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00