Commit Graph

4775 Commits

Author SHA1 Message Date
Nicolas Formenton
c75d835555 feat(desktop): mark a session as unread/read with a persisted watermark 2026-08-15 01:21:40 -07:00
Moisés Valero
fef9c537d7 fix(cli): convert Alt key shortcuts to sequence tuple for prompt_toolkit (#74169) 2026-08-15 01:05:39 -07:00
John Lussier
2ae7884ffa fix: make CLI multiline shortcuts work by default 2026-08-15 01:05:39 -07:00
Teknium
77fcc2ea31 feat(display): honor display.timestamps across desktop transcript and TUI
One config key everywhere (#41531): the same display.timestamps that stamps
[HH:MM] on classic-CLI labels now gates the desktop transcript's timeline
timestamps and renders dim [HH:MM] labels on TUI user/assistant rows.

- desktop: $displayTimestamps store fed from config.yaml via
  use-hermes-config; TimelineTimestamp renders nothing while the key is off
  (the default). Hover tooltips with the exact time stay ungated (#70450).
- TUI: tui_gateway forwards each persisted row's timestamp in the display
  projection; toTranscriptMessages threads it as Msg.createdAt; live rows
  are stamped at append (the #82840 rule); MessageLine shows a dim [HH:MM]
  above user/assistant rows when display.timestamps is on.
- No new config keys, no HERMES_* env vars; display-only, prompt-cache safe.
2026-08-15 01:04:19 -07:00
doncazper
6a375a24a5 fix(gateway): timestamp shared stderr output 2026-08-15 01:04:19 -07:00
doncazper
1db9273584 fix(gateway): timestamp launchd error log lines 2026-08-15 01:04:19 -07:00
Paul BlackSwan
d5a865882a fix(desktop): save remote gateway files natively 2026-08-15 00:37:00 -07:00
jackoconner55
2f54ad4023 perf(cli): skip launcher-side plugin discovery for TUI handoff
(cherry picked from commit e2e0edd6b8ec10d02e867ff11dd1f9b4961a1ba7)
2026-08-15 00:36:03 -07:00
spfcraze
2a5093eeff docs(auth): note get_default_hermes_root in the global-store memo docstring
(cherry picked from commit 64de1a13f967c5735d861b972fe6c65e89f5af17)
2026-08-15 00:36:03 -07:00
spfcraze
4bd746c6e9 perf(cli): memoise default-hermes-root resolution and global auth-store read
get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).

_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.

Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).

Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).

(cherry picked from commit be348f32e5bd7479c26fabb652de549fb9c8a1e1)
2026-08-15 00:36:03 -07:00
spfcraze
473490a2e9 perf(cli): stop per-keystroke config re-reads in slash completers
The /tools and /personality completers run on every keystroke while the
user types those commands (complete_while_typing), and both re-read +
re-parse the full config on each keypress:

- _tools_completions called load_config() — the defensive deepcopy
  (~340us/call on cache hit) even though it only reads toolset enable
  state + MCP server names. Switched to load_config_readonly() (the
  perf(agent) #74322 pattern; this per-keystroke site was missed).
- _personality_completions called load_cli_config() — a full YAML parse
  + deep merge of the built-in defaults (~110us) — on every keystroke.
  Memoised keyed on the config file path+mtime (same pattern as load_env
  / _nous_auth_status_cache), so the parse runs once per config state.

Measured: /tools 357us -> 18us per keystroke; /personality parse drops
from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse).

Regression tests: _tools_completions uses the readonly loader (deepcopy
loader never called); personality memo parses once across repeated
completions and re-parses once after a config mtime change.

(cherry picked from commit 2b3f897171f93dc6b099848ec5ab3763b7294870)
2026-08-15 00:36:03 -07:00
Adolanium
ee1731b13c perf(cli): re-export decomposed command modules lazily, ~60ms off every CLI start
The main.py decomposition re-exported the sessions/update/dashboard command
surface with eager from-imports, so every hermes invocation (including
hermes --version) paid for update_cmd's dependency chain (jwt, click,
cryptography). Resolve the re-exports through the existing PEP 562 module
__getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on
first actual use. Internal call sites go through a _self() helper because
bare-name lookups do not trigger __getattr__; _self() imports sys locally
since update tests patch hermes_cli.main.sys. The
_warn_stale_dashboard_processes back-compat alias moves into the lazy
surface, and the sessions argparse dispatch defers sessions_cmd to call
time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a
real module attribute, which shadows __getattr__.

Measured (Windows 11, Python 3.11, median of 7 warm runs):
import hermes_cli.main 253ms -> 196ms (-22%).

(cherry picked from commit cad1083b71635f98815b698d69a18f7f58e15517)
2026-08-15 00:36:03 -07:00
lepetitprince716-prog
7e439dbb1b perf: parallelize provider model-list fetches in model picker
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.

Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
  that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
  concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
  with threading.Lock to prevent concurrent write races

Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)

Closes #80413

(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
2026-08-15 00:34:29 -07:00
briandevans
4c24629bc9 fix(console): skip the checkpoints prune confirmation the console already took
Hermes Console registers `checkpoints prune`, `clear` and `clear-legacy` as
mutating, so it takes a console-level confirmation before dispatching any of
them. `_apply_confirmed_defaults` then exists to keep the CLI layer from
asking a second time — its docstring says so — but it only force-defaults
`clear` and `clear-legacy`. `prune` was left out, even though `cmd_prune`
gates its orphan preview on the identical `not args.force` shape.

`_capture_output` redirects stdout and stderr but never stdin, so the
unskipped `_confirm()` call hits `input()` with no terminal behind it:
`EOFError` propagates into `_confirm`, which returns False, and `cmd_prune`
prints "Aborted." and returns 1. The console turns that non-zero exit into a
ConsoleCommandError, so `checkpoints prune` fails outright for any user who
has at least one orphan checkpoint project — after that user already
confirmed. When the server does happen to inherit a foreground terminal, the
same call instead blocks a console worker thread and eats the operator's
keystrokes.

Forcing the flag is the documented behavior here rather than a weakening of
the recent orphan-allowlist hardening. `orphan_allowlist` binds a deletion to
the identities shown in the preview, guarding the window where a workdir
disappears while the command waits on `input()`. Under the console there is
no preview and no wait, which is exactly the `--force` case the comment on
`cmd_prune` describes as "no restriction".
2026-08-15 00:33:32 -07:00
YuYigeng
7d34d7d8d1 docs(tui): clarify destructive confirm controls 2026-08-15 00:33:32 -07:00
Teknium
f378a8fb3b fix(projects): dedup project_create by primary_path (#75820)
Creating a project whose resolved primary path already belongs to a
non-archived project now raises a clear ValueError naming the existing
project (create_project) — duplicated projects each seeded an identical
copy of the repo subtree, multiplying the duplicate-lane bug per copy.
The agent-facing project_create tool is idempotent instead: it re-activates
the existing project rather than erroring. allow_duplicate_path=True keeps
deliberate duplicates possible. Also updates the legacy non-git lane-id
expectation to the branch-style id introduced for #53329.
2026-08-15 00:33:11 -07:00
Teknium
94ce8396e8 fix(sessions): release active-session leases against their acquisition registry
A gateway active-session lease is acquired against the root HERMES_HOME,
but release_active_session()/transfer_active_session() re-resolved the
registry path from the *current* HERMES_HOME. Under native multiplex a
routed turn runs agent cleanup inside _profile_runtime_scope, so the
release looked under the named profile while the root entry stayed
alive — after max_concurrent_sessions routed turns every new session was
rejected with 'Hermes is at the active session limit' (#85431).

Pin state/lock paths on the lease at acquisition time and prefer them on
release and transfer. Fixes #85431.
2026-08-15 00:33:01 -07:00
rainbowgits
aba4934274 fix(agent): omit unsupported metadata on Relay scope.pop
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
2026-08-15 00:33:01 -07:00
Teknium
67a1c1ed1a fix(dashboard): use a fixed sidebar cache TTL (no HERMES_* env var for non-secret config) 2026-08-15 00:32:53 -07:00
Christopher
5bceb3e84b fix(dashboard): add idle back-off to PTY pump loop (#42627) 2026-08-15 00:32:53 -07:00
Lucas Oliveira
f0cfe5a56f perf(dashboard): bound multi-profile sidebar polling 2026-08-15 00:32:53 -07:00
thatssoheil
ce02f0ab8a fix(pets): cover the pets CLI (doctor, has-active, /pet toggle) too
Second review follow-up: hermes_cli/pets.py still read display.pet.enabled
with bare bool()/truthiness in _cmd_doctor (misreported quoted 'false' as
enabled in 'hermes pets doctor'), _has_active_pet (quoted 'false' treated
as active, so /pet install skipped the selection prompt), and
toggle_pet_display (/pet toggle flipped the WRONG way). All three now go
through is_truthy_value(default=False); the module imports the shared
helper at the top.
2026-08-15 00:32:44 -07:00
Teknium
fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
Teknium
ce996d4057 feat(delegation): raise max_concurrent_children default 3 -> 10 (+migration) (#86745)
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/​investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.

Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.

- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
  default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.

Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:20:32 -07:00
EvanProgramming
30c469b153 fix(gateway): spare pidfile-less Scheduled-Task gateways from the orphan reaper on Windows (#83683)
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).

Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.

The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.

Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.

Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
2026-08-15 12:04:29 +05:30
Teknium
471c687c2b test(managed_uv): cover explicit-patch fallback on the next minor line; dedupe retried versions
Follow-up to the salvaged #76252 addressing both review gaps:

- New TestMinorLineFallForward class with a direct test of the
  explicit-patch fallback branch: bare '3.12' resolves to a VULNERABLE
  build while an explicit 3.12.x patch is fixed, so recovery must go
  through _list_available_patches on the next minor line. Asserts the
  exact `uv python install` request sequence.
- New all-minors-exhausted test: everything vulnerable on 3.11-3.13
  returns None with per-line attempts bounded by _MAX_PATCH_RETRIES and
  no requests beyond 3.13 (requires-python is <3.14).
- test_retry_is_bounded_by_max_retries_constant now actually uses its
  counting wrapper and asserts the collected install calls (the
  previous version collected them into a dead variable).

Also dedupes the fallback loop the same way the same-minor loop does:
_attempt_install_generation can now record the probed candidate version
into a caller-supplied tried_versions set, so the explicit-patch pass
skips the version the bare-minor request already resolved to and
rejected -- previously that wasted a full download+install+probe+delete
cycle per minor line re-trying a known-vulnerable build.
2026-08-14 22:37:46 -07:00
RelaxJonh
2bccd6ad08 fix(managed_uv): fall forward to next Python minor when current line has no fixed SQLite build
When every patch on the current minor line (e.g. 3.11) still links a
vulnerable SQLite (e.g. 3.50.4 on Windows), the provisioner now tries
the next supported minor line (3.12, then 3.13) before giving up.

Previously, _install_safe_python_generation only tried patches within
the same minor line. On Windows, where python-build-standalone may not
publish a fixed build for the installed patch, users were stuck with a
repeated warning on every `hermes update` with no path forward.

The requires-python constraint (>=3.11,<3.14) and the downstream
import smoke test already gate compatibility, so the minor-line
upgrade is safe.

Adds allow_minor_upgrade parameter to _attempt_install_generation to
relax the same-minor-line version guard when called from the fallback
path.

Fixes #76106
2026-08-14 22:37:46 -07:00
Tachi
d1df111ccd fix(update): restore Hermes Tools dependencies 2026-08-14 22:33:44 -07:00
Tachi
979a20052f fix(update): preserve activated extras across runtime rebuilds 2026-08-14 22:33:44 -07:00
konsisumer
4aa9f738ce fix(update): rebuild Desktop after release artifact loss 2026-08-14 22:27:37 -07:00
joaomarcos
fa72a1edf8 fix(kanban): re-create the schema when a cached DB path loses it (#83445)
`connect()` caches every path it has initialized in the process-local
`_INITIALIZED_PATHS` set and then skips all first-open work for it —
header validation, integrity probe, `SCHEMA_SQL`, additive migrations.
That cache is keyed on a path, but the schema it stands for lives in a
file, and the two can drift apart: delete or replace `kanban.db` under a
live gateway/dispatcher/dashboard process and the next `connect()` takes
the fast path, lets SQLite create a fresh empty database, and hands back
a connection with no tables in it.

Nothing notices. Every query then fails with `no such table: tasks`,
`plugin_api._conn()` logs its init warning and carries on, and the board
renders empty. Because the cache entry survives, the process re-creates
the same schema-less ~4 KB file on every restart of the desktop app in
front of it — only killing the backing process clears it.

Verify the sentinel table on the fast path and self-heal when it is gone:
drop the stale cache entry and fall through to the existing init path,
which re-runs the probes and the schema script under the cross-process
init lock. The check is one `sqlite_master` lookup on the already-resident
page 1, so the steady-state path stays lock-free (#36644) and does no
schema work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:25:23 -07:00
Teknium
4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
RelaxJonh
198e2f2746 fix(routing): use hostname match for azure.com endpoint detection (#74312)
Replace raw substring checks ("azure.com" in full_url) with the existing
base_url_host_matches() helper at two sites in runtime_provider.py.

The substring approach misclassified URLs whose path (not hostname)
contained "azure.com" — e.g. https://example.invalid/proxy/azure.com/v1 —
causing the wrong credential (Azure key instead of explicit Anthropic token)
to be selected, and potentially leaking a more-privileged Azure key across
a trust boundary.

base_url_host_matches() parses the URL and validates only the hostname
against allowed Azure suffixes with proper boundary rules.

Fixes #74312
2026-08-14 22:04:16 -07:00
Teknium
42a1db4c64 fix(update): use canonical venv_bin_dir in _install_repair (no open-coded Scripts/bin) 2026-08-14 22:03:56 -07:00
Halldrix
19cff89300 fix(update): complete pending core install before any native import (self-lock loop fix)
Reviewer egilewski found the original defer was circular (#83590 comment):
the self-lock preflight wrote .update-incomplete and exited, but the next
launch only ran the full recovery AFTER main.py's third-party imports —
so a healthy venv's probes made the early pass a no-op, main.py imported
cryptography eagerly, the .pyd got mapped again, and the deferred install
re-hit the exact self-lock it was meant to escape.

Close the loop by making the marker guarantee the install runs BEFORE any
native extension can be imported:

- hermes_cli/_install_repair.py (new, stdlib-only): single source of truth
  for the core .[all] reinstall — ensurepip bootstrap, uv-pip/pip
  resolution with VIRTUAL_ENV, Termux env stripping, Windows hermes*.exe
  quarantine, per-extra fallback ladder, and fd1→fd2 routing for acp
  safety.  Deliberately free of managed_uv/hermes_constants imports so it
  stays importable in the corrupted-venv state it exists to repair.
- hermes_cli/_early_recovery.py: recover_if_needed now completes a pending
  .update-incomplete install BEFORE the import probes, on every launch
  that sees the marker (unless argv is update).  Success clears the
  marker; failure bumps an attempts counter inside the marker body and
  keeps it.  A 3-attempt ceiling stops a persistently-failing install
  from reinstall-hammering every launch (hermes acp included) — past the
  ceiling the late post-import recovery takes over with its manual
  recovery instructions.  Single-flight lock shared with the late path.
- hermes_cli/main.py: _recover_core_update_marker_locked delegates the
  install to the shared executor (no duplicated logic); ensure_uv stays
  in the late path so a venv whose uv vanished mid-update still
  bootstraps it.
- tests: 7 new regressions — the reviewer's exact case (marker + healthy
  venv → install runs while sys.modules has no cryptography), failure
  keeps marker + increments attempts, retry ceiling, lazy marker does
  not trigger core install (#58004 invariant), argv-update skip, and
  corrupt/missing marker bodies.  The key test was sabotage-verified:
  removing the pre-import branch makes it fail with zero install calls,
  while a lone-lazy-marker test still passes; restoring the branch makes
  it pass again.

Refs #83569
2026-08-14 22:03:56 -07:00
Halldrix
c6a71294b6 fix(update): detect updater self-lock on Windows + repair venvs whose base interpreter is uv-managed
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:

1. Self-lock detection. _detect_venv_python_processes() always excludes
   the calling process by design — a CLI hermes update IS the venv python.
   An updater that had already imported a native venv extension (the
   canonical one being cryptography.hazmat.bindings._rust, mapped while
   hermes_cli.main resolved external secret sources) passed every
   preflight and then died mid-sync with os error 5 when uv tried to
   rewrite the mapped .pyd, stranding the venv half-updated. A new
   preflight now refuses the sync before touching the checkout, writes
   the update-incomplete marker so the next fresh launch completes the
   install, and exits 2. Verified on a live Windows 11 host: after
   importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
   in the caller, and a peer process cannot open it read-write
   (Permission denied) — while a rename succeeds, matching how uv/pip
   actually fail (truncate+write, not rename).

2. Early-recovery install path. _early_recovery._run_repair_install used
   sys.executable -m pip unconditionally. Windows git checkouts install
   on a uv-managed base interpreter (python-build-standalone), whose
   EXTERNALLY-MANAGED marker makes plain pip abort with
   externally-managed-environment — the repair no-oped and the venv
   stayed broken. The repair now detects the PEP 668 marker, prefers
   uv pip install with VIRTUAL_ENV pointed at the project venv, and
   falls back to pip --break-system-packages when no uv binary exists.

Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.

Fixes #83569
2026-08-14 22:03:56 -07:00
chelsealong
49d72a02f6 fix(update): verify Windows gateway cold-start survives before reporting success
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).

Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
2026-08-14 22:03:56 -07:00
PRATHAMESH75
517151ee4a fix(install): fail closed when a stopped gateway leaves an empty survivor probe
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.

Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.
2026-08-14 22:03:56 -07:00
PRATHAMESH75
95018b6bba fix(install): surface aborted gateway restart during hermes update
The gateway auto-restart phase in `hermes update` was wrapped in a blanket
`except Exception` that only logged at debug level. When the phase raised
early — e.g. importing `hermes_cli.gateway` from the freshly pulled checkout
inside a process that already loaded pre-update modules — every drain and
restart line vanished from the update output, the update printed
"Update complete!" and exited 0, and the still-running gateway kept serving
pre-update modules against replaced source files. The next Telegram turn died
with `ImportError: cannot import name 'is_trivial_prompt'`.

The handler now probes for surviving gateway processes and, unless it can
positively prove none are running, prints the cause plus a manual recovery
command and marks the fleet restart incomplete — which exits nonzero and
writes the gateway-mode exit-code marker, matching the existing
failed-or-stale-unit path.

Fixes #78574
2026-08-14 22:03:56 -07:00
Soheil Fakour
bdfdd4392f fix(update): gate 'Code updated!' on HEAD actually moving (#79678)
A detached/pinned checkout can report 'N new commit(s)' against origin,
run the ff-only merge successfully, and still sit on the old commit
afterward (the branch-switch step re-detaches to the raw SHA). Before
this guard 'hermes update' printed '✓ Code updated!' and reinstalled
deps + rebuilt the desktop app against the stale tree - no error, no
warning, 'hermes doctor' healthy.

Compare pre-pull and post-pull HEAD; if they match, fail loudly with a
reattach hint instead of claiming success.
2026-08-14 22:03:56 -07:00
kshitij
b58fa89cd7 refactor: centralize dict-valued model.default coercion via shared helper
Promote _split_model_config_default to hermes_cli/config.py as the single
shared helper for flattening dict-valued model.default/model.model config.
All 8 defense-in-depth sites now route through it instead of inlining
their own isinstance checks with inconsistent key orders.

Changes:
- Add split_model_config_default() to hermes_cli/config.py (public)
- cli.py: _split_model_config_default delegates to shared helper
- Fix key extraction order: agent_runtime_helpers.py was reversed
  (default->model); now consistent (model->default) across all sites
- Remove provider-as-model-name fallback from main.py, oneshot.py,
  model_tools.py, cli.py — provider is a routing key, not a model ID
- Add 'name' to _normalize_root_model_keys flattening loop and
  _has_nested_default detection to cover the deprecated model.name alias

Tests: 118 passed + 1 skipped (cli_init, managed_scope, config).
E2E: 31/31 passed (config chokepoint, managed scope, crash site,
negative cases, edge cases).
2026-08-15 10:29:05 +05:30
mariobgsp
be708ff1b9 fix: normalize managed config overlay before merge in load_config
The shared load-boundary flatten added for dict-valued model.default only
ran on the user/default merge; _load_config_impl then deep-merged the raw
managed overlay without normalizing, so a managed model.default:
{provider, model} still reached status/fallback/runtime readers as a dict.

Normalize the managed overlay (same _normalize_root_model_keys pass, plus
the bare model-string -> model.default promotion used by
managed_scope.apply_managed_overlay) before expanding and merging, so
every overlay is canonical before load_config returns.

Adds load_config() regressions for a nested managed default and a bare
managed model string.
2026-08-15 10:29:05 +05:30
mariobgsp
998329a621 fix: flatten dict-valued model.default at config load boundary
Extends the fix to the config-load chokepoint so every reader sees plain
strings, not just the interactive CLI paths. _normalize_root_model_keys
now flattens a dict-valued model.default/model.model into a string default
plus the nested provider (promoted to model.provider when no explicit
outer provider or "auto" is set), covering the residual readers the
review flagged: doctor, status/dump, fallback picker, prompt-size, and
the context-switch guard — all of which called .strip()/flowed the raw
value and would crash or misroute on a nested dict.

Adds _normalize_root_model_keys regression coverage for the flatten
(precedence, auto-override, explicit-provider-win, alias shape, flat
strings untouched).
2026-08-15 10:29:05 +05:30
Ario Bagus Prakusa
cb4daf23f3 fix: coerce dict-valued model/default config back to string across resolution paths
A dict-valued model.default (e.g. {provider:..., model:...}) in config.yaml
was leaking into agent.model and crashing the agent at init:

  AttributeError: 'dict' object has no attribute 'lower'
    agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy

This manifested on the Telegram gateway as an infinite reset loop: every
turn built an agent with model=dict, crashed during init, the gateway
treated the failed turn as a session needing reset, and /reset rebuilt the
agent and crashed again.

Coerce dict -> string at every model-resolution entry point so the value
is normalized once and never reaches a .lower() call as a dict:
- agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy (the crash site)
- agent/agent_init.py: configured default model resolution
- cli.py: CLI config model + _normalize_model_for_provider
- hermes_cli/main.py: _has_any_provider_configured
- hermes_cli/oneshot.py: _run_agent model resolution
- hermes_cli/runtime_provider.py: _get_model_config default handling
- model_tools.py: _resolve_active_context_length
2026-08-15 10:29:05 +05:30
HexLab98
18b442cdeb fix(install): abort Windows venv recreate when rename-aside fails
When Rename-Item on the live venv is denied, do not fall back to an
in-place Remove-Item that can gut site-packages and leave no rollback.
Also mark venv-blocker probe failures with probe_failed so they cannot
be read as a clear scan (#83149).
2026-08-14 21:58:09 -07:00
tachyon-r
7a6b8917f7 fix(tools): recognize discovered plugin platforms 2026-08-14 21:56:33 -07:00
Chen Jin
7224301856 fix(toolsets): admit explicitly-configured plugin toolset keys in _get_platform_tools (#81163)
Layer 2 of the #81163 / #78050 fix: _get_platform_tools computed
plugin_ts_keys = _get_plugin_toolset_keys() but only used
CONFIGURABLE_TOOLSETS in the explicit-config filter, so a user-listed
plugin key like `a2a` in `platform_toolsets.cli: [hermes-cli, a2a]` was
silently dropped. The filter now unions configurable and plugin toolset
keys when evaluating has_explicit_config and when admitting per-key
entries.

Cherry-picked from PR #81190 (Layer 2 hunks only; Layer 1 is covered by
the provides_tools mechanism from PR #78842).
2026-08-14 21:56:33 -07:00
Eman
e42db348c9 fix(plugins): register deferred platform client tools at discovery (#78050)
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.

Re-anchored accordingly:

- Discovery-time pre-registration, module reuse, and the `provides_tools`
  opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
  since those tools registered before `registration_start` and the slice
  cannot see them.
- A failed materialization no longer carries attribution across. The
  failure path now sweeps the whole ownership ledger for the plugin key,
  not just the `registration_start:` slice, so the pre-registered tools
  are disposed along with the adapter. Attribution and the registry now
  agree at zero instead of reporting tools the process is not serving.

tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:56:33 -07:00
webtecnica
2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
Jack Lau
6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00