Commit Graph

9451 Commits

Author SHA1 Message Date
kshitijk4poor
41303b379e refactor(sessions): export guard always counts every row; export_all reuses _active_clause
assert_export_safe grew an include_inactive flag whose only production caller
(the console export guard) always passes True, leaving the live-only branch
and its default unused. Drop the parameter and count every row, which is what
the transfer export materializes; the console call and docstring follow. The
existing guard tests seed live rows only and are unchanged.

export_all re-derived the live-row clause by hand; use the existing
_active_clause(include_inactive, False) so "live row" stays defined in one
place. Behaviour is unchanged.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-26 22:27:13 +05:30
kshitijk4poor
7d3c7caf99 fix(sessions): web streaming Export carries compaction-archived turns
The console export guard tells users whose session exceeds the in-memory cap
to use the Sessions page's streaming Export instead, but that endpoint read
get_messages() with the default live-only projection. Its JSON therefore
dropped every compaction-archived turn, and re-importing it collapsed a
9-turn compacted session to its 3 live rows - the exact loss this stack fixes
for the console path.

Page with include_inactive=True (keyset after_id is only incompatible with
include_compacted). Each row now carries its active/compacted flags, which
import_sessions re-archives. Probe: 9-turn compacted session -> web export
(11 rows) -> web import -> display 9, live 3, message_count 3 (was 3/3).
2026-09-26 22:27:13 +05:30
joaomarcos
85deb4553a fix(sessions): transfer export round-trips compaction-archived turns
The console `sessions export` backup went through export_session /
export_all, whose batched read filters `AND active = 1`, so an
export-then-import restore silently dropped every compaction-archived
turn (display 9 -> 3 in the C13 probe). export_all now takes
include_inactive (storage-order rows with their active/compacted flags,
which import_sessions restores as archived), the console export passes
it for both single-session and all-session exports, and
assert_export_safe sizes the same projection it guards.

Ported from #123267's transfer/backup hunks; its import-normalisation
hunks are superseded by #122680's checkpoint-correct import.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:27:13 +05:30
kshitijk4poor
d24aadfdd1 refactor(copilot): one GitHub effort clamp for the profile and main agent
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.

The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
2026-09-26 22:21:47 +05:30
kshitijk4poor
e991bfe24a fix(copilot): offline Astra effort fallback matches the exact Astra slugs
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
2026-09-26 22:21:47 +05:30
Jash Lee
7a597324a1 fix(copilot): preserve Astra reasoning effort
(cherry picked from commit 108f20c553852246deb0ce86fd5474d5d7992dc1)
2026-09-26 22:21:47 +05:30
kshitijk4poor
48ac94872a fix(models): canonical-URL check covers OpenRouter and equivalent spellings
Follow-up to the #123592 salvage. The canonical endpoint now comes from the
provider profile, which also covers OpenRouter (absent from PROVIDER_REGISTRY),
so pinning https://openrouter.ai/api/v1 keeps the native catalog too. Both
sides go through normalize_route_base_url, the helper the rest of the route
comparisons use, so scheme/host case and default ports no longer read as a
relay. The provider is normalized once.
2026-09-26 22:15:19 +05:30
tarkilhk
6e74191b87 fix(models): preserve native discovery for canonical provider URLs
(cherry picked from commit dbbe9a98e238d7c01177d0d3ac205deaa94e4675)
2026-09-26 22:15:19 +05:30
kshitijk4poor
36cbf4c076 fix(backup): stop blaming a live holder for every snapshot restore failure
_safe_restore_db now also returns False when the snapshot copy fails its
SQLite integrity check, but restore_quick_snapshot still logged every False
as "live-safe restore refused", pointing users at running processes when
the real cause was a corrupt snapshot. Make the log line and comment
neutral and defer to the preceding detailed log line from
_safe_restore_db; no second integrity check is added on this path.
2026-09-26 21:38:28 +05:30
kshitijk4poor
e9f3ea65ff fix(backup): tell the user when an imported database fails its integrity check
The integrity gate added to _safe_restore_db makes it return False for a
corrupt archive member, but _import_db_member still reported every False as
a live-holder refusal ("Stop the gateway/dashboard processes ..."), sending
users after the wrong fix while the real cause only reached logger.error.
On failure, re-run the bounded integrity check on the extracted temp file
and raise a message naming the actual cause; the holder message is kept for
genuine refusals. The check runs only on the failure path, so successful
imports pay nothing extra. Also document the new False case in the
_safe_restore_db docstring.

Co-authored-by: liuzikaii <2319582736@qq.com>
2026-09-26 21:38:28 +05:30
kshitijk4poor
48f1bf7251 refactor(backup): route read-only restore opens through read_only_db_uri
The three percent-encoded mode=ro connects added earlier in this stack
re-implemented hermes_state_holders.read_only_db_uri, which already exists
for exactly this '#'/'?' truncation bug and is the canonical builder used by
hermes_state and doctor_state. Using it keeps a single definition of the
encoding so a future fix lands everywhere. Behaviour is unchanged
(timeout=1.0 kept at the backup.py site).
2026-09-26 21:38:28 +05:30
salch-cred
793c07ffaf fix(backup): percent-encode path in _query_ro_sqlite URI
verify_sqlite_integrity() (now the source gate in _safe_restore_db) opens
the file through _query_ro_sqlite with a raw f"file:{path}?mode=ro" URI.
Under a path containing '#', SQLite treats the rest as a fragment: the
?mode=ro is dropped, the pre-'#' prefix is opened read-write (and created
empty), and the integrity check passes a corrupt source. Use
Path.as_uri(), matching the other restore-flow sites.

Salvaged from PR #123374.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-26 21:38:28 +05:30
liuzikaii
a609bad294 fix(backup): validate SQLite sources before restoring live databases
(cherry picked from commit e2767f7b00af3aadc1f87625d431653b402eb4d6)
2026-09-26 21:38:28 +05:30
liuzikaii
37a0857a33 fix(backup): escape SQLite restore source paths
(cherry picked from commit ff093bf8671cf5a857f4283f978ac685cb5a7c84)
2026-09-26 21:38:28 +05:30
kshitijk4poor
8afaab3703 fix(gateway): restart wait survives a non-finite drain; tighten its tests
Follow-up to the salvaged restart-wait commits:

- A drain or cron timeout of .inf now means "wait indefinitely" instead of
  an OverflowError from the integer stop envelope, which crashed
  `hermes gateway restart` and made `hermes update` silently fall back to
  its 45s floor. The fleet "draining (up to Ns)" lines and the drain
  progress report format the budget instead of int()-ing it, so an
  unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
  the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
  replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
  envelope and that configured cron reaches the CLI wait, instead of
  re-deriving the formula; the negative wording assertion on the pending
  footer is dropped (change-detector).
2026-09-26 20:07:47 +05:30
Brian Le
5aad192995 fix(gateway): include cron drain in restart exit wait
(cherry picked from commit aefbb0741ed3d8491ede6e2d26689022c69a85bc)
2026-09-26 20:07:47 +05:30
emozilla
9c1ef17550 fix(local-runtime): move a pre-PM llama.cpp engine into the PM store and pin b10964
The package-manager switch (#102765) made engine detection read only PM's
store and pinned every llama.cpp backend to b10362. A machine with an
engine under runtimes/llamacpp/b<tag>/<backend>/ reported no engine, so
the local models pane showed one-click setup with models already on disk.

installed_engine() now moves the newest pre-PM install into the store
once. The old manifest's archive digests become the PM identity, the
package verifier runs llama-server --version, and os.rename puts the
directory in place before facts.json records it. A store that already
holds an engine is left alone, a store lock held by another PM operation
defers the move, and an install that fails verification stays where it
is for the rest of the process.

The lock returns to b10964 on all five backends, the default before the
PM switch. With b10362 pinned, the pane offered a downgrade from the
moved engine as an update. Upstream renamed the ROCm archives to 10.0 at
b10767, and the hip assets follow. Only the CUDA build is measured at
b10964.
2026-09-25 22:07:48 -04:00
ethernet
3582c17cbc fix(local-runtime): move models from the old per-profile models dir into the machine dir
Models used to live at <profile home>/models (adb2fdbf5c) before the managed
runtime moved them to the machine-scoped <root>/models (43e67d872f). GGUFs
staged under a named profile's old dir silently stopped being served.

adopt_legacy_models() renames them (and their assets/) into the current dirs,
so everything downstream keeps reading one directory: listing, presets,
delete, the --models-dir fallback. It runs at boot, before the "anything
staged?" check, and in the Local Models status route so the pane lists them
while the runtime is off.

os.rename only: instant within a filesystem, while shutil.move would silently
copy tens of GB across devices at session start; a cross-device dir stays put
with a warning. An existing destination name is never replaced.
2026-09-25 22:06:58 -04:00
kshitijk4poor
98054af9a7 fix(models): keep Sonnet 5 as the Bedrock default after adding Opus 5.5
get_default_model_for_provider('bedrock') returns the static list's first
entry, so putting Opus 5.5 at [0] silently moved unconfigured Bedrock
users, gateway/API-server runs with no model, and '/model bedrock' from
Sonnet 5 to the priciest flagship. Opus 5.5 now sits second.
2026-09-26 07:14:26 +05:30
Deyan Dimitrov
87a9b4820c feat(models): list Claude Opus 5.5 in the Bedrock static fallback
The static Bedrock list is what the picker shows when live discovery
(ListFoundationModels + ListInferenceProfiles) is unavailable; it had no
Opus 5.5. Its 1M window comes from the anthropic.claude-opus-5 key in
BEDROCK_CONTEXT_LENGTHS (longest-substring match).

Salvaged from #119444 (catalog hunk only).

(cherry picked from commit 22b7e61f6b)
2026-09-26 07:14:26 +05:30
Deyan Dimitrov
42ba854db5 feat(models): add Claude Opus 5.5 to the native Anthropic picker
The native Anthropic curated list had no Opus 5.5 even though the
OpenRouter table in the same file already lists anthropic/claude-opus-5.5,
so the native picker showed it only as a live-discovered extra at the
bottom. Uses the hyphenated id that /v1/models returns.

Salvaged from #119444 (catalog hunk only; the thinking/OAuth/tool_choice
halves are covered by #121016 and #106414).

(cherry picked from commit 757ebf625ae0fee45aae8df29b9d140c16c27a8c)
2026-09-26 07:14:26 +05:30
David Metcalfe
c60ba75199 fix(desktop): pass skin customCSS through to desktop theme context
The web dashboard supports customCSS in theme YAML (PR #14776) but
the desktop app's skin pipeline dropped the field. Users who wanted
custom styling had to hack app.asar, which gets overwritten on every
update.

This adds customCSS passthrough through the full skin pipeline:

- HermesSkin (apps/shared) and DesktopTheme types gain customCSS?: string
- skinToDesktopTheme() passes skin.customCSS through
- applyTheme() injects a scoped <style id="hermes-desktop-custom-css">
  tag on theme apply, and removes it when switching to a CSS-less theme
- SkinConfig gets custom_css field, read from YAML as "customCSS"
- _build_skin_config() caps at 32 KiB (same as the web dashboard)
- resolve_skin() emits "customCSS" in the gateway JSON-RPC payload

Users now put CSS in ~/.hermes/skins/<name>.yaml under the customCSS
key — it persists across updates because the skin dir is outside app.asar.

Closes #53013
Refs #53012
2026-09-25 20:00:57 -05:00
kshitijk4poor
4c1fd5bc05 fix(update): stop rendering 'managed outside dashboard' as a shell command
Containerized dashboards refused updates with update_command set to the
prose 'managed outside dashboard'. Desktop renders update_command verbatim
as a copyable '$ <cmd>' line, so users got a fake command that fails with
'command not found: managed'.

- backend: managed-externally refusal and check return update_command=''
  (no runnable command), matching the commit-build refusal.
- desktop: treat an empty update_command as 'no command' (message-only
  manual view); fall back to 'hermes update' only when an older backend
  omits the field. Previously '' || 'hermes update' also showed a wrong
  command for commit builds.
2026-09-26 06:23:43 +05:30
Hermes Agent
3ddf82fe24 fix(desktop): detach the Windows packaged launch and swallow launcher Ctrl-C
Two launcher defects in cmd_gui's packaged launch path:

- On Windows the packaged Desktop was spawned with a console-inheriting
  subprocess.run, so closing the launching shell sent CTRL_CLOSE_EVENT
  down the process group and killed the app with it, while Electron/Node
  stdout flooded (and, under cp936, mojibaked) the parent terminal.
  Spawn it detached instead — CREATE_NEW_PROCESS_GROUP|DETACHED_PROCESS
  (+ CREATE_BREAKAWAY_FROM_JOB with a no-breakaway fallback), DEVNULL
  stdio — and exit 0 immediately, like the bundled launcher and
  gateway_windows._spawn_detached. macOS/Linux keep the foreground run.

- On the foreground platforms, Ctrl-C in the attached terminal raised a
  raw KeyboardInterrupt traceback out of the CLI instead of a clean
  close. Catch it and exit 0 with a short message.

Co-authored-by: briandevans <252620095+briandevans@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-25 19:25:40 -05:00
Hermes Agent
67c6fdc092 fix(desktop): stop home-directory repo scans when no roots are configured
An empty repo_scan_roots silently expanded to a bounded scan of the
user's entire home directory on every Desktop launch, with no way to
restrict the traversal short of disabling discovery outright. Empty
roots are now a safe no-op: users must explicitly configure
desktop.repo_scan_roots for filesystem scanning, and session-derived
projects remain available. The default config comment documents the
opt-in.

Fixes #53328

Co-authored-by: John Kim Querobines <jkim.querobines@gmail.com>
2026-09-25 19:09:03 -05:00
Hermes Agent
f84db42a32 fix(cron): deduplicate cross-profile jobs in GET /api/cron/jobs
profile=all aggregated every profile's list without deduplication, so a job
copied into a second profile's cron/jobs.json during profile creation appeared
twice — inflating the desktop sidebar count and rendering duplicate rows.

Collect all jobs first, then resolve duplicates by id with default-profile
priority, instead of keeping whichever copy the profile loop happened to append
first. cron.jobs._normalize_job_record() fills a missing id with the literal
sentinel string "unknown", which is truthy — treat it like no id so two
genuinely different id-less legacy records from different profiles are never
collapsed into one.

Fixes #51721

Salvaged from #69132 by @ygd58 (kept its dedup semantics and regression tests,
ported onto the web_routers/cron.py seam after the web_server refactor).

Co-authored-by: ygd58 <buraysandro9@gmail.com>
2026-09-25 18:28:06 -05:00
Hermes Agent
a975615869 fix(uninstall): sweep macOS launchd dashboard jobs and caches in the full wipe
A full uninstall only removed the gateway launchd label, ~/.hermes, the
checkout, and the desktop's Application Support userData. Left behind on
macOS: every dashboard/serve LaunchAgent (launchd kept respawning its
backend), and the Electron/Chromium + setup cache dirs written outside
HERMES_HOME (~/Library/Caches/Hermes, com.nousresearch.hermes,
hermes-setup, com.nousresearch.hermes.setup) — all of which survived the
"uninstall complete" screen and degraded the next install (#62209).

Extend the full-wipe step: boot out and delete every launchd job whose
ProgramArguments runs a hermes dashboard / serve backend (mirroring the
enumeration hermes_cli.main_dashboard uses to find those jobs, including
the malformed-plist skip contract), and rmtree the four cache dirs. Both
steps are macOS-only and run only in the full wipe; keep-data mode is
unchanged. The dry-run plan now lists them.

Fixes #62209
2026-09-25 18:08:49 -05:00
Brooklyn Nicholson
87d1c5b0c3 refactor(tui): drop redundant native-mode guards 2026-09-25 17:57:38 -05:00
brooklyn!
93f7617406 feat(tui): add native terminal mode defaults 2026-09-25 17:57:38 -05:00
Hermes Agent
6d70abc871 fix(desktop): label the Anthropic OAuth/Pro card 'Anthropic Account'
The accounts page and onboarding picker labeled the anthropic entry
'Anthropic API Key' even though that flow (hermes auth add anthropic)
connects the Claude Pro/Max subscription via OAuth/PKCE. Users with a
subscription read the label literally, concluded they needed an API key,
and went key-hunting. The entry that actually wants a pasted key is
'claude-code' (claude setup-token), which keeps its own name.

Rename the shared PROVIDER_DISPLAY_NAMES entry and the dashboard
provider-catalog name to 'Anthropic Account', and update the onboarding
test assertions.

Fixes #59071
Salvages #59073 (same rename, authored by Kailigithub)

Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
2026-09-25 17:18:06 -05:00
Hermes Agent
43c1eee1fa fix(web): bound the /api/model/info context-length probe
get_model_info called get_model_context_length inline. The resolver
chain runs several sequential provider probes, each with its own
multi-second timeout, so an unreachable model.base_url held the
response for tens of seconds. Bound the whole chain to a 5s budget on
a throwaway worker thread; on timeout the response degrades to
auto_context_length=0 while the abandoned probe finishes on its own.

Fixes https://github.com/NousResearch/hermes-agent/issues/63214 (backend half)
2026-09-25 17:17:48 -05:00
Hermes Agent
954d94c3b0 fix(cli): name Windows Smart App Control when it blocks the dashboard runtime
When Windows Smart App Control or an Application Control policy blocks the
embedded Python runtime's _ssl module, the fastapi/uvicorn import in the
dashboard startup fails with the "DLL load failed ... _ssl" signature, but
the handler printed the generic missing-deps message — so users looped on
`hermes pm install` / `hermes pm repair`, which can never lift a policy
block.

Detect the signature and say so instead: the runtime's DLL is blocked,
repair cannot fix it, and the recovery options are a trusted system
Python, an IT exemption, or the CLI/gateway (which use the system Python).

Fixes #63796

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-25 16:28:14 -05:00
Austin Pickett
fee49791a3 fix(desktop): handle package-manager installs with no desktop source tree
A Homebrew (or other pip/package-manager) install lives in a
site-packages tree that does not ship apps/desktop, so `hermes desktop`
failed with a generic missing-source error implying a broken checkout.
Detect the install kind (Cellar/site-packages), prefer launching a
separately installed /Applications/Hermes.app, and otherwise print a
brew-specific instruction instead.

Fixes #61056
2026-09-25 16:57:57 -04:00
brooklyn!
331a230da6 fix(accounts): stop stale Claude Code connections and false removal success
Accounts Connected now follows token validity instead of access-token
presence. Windows removal uses unambiguous PowerShell, a clear that
removes nothing is an error instead of a success toast, and the
connected-row terminal control runs disconnect.
2026-09-25 14:25:49 -05:00
Hermes Agent
70f7fa05ca fix(providers): keep llamacpp as the model table's managed-runtime id
Mapping the llamacpp aliases to custom in hermes_cli.models sent the
managed local runtime's /model validation down the custom branch before
the staged-library check, so a downloaded-but-not-running GGUF lost its
recognized verdict and an unstaged one was accepted. Keep local and vllm
on custom there, leave the llamacpp aliases on their runtime id, and
move the orphaned 'local' display label to 'custom'.

The parity test now reads the aliases from the custom provider profile
instead of a hand-written list.
2026-09-25 14:25:16 -05:00
PRATHAMESH75
7a5c715844 fix(providers): unify local OpenAI-compatible provider aliases to custom (#62213)
Local OpenAI-compatible server aliases (local, vllm, llamacpp, llama.cpp,
llama-cpp) were normalized inconsistently across the three provider-name
tables: hermes_cli.providers mapped them to the orphan "local" id,
hermes_cli.models left them unmapped, while hermes_cli.auth mapped them to
"custom". Align all three on the generic "custom" provider so routing,
the model picker, and credential resolution agree, matching what
resolve_provider already did for vllm/llamacpp. Add a parity contract test.
2026-09-25 14:25:16 -05:00
Hermes Agent
0cf900ea17 fix(skills): keep built-in provenance for skills the catalog dropped
Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-25 14:24:56 -05:00
Austin Pickett
1b57acf94a fix(mcp): budget discovery-thread GIL so agent build fits the 30s wait
Background MCP discovery does bursty CPU work (SDK/pydantic imports, JSON-RPC
schema parsing, tool registration) that at the default 5ms switch interval rides
the GIL convoy effect and starves the concurrent agent-build thread and the main
event loop for tens of seconds after a serve restart — _wait_agent(30s) times
out with 'agent initialization timed out' (error 5032).

Lower the interpreter switch interval while discovery runs (restored afterwards)
so its CPU bursts are sliced finely enough that waiters stay responsive.

Fixes #60371
2026-09-25 15:01:16 -04:00
Hermes Agent
ec5c9c738a fix(processes): persist_on_release keeps background jobs alive across lifecycle kill sweeps (#41225)
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.

Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
  carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
  (_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
  explicit operator stops (process_manage kill, /stop slash + RPC mirror,
  CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
  persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
  the host is going away and survivors would become PPID=1 orphans

Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
2026-09-25 13:49:31 -05:00
Austin Pickett
b00b4bb7f2 fix(cli): pin utf-8 decoding on all text-mode subprocess readers
On Windows, subprocess text=True without an explicit encoding decodes
child output with the ANSI code page (e.g. 'gbk'); non-ASCII bytes then
raise UnicodeDecodeError inside subprocess._readerthread, killing the
Hermes backend before it becomes ready and surfacing as the desktop boot
timeout.

Sweep every hermes_cli text=True subprocess call to encoding='utf-8',
errors='replace', and add an AST-based regression test that fails when a
future text-mode call omits the encoding.

Fixes #55658
2026-09-25 14:22:12 -04:00
Austin Pickett
fae9e5677a fix(update): stop reading gateway identity off the Windows restart watcher's argv (#107002) (#121635)
* fix(update): stop reading gateway identity off the restart watcher's argv (#107002)

The detached restart watcher is spawned as
`python -c <watcher source> <old_pid> <python> -m hermes_cli.main gateway run`.
Its trailing argv is the command it will spawn LATER, but the canonical matchers
read identity straight off the joined command line, so the watcher itself was
classified as a live `gateway run` process — the documented "never infer process
identity from argv substrings" bug class, on the exact surface `hermes update`
uses to verify a post-update relaunch.

Also budget the post-relaunch liveness poll against the watcher's own deadline:
the watcher respawns the gateway only after the PID it was handed exits, so a
30 s window can expire before the relaunch it verifies was scheduled to start.

* test(windows-live): run the gateway-ancestor harness parent from a script file

A `python -c <src>` parent is an interpreter running inline source and carries
no readable Hermes identity, so it is no longer a gateway to any classifier —
the harness's own comment already said a realistic gateway argv is not a -c blob.

* test(windows-live): share one sleeper SCRIPT across the live process-topology fixtures

Four live Windows E2E files stood processes up as `python -c "sleep" <hermes argv tail>`.
That shape no longer carries a readable Hermes identity, so the fixtures stopped standing
in for the gateways they simulate. One shared sleeper script replaces the -c spelling.

* fix(gateway): drop the duplicated _INLINE_SOURCE_FLAG_RE definition

The constant was emitted twice around command_line_runs_inline_source. Same
pattern both times, so behaviour is unchanged — but one definition is enough.

* test(stderr-timestamp): run the gateway-lookalike children from script files

Both lookalikes stood a gateway child up as `python -c <src> <gateway tail>`.
That shape no longer carries a readable Hermes identity (#107002), so the
wrapper correctly stopped treating them as gateway spawns and the tests failed.
A `-c` tail is data for a program the inline source may spawn LATER, never the
child's own identity — the real wrapper child is `python -m hermes_cli.main
gateway run`, which has no `-c`. Running the stand-ins from a real script file
restores what the tests mean to assert without depending on the misread.

* test(windows-live): restore the tempfile import dropped with the local sleeper helper

* test(windows-live): wait on the sleeper SCRIPT name, not its source text

The live fixtures proved argv visibility by waiting for `time.sleep(120)` in the
spawned process's command line. That string only ever appeared there because the
sleeper was spelled `python -c "import time; time.sleep(120)"`; now that it runs
from a file the source is in the file, so the probe timed out ("sleeper argv
never visible") even though the argv was perfectly visible.

Wait on the script name instead, exported as SLEEPER_MARKER next to the script
so the probe and the spelling cannot drift apart again.

* fix(tests,gateway): keep the live-system guard blocking -c-wrapped gateway spawns

The #107002 identity fix made _gateway_command_subcommand return None for
'python -c <src> … -m hermes_cli.main gateway run'. tests/_fixtures/live_system_guard.py
shares that matcher, so the autouse guard stopped blocking the detached restart
watcher: real gateways leaked out of the e2e run and squatted the webhook port.

Add gateway.status.gateway_spawn_intent_subcommand — the spawn-intent mirror of the
identity matcher, peeling the inline-source wrapper token-wise and re-running the same
canonical matcher on each suffix (still no substring matching) — and point the guard at
it. Read-only subcommands stay spawnable.

* fix(gateway): make the inline-source option walk value-aware so -X utf8 -c is not read as a gateway

The walk that decides whether a command line is an interpreter running inline
source (`python -c <src> ...`) treated every token starting with `-` as a flag
and the first non-flag token as the end of the option block. CPython options
that take a SEPARATE operand (-X/-W/-Q, --check-hash-based-pycs, --jit) break
that model: the operand was mistaken for the end of the block, so the walk
never reached the -c behind it and the watcher was read as a live gateway
again -- exactly the #107002 misclassification, one shape further out.

- Reuse the canonical operand sets from hermes_state_holders rather than
  hand-rolling a second copy (AGENTS.md: parser-derived flag sets).
- Walk case-preserving tokens: operand-taking -Q/-W/-X must not be conflated
  with operand-less -q/-b, so callers no longer lowercase before the walk.
- Handle clustered short options precisely (-uc is inline source, -Xc is -X c).
- Replace the ad-hoc _INLINE_SOURCE_FLAG_RE rescan in
  gateway_spawn_intent_subcommand with the index the same walk returns; the
  regex could not find spellings the walk accepts and would raise
  StopIteration.

Reported by an automated review on PR #121635 and reproduced here.

Refs #107002

---------

Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
2026-09-25 14:15:22 -04:00
Austin Pickett
9df628615c Merge pull request #121614 from NousResearch/austin/fix/121347-base-url-resolution
fix(providers): honor a configured base_url instead of the vendor host across runtime, fallback, listing and validation
2026-09-25 14:15:01 -04:00
Austin Pickett
4a3b54ac7e fix(desktop): diagnose esbuild ignore-scripts build failures
When a desktop build fails because esbuild's platform binary
(@esbuild/<platform>) was never staged - typically because
ignore-scripts=true skips esbuild's postinstall - print an actionable
diagnosis with the fix instead of leaving only the raw build error.

Fixes #53082
2026-09-25 14:12:27 -04:00
Hermes Agent
94f3dbec9b fix(agent): close one-shot AIAgents on every exit path
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:

- batch_runner._process_single_prompt: one agent per prompt, N prompts
  per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
  gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.

Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.

Fixes #50197
2026-09-25 12:49:56 -05:00
Austin Pickett
4661892752 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution
# Conflicts:
#	hermes_cli/models.py
2026-09-25 13:20:26 -04:00
calvinnwq
979c7baeef fix(serve): retire an SSH-isolated backend once the install moves to new code
The serve a remote Desktop spawns over SSH is never restarted by the
host's updater (only its client holds the token and owner nonce), and
the idle watchdog stays quiet while that client is connected. After an
update it therefore kept running the pre-update code against the new
tree until the client happened to reconnect.

Poll the checkout sha against the sha this process loaded. On two
consecutive mismatches, with no update in flight, retire through the
existing retirement fence (which proves idle and closes admission), and
exit cleanly so the client respawns the backend on the new code. Unknown
shas, busy or unreadable ledgers and a live update all keep it up.
2026-09-25 12:09:37 -05:00
calvinnwq
074ead267d fix(update): classify a Desktop SSH serve as its remote client's, not manual-serve
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.

Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
2026-09-25 12:09:37 -05:00
calvinnwq
2e0243b158 fix(desktop): never attach to an isolated serve found in the spawn ledger
hermes serve --isolated (the backend another machine's Desktop spawns
over SSH) opts out of the host singleton on the CLI side, but its spawn
ledger row carried no structured marker, so the local Desktop's
attach-first discovery adopted it. Nothing on this host owns that
process, so a Desktop-driven update left the local app on stale code.

Record isolated=True in the ledger row and skip such rows in
parseSpawnLedger. An ordinary serve is still attached even when an
isolated record is newer.
2026-09-25 12:09:37 -05:00
brooklyn!
296ec08dcf fix(models): keep generation models out of chat
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
2026-09-25 12:07:32 -05:00
Hermes Agent
d0cb567273 fix(gateway): take host identity from the live host record on replay
The settled-flag fix covers a process that multiplexes itself. The update
and fleet processes replay a FOREIGN gateway's captured argv with no
settled flag of their own, so a selector-less argv fell back to the
ambient HERMES_HOME comparison — the exact coordinate the review rejects
(#93943): a host launched from a named profile was replayed as that
profile, donating the named credentials to the respawned host.

Both restart edges now consult, in order: this process's settled
multiplex verdict, then the live host gateway's published rendezvous
record (its SETTLED served set, proven live), and only then the
compatibility default-root comparison. The raw config re-read stays last
so no settled identity exists => unchanged compatibility behavior.

Regressions: selector-less replay from a named home with a live host
record is host; without one it stays profile-scoped; the restart watcher
env takes the default root and drops the named token when only the host
record proves hostness.
2026-09-25 12:01:44 -05:00