Commit Graph

9426 Commits

Author SHA1 Message Date
Hermes Agent
f84db42a32 fix(cron): deduplicate cross-profile jobs in GET /api/cron/jobs
profile=all aggregated every profile's list without deduplication, so a job
copied into a second profile's cron/jobs.json during profile creation appeared
twice — inflating the desktop sidebar count and rendering duplicate rows.

Collect all jobs first, then resolve duplicates by id with default-profile
priority, instead of keeping whichever copy the profile loop happened to append
first. cron.jobs._normalize_job_record() fills a missing id with the literal
sentinel string "unknown", which is truthy — treat it like no id so two
genuinely different id-less legacy records from different profiles are never
collapsed into one.

Fixes #51721

Salvaged from #69132 by @ygd58 (kept its dedup semantics and regression tests,
ported onto the web_routers/cron.py seam after the web_server refactor).

Co-authored-by: ygd58 <buraysandro9@gmail.com>
2026-09-25 18:28:06 -05:00
Hermes Agent
a975615869 fix(uninstall): sweep macOS launchd dashboard jobs and caches in the full wipe
A full uninstall only removed the gateway launchd label, ~/.hermes, the
checkout, and the desktop's Application Support userData. Left behind on
macOS: every dashboard/serve LaunchAgent (launchd kept respawning its
backend), and the Electron/Chromium + setup cache dirs written outside
HERMES_HOME (~/Library/Caches/Hermes, com.nousresearch.hermes,
hermes-setup, com.nousresearch.hermes.setup) — all of which survived the
"uninstall complete" screen and degraded the next install (#62209).

Extend the full-wipe step: boot out and delete every launchd job whose
ProgramArguments runs a hermes dashboard / serve backend (mirroring the
enumeration hermes_cli.main_dashboard uses to find those jobs, including
the malformed-plist skip contract), and rmtree the four cache dirs. Both
steps are macOS-only and run only in the full wipe; keep-data mode is
unchanged. The dry-run plan now lists them.

Fixes #62209
2026-09-25 18:08:49 -05:00
Brooklyn Nicholson
87d1c5b0c3 refactor(tui): drop redundant native-mode guards 2026-09-25 17:57:38 -05:00
brooklyn!
93f7617406 feat(tui): add native terminal mode defaults 2026-09-25 17:57:38 -05:00
Hermes Agent
6d70abc871 fix(desktop): label the Anthropic OAuth/Pro card 'Anthropic Account'
The accounts page and onboarding picker labeled the anthropic entry
'Anthropic API Key' even though that flow (hermes auth add anthropic)
connects the Claude Pro/Max subscription via OAuth/PKCE. Users with a
subscription read the label literally, concluded they needed an API key,
and went key-hunting. The entry that actually wants a pasted key is
'claude-code' (claude setup-token), which keeps its own name.

Rename the shared PROVIDER_DISPLAY_NAMES entry and the dashboard
provider-catalog name to 'Anthropic Account', and update the onboarding
test assertions.

Fixes #59071
Salvages #59073 (same rename, authored by Kailigithub)

Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
2026-09-25 17:18:06 -05:00
Hermes Agent
43c1eee1fa fix(web): bound the /api/model/info context-length probe
get_model_info called get_model_context_length inline. The resolver
chain runs several sequential provider probes, each with its own
multi-second timeout, so an unreachable model.base_url held the
response for tens of seconds. Bound the whole chain to a 5s budget on
a throwaway worker thread; on timeout the response degrades to
auto_context_length=0 while the abandoned probe finishes on its own.

Fixes https://github.com/NousResearch/hermes-agent/issues/63214 (backend half)
2026-09-25 17:17:48 -05:00
Hermes Agent
954d94c3b0 fix(cli): name Windows Smart App Control when it blocks the dashboard runtime
When Windows Smart App Control or an Application Control policy blocks the
embedded Python runtime's _ssl module, the fastapi/uvicorn import in the
dashboard startup fails with the "DLL load failed ... _ssl" signature, but
the handler printed the generic missing-deps message — so users looped on
`hermes pm install` / `hermes pm repair`, which can never lift a policy
block.

Detect the signature and say so instead: the runtime's DLL is blocked,
repair cannot fix it, and the recovery options are a trusted system
Python, an IT exemption, or the CLI/gateway (which use the system Python).

Fixes #63796

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-25 16:28:14 -05:00
Austin Pickett
fee49791a3 fix(desktop): handle package-manager installs with no desktop source tree
A Homebrew (or other pip/package-manager) install lives in a
site-packages tree that does not ship apps/desktop, so `hermes desktop`
failed with a generic missing-source error implying a broken checkout.
Detect the install kind (Cellar/site-packages), prefer launching a
separately installed /Applications/Hermes.app, and otherwise print a
brew-specific instruction instead.

Fixes #61056
2026-09-25 16:57:57 -04:00
brooklyn!
331a230da6 fix(accounts): stop stale Claude Code connections and false removal success
Accounts Connected now follows token validity instead of access-token
presence. Windows removal uses unambiguous PowerShell, a clear that
removes nothing is an error instead of a success toast, and the
connected-row terminal control runs disconnect.
2026-09-25 14:25:49 -05:00
Hermes Agent
70f7fa05ca fix(providers): keep llamacpp as the model table's managed-runtime id
Mapping the llamacpp aliases to custom in hermes_cli.models sent the
managed local runtime's /model validation down the custom branch before
the staged-library check, so a downloaded-but-not-running GGUF lost its
recognized verdict and an unstaged one was accepted. Keep local and vllm
on custom there, leave the llamacpp aliases on their runtime id, and
move the orphaned 'local' display label to 'custom'.

The parity test now reads the aliases from the custom provider profile
instead of a hand-written list.
2026-09-25 14:25:16 -05:00
PRATHAMESH75
7a5c715844 fix(providers): unify local OpenAI-compatible provider aliases to custom (#62213)
Local OpenAI-compatible server aliases (local, vllm, llamacpp, llama.cpp,
llama-cpp) were normalized inconsistently across the three provider-name
tables: hermes_cli.providers mapped them to the orphan "local" id,
hermes_cli.models left them unmapped, while hermes_cli.auth mapped them to
"custom". Align all three on the generic "custom" provider so routing,
the model picker, and credential resolution agree, matching what
resolve_provider already did for vllm/llamacpp. Add a parity contract test.
2026-09-25 14:25:16 -05:00
Hermes Agent
0cf900ea17 fix(skills): keep built-in provenance for skills the catalog dropped
Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-25 14:24:56 -05:00
Austin Pickett
1b57acf94a fix(mcp): budget discovery-thread GIL so agent build fits the 30s wait
Background MCP discovery does bursty CPU work (SDK/pydantic imports, JSON-RPC
schema parsing, tool registration) that at the default 5ms switch interval rides
the GIL convoy effect and starves the concurrent agent-build thread and the main
event loop for tens of seconds after a serve restart — _wait_agent(30s) times
out with 'agent initialization timed out' (error 5032).

Lower the interpreter switch interval while discovery runs (restored afterwards)
so its CPU bursts are sliced finely enough that waiters stay responsive.

Fixes #60371
2026-09-25 15:01:16 -04:00
Hermes Agent
ec5c9c738a fix(processes): persist_on_release keeps background jobs alive across lifecycle kill sweeps (#41225)
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.

Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
  carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
  (_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
  explicit operator stops (process_manage kill, /stop slash + RPC mirror,
  CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
  persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
  the host is going away and survivors would become PPID=1 orphans

Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
2026-09-25 13:49:31 -05:00
Austin Pickett
b00b4bb7f2 fix(cli): pin utf-8 decoding on all text-mode subprocess readers
On Windows, subprocess text=True without an explicit encoding decodes
child output with the ANSI code page (e.g. 'gbk'); non-ASCII bytes then
raise UnicodeDecodeError inside subprocess._readerthread, killing the
Hermes backend before it becomes ready and surfacing as the desktop boot
timeout.

Sweep every hermes_cli text=True subprocess call to encoding='utf-8',
errors='replace', and add an AST-based regression test that fails when a
future text-mode call omits the encoding.

Fixes #55658
2026-09-25 14:22:12 -04:00
Austin Pickett
fae9e5677a fix(update): stop reading gateway identity off the Windows restart watcher's argv (#107002) (#121635)
* fix(update): stop reading gateway identity off the restart watcher's argv (#107002)

The detached restart watcher is spawned as
`python -c <watcher source> <old_pid> <python> -m hermes_cli.main gateway run`.
Its trailing argv is the command it will spawn LATER, but the canonical matchers
read identity straight off the joined command line, so the watcher itself was
classified as a live `gateway run` process — the documented "never infer process
identity from argv substrings" bug class, on the exact surface `hermes update`
uses to verify a post-update relaunch.

Also budget the post-relaunch liveness poll against the watcher's own deadline:
the watcher respawns the gateway only after the PID it was handed exits, so a
30 s window can expire before the relaunch it verifies was scheduled to start.

* test(windows-live): run the gateway-ancestor harness parent from a script file

A `python -c <src>` parent is an interpreter running inline source and carries
no readable Hermes identity, so it is no longer a gateway to any classifier —
the harness's own comment already said a realistic gateway argv is not a -c blob.

* test(windows-live): share one sleeper SCRIPT across the live process-topology fixtures

Four live Windows E2E files stood processes up as `python -c "sleep" <hermes argv tail>`.
That shape no longer carries a readable Hermes identity, so the fixtures stopped standing
in for the gateways they simulate. One shared sleeper script replaces the -c spelling.

* fix(gateway): drop the duplicated _INLINE_SOURCE_FLAG_RE definition

The constant was emitted twice around command_line_runs_inline_source. Same
pattern both times, so behaviour is unchanged — but one definition is enough.

* test(stderr-timestamp): run the gateway-lookalike children from script files

Both lookalikes stood a gateway child up as `python -c <src> <gateway tail>`.
That shape no longer carries a readable Hermes identity (#107002), so the
wrapper correctly stopped treating them as gateway spawns and the tests failed.
A `-c` tail is data for a program the inline source may spawn LATER, never the
child's own identity — the real wrapper child is `python -m hermes_cli.main
gateway run`, which has no `-c`. Running the stand-ins from a real script file
restores what the tests mean to assert without depending on the misread.

* test(windows-live): restore the tempfile import dropped with the local sleeper helper

* test(windows-live): wait on the sleeper SCRIPT name, not its source text

The live fixtures proved argv visibility by waiting for `time.sleep(120)` in the
spawned process's command line. That string only ever appeared there because the
sleeper was spelled `python -c "import time; time.sleep(120)"`; now that it runs
from a file the source is in the file, so the probe timed out ("sleeper argv
never visible") even though the argv was perfectly visible.

Wait on the script name instead, exported as SLEEPER_MARKER next to the script
so the probe and the spelling cannot drift apart again.

* fix(tests,gateway): keep the live-system guard blocking -c-wrapped gateway spawns

The #107002 identity fix made _gateway_command_subcommand return None for
'python -c <src> … -m hermes_cli.main gateway run'. tests/_fixtures/live_system_guard.py
shares that matcher, so the autouse guard stopped blocking the detached restart
watcher: real gateways leaked out of the e2e run and squatted the webhook port.

Add gateway.status.gateway_spawn_intent_subcommand — the spawn-intent mirror of the
identity matcher, peeling the inline-source wrapper token-wise and re-running the same
canonical matcher on each suffix (still no substring matching) — and point the guard at
it. Read-only subcommands stay spawnable.

* fix(gateway): make the inline-source option walk value-aware so -X utf8 -c is not read as a gateway

The walk that decides whether a command line is an interpreter running inline
source (`python -c <src> ...`) treated every token starting with `-` as a flag
and the first non-flag token as the end of the option block. CPython options
that take a SEPARATE operand (-X/-W/-Q, --check-hash-based-pycs, --jit) break
that model: the operand was mistaken for the end of the block, so the walk
never reached the -c behind it and the watcher was read as a live gateway
again -- exactly the #107002 misclassification, one shape further out.

- Reuse the canonical operand sets from hermes_state_holders rather than
  hand-rolling a second copy (AGENTS.md: parser-derived flag sets).
- Walk case-preserving tokens: operand-taking -Q/-W/-X must not be conflated
  with operand-less -q/-b, so callers no longer lowercase before the walk.
- Handle clustered short options precisely (-uc is inline source, -Xc is -X c).
- Replace the ad-hoc _INLINE_SOURCE_FLAG_RE rescan in
  gateway_spawn_intent_subcommand with the index the same walk returns; the
  regex could not find spellings the walk accepts and would raise
  StopIteration.

Reported by an automated review on PR #121635 and reproduced here.

Refs #107002

---------

Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
2026-09-25 14:15:22 -04:00
Austin Pickett
9df628615c Merge pull request #121614 from NousResearch/austin/fix/121347-base-url-resolution
fix(providers): honor a configured base_url instead of the vendor host across runtime, fallback, listing and validation
2026-09-25 14:15:01 -04:00
Austin Pickett
4a3b54ac7e fix(desktop): diagnose esbuild ignore-scripts build failures
When a desktop build fails because esbuild's platform binary
(@esbuild/<platform>) was never staged - typically because
ignore-scripts=true skips esbuild's postinstall - print an actionable
diagnosis with the fix instead of leaving only the raw build error.

Fixes #53082
2026-09-25 14:12:27 -04:00
Hermes Agent
94f3dbec9b fix(agent): close one-shot AIAgents on every exit path
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:

- batch_runner._process_single_prompt: one agent per prompt, N prompts
  per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
  gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.

Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.

Fixes #50197
2026-09-25 12:49:56 -05:00
Austin Pickett
4661892752 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution
# Conflicts:
#	hermes_cli/models.py
2026-09-25 13:20:26 -04:00
calvinnwq
979c7baeef fix(serve): retire an SSH-isolated backend once the install moves to new code
The serve a remote Desktop spawns over SSH is never restarted by the
host's updater (only its client holds the token and owner nonce), and
the idle watchdog stays quiet while that client is connected. After an
update it therefore kept running the pre-update code against the new
tree until the client happened to reconnect.

Poll the checkout sha against the sha this process loaded. On two
consecutive mismatches, with no update in flight, retire through the
existing retirement fence (which proves idle and closes admission), and
exit cleanly so the client respawns the backend on the new code. Unknown
shas, busy or unreadable ledgers and a live update all keep it up.
2026-09-25 12:09:37 -05:00
calvinnwq
074ead267d fix(update): classify a Desktop SSH serve as its remote client's, not manual-serve
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.

Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
2026-09-25 12:09:37 -05:00
calvinnwq
2e0243b158 fix(desktop): never attach to an isolated serve found in the spawn ledger
hermes serve --isolated (the backend another machine's Desktop spawns
over SSH) opts out of the host singleton on the CLI side, but its spawn
ledger row carried no structured marker, so the local Desktop's
attach-first discovery adopted it. Nothing on this host owns that
process, so a Desktop-driven update left the local app on stale code.

Record isolated=True in the ledger row and skip such rows in
parseSpawnLedger. An ordinary serve is still attached even when an
isolated record is newer.
2026-09-25 12:09:37 -05:00
brooklyn!
296ec08dcf fix(models): keep generation models out of chat
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
2026-09-25 12:07:32 -05:00
Hermes Agent
d0cb567273 fix(gateway): take host identity from the live host record on replay
The settled-flag fix covers a process that multiplexes itself. The update
and fleet processes replay a FOREIGN gateway's captured argv with no
settled flag of their own, so a selector-less argv fell back to the
ambient HERMES_HOME comparison — the exact coordinate the review rejects
(#93943): a host launched from a named profile was replayed as that
profile, donating the named credentials to the respawned host.

Both restart edges now consult, in order: this process's settled
multiplex verdict, then the live host gateway's published rendezvous
record (its SETTLED served set, proven live), and only then the
compatibility default-root comparison. The raw config re-read stays last
so no settled identity exists => unchanged compatibility behavior.

Regressions: selector-less replay from a named home with a live host
record is host; without one it stays profile-scoped; the restart watcher
env takes the default root and drops the named token when only the host
record proves hostness.
2026-09-25 12:01:44 -05:00
Hermes Agent
8bde3a72e8 fix(gateway): preserve settled host identity across restart 2026-09-25 12:01:44 -05:00
brooklyn!
1ac24fa209 fix(gateway): do not let a launching profile own the host gateway
A profile-scoped parent donated its environ to the host multiplexer, so a
named launcher was treated as the primary adapter owner and its platform
token became the primary claim. Spawn the host with served_profile_child_env
for the default root, mark multiplex active before that primary load, and
name the env-derived side in a duplicate-credential refusal.
2026-09-25 12:01:44 -05:00
Austin Pickett
6926997f05 fix(providers): make a configured relay terminal for catalog egress; refuse a foreign Anthropic endpoint
Addresses two P1 review findings on #121614.

1. `provider_model_ids`: a failed or empty relay probe fell through to the
   canonical per-provider fetcher, sending the provider credential to exactly
   the vendor host the user routed away from — recreating #121387 on the
   failure path. A configured `model.base_url` relay is now TERMINAL for live
   catalog egress and degrades to the local curated list instead. The curated
   tail is extracted as `_static_catalog` and shared by both paths.

   Fetchers that already resolve `model.base_url` themselves and degrade
   locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
   simple api-key fetchers) are excluded from interception via
   `_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
   produce a better-merged catalog.

2. `_try_anthropic`: an `explicit_base_url` that failed
   `_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
   the ambient/canonical host and continuing with the explicit credential —
   a silent retarget of authority, not the refusal the PR body claimed. It now
   returns unavailable before client construction.

Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
2026-09-25 12:56:32 -04:00
Austin Pickett
0cc91bb396 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 12:50:14 -04:00
kshitijk4poor
56490ca109 refactor(aux): drop the now-uncalled _read_codex_access_token and retarget its test seams
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
2026-09-25 21:27:06 +05:30
kshitijk4poor
5e2d55ca9c fix(codex): quota probe and /usage pool paths use the pool route base
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).

Refs #121486
2026-09-25 21:27:06 +05:30
kshitijk4poor
e326520d50 refactor(codex): one credential/route authority for aux + image paths
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
  fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
  directly when the pool yields no token (no second uncached pool load,
  no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
  is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
  helper's error fallback reads the profile-scoped override, not the raw
  process env.
- model setup flow: the confirm guards get the resolved Codex base, not
  the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e87f673faa fix(codex): send catalog/image credentials only to their own route
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.

- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
  credential actually routes to (runtime_provider._pool_entry_mode_and_url:
  env > model.base_url while the row is canonical > row URL) instead of the
  ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
  resolved with the token from hermes_cli/models.py, the CLI default-model
  swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
  one pool selection; the image plugin, _build_codex_client and the raw
  Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
  chatgpt.com; a gateway key may probe its own gateway's /models.

Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.

Addresses @andrexibiza's review on #121508.
2026-09-25 21:27:06 +05:30
liuhao1024
602aa9a55b fix(agent): keep custom-Codex-base credentials off chatgpt.com
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):

- the OAuth context-length probe (agent/model_metadata.py) and the
  /model picker's live discovery (hermes_cli/codex_models.py) both GET
  https://chatgpt.com/backend-api/codex/models with
  Authorization: Bearer <gateway key> whenever model.context_length is
  not pinned — the key is sent to a service it does not belong to and
  cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.

Fix, mirroring the quota probe's existing gate in auth_codex:

- both catalog sites now decline to probe non-JWT credentials (real
  Codex access tokens are JWTs; a gateway key is not one) and fall
  back to the static table / offline sources — same outcome as the
  doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
  /models instead of chatgpt.com (catalog URLs are built from the
  resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
  same way the text client does.

Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.

(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
2026-09-25 21:27:06 +05:30
kshitijk4poor
131d4a4634 fix(plugins): picker tick purges every alias of the disable; report real flips
Follow-up to the ported picker fix (#121623):

- Ticking a row re-enables it even when plugins.disabled holds the manifest
  name (e.g. telegram-platform, written by pickers before #40190). The save
  only dropped the key and its bare leaf, so the gate still matched the
  manifest name and the plugin stayed off. It now purges every alias like
  `hermes plugins enable/disable` (_apply_activation, shared with
  _set_plugin_enabled), with one discovery scan per save.
- The success line counts the rows actually turned on/off instead of treating
  every unticked row as "disabled"; the unused new_enabled return is gone.
- _entry_status shares the per-row status call between the picker
  preselection and `hermes plugins list --enabled`.
2026-09-25 21:19:04 +05:30
Victor Kyriazakos
bfb8d3bd64 fix(plugins): picker preselects the effective state and persists only flipped rows
The bare `hermes plugins` picker preselected only rows listed in
plugins.enabled. Bundled platforms, backends and model providers are active
without a list entry, so they opened unticked, and the save on exit wrote every
unticked row into plugins.disabled: opening the picker and leaving without a
change disabled every messaging adapter on the next gateway restart.

Rows now open ticked by the load-time rule (_plugin_status), and only rows the
user flipped are written.

Ported onto the plugins_cmd_toggle sibling and the admission-authority save
(the original targeted the pre-split plugins_cmd.py). The nested
canonical-key composite test now ticks its row explicitly: it handed the menu
a checkbox state that contradicted the config and relied on the old
rebuild-everything save. The original helper-level tests are rewritten
against real admission in a follow-up commit.

(cherry picked from commit 3a0d5f1b8bce3b23b115ba4995b8c23dfbba9ad0)
2026-09-25 21:19:04 +05:30
Hermes Agent
f8d35574ab fix(recovery): repair out-of-window timestamp cells in the recovered database 2026-09-25 10:46:54 -05:00
Austin Pickett
5af33f8375 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 09:09:44 -04:00
ethernet
268820136f fix(update): repair an install whose venv runs another checkout
An install's in-tree venv whose editable record names another checkout
(what project_venv_dir used to cause) turns <install>/venv/bin/hermes
into that checkout's CLI. Every update Desktop hands to the launcher
then pulls the other tree: the install never moves, the hand-off
reports success, and posix.sh swaps the install's stale release/ build
back over the app. Pressing Update can never get out of that loop.

`hermes update` now notices it is running on another checkout's in-tree
venv and re-runs itself with that install's own code (PYTHONPATH=<install>,
cwd=<install>). The install's updater pulls the install and reinstalls it
into its venv, which points the editable record home again, so the next
relaunch runs a fresh build. A wrapper that put the running checkout on
PYTHONPATH on purpose chose that tree and is left alone.
2026-09-25 01:21:28 -04:00
ethernet
138e33d51f Merge pull request #122244 from NousResearch/fix/restore-pm-merge-drops-current
Restore behavior lost during PM integration
2026-09-25 00:36:50 -04:00
ethernet
af3299a22c fix(gateway): don't ask the Windows login question when stdout is captured
#122234 gave Desktop update steps NUL stdin, but the hand-off script
that runs is the one from the checkout being updated FROM. Every update
that starts on an older commit still runs the old script, which gives
steps the hand-off console as stdin and captures their stdout until
they exit. Only the steps after `hermes update` run new code.

So `gateway start --all` from the new checkout still saw an interactive
console, asked "Install it now so the gateway starts on login?" into
the captured stdout, and waited forever. The update never relaunched.

start() now asks only when stdout is a terminal too. Nobody can answer
a question they cannot see.
2026-09-25 00:33:42 -04:00
ethernet
2ef41d2b58 Merge pull request #122103 from NousResearch/ethie/pm-evict-incompatible-plugins
fix(pm): hermes update disables plugins that no longer fit instead of failing
2026-09-25 00:08:10 -04:00
ethernet
a75d8b420e Merge pull request #122234 from NousResearch/fix/handoff-noninteractive-steps
fix(desktop-update): Windows update steps get NUL stdin; installer asks gateway questions once
2026-09-25 00:05:41 -04:00
ethernet
ccf631701a fix(pm): disable plugins that no longer fit instead of failing the update
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.

Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
2026-09-25 00:00:09 -04:00
ethernet
abe76b5176 fix: restore upstream behavior lost in PM conflict resolutions 2026-09-24 23:59:55 -04:00
ethernet
97c4fa022c fix(update): refresh release tags before stamping source versions 2026-09-24 23:52:31 -04:00
ethernet
746d861504 fix(install): don't ask the gateway install questions twice
The setup stage installs the gateway service through
ensure_gateway_service. On Windows that asks the start-now, Scheduled
Task and UAC questions. The gateway stage then ran `hermes gateway
install`, which asked them all again.

`gateway install --if-missing` does nothing when a service is already
installed. Both installers' gateway stages use it, so they ask only
when setup did not install the service.
2026-09-24 23:49:18 -04:00
ethernet
6ce9ed4223 fix(update): key the early-spawn sync on currency, not on a commit
Under the updater's claim, sync whenever dependencies are not current
rather than only when nothing is committed. The tail's own children
and post-sync verification children are current and stay no-ops. A
stale generation still committed from the previous Python pin is the
same ABI trap as the pre-PM venv, and it now syncs too. If the sync
still leaves the tree out of date, raise instead of relaunching into
another sync.
2026-09-24 22:56:59 -04:00
ethernet
3a42f0fe10 fix(update): commit dependencies for processes the update spawns early
prepare_launch returned early for any process running under the
updater's own claim, so it would not re-run the completion tail. That
also covered processes the updater spawns before PM commits a
generation (a restarted gateway), which then booted with no
environment: previously on the pre-PM venv, now refused.

Under the updater's claim with nothing committed, sync the dependency
generation (carrying the legacy venv's extras, as the first sync
always has), skip the tail since that belongs to the updater, and
relaunch on the store Python. The relaunched process sees the commit
and returns early as before, so the no-recursion guard still holds.
2026-09-24 22:53:06 -04:00
brooklyn!
20bcc9bd14 fix(tts): stream Edge speak-stream per sentence instead of whole-text fallback
Desktop speak-stream sent type=fallback whenever the provider had no
chunked PCM API. Edge is that case, so the client waited for the full
reply and POSTed it. Cut sentences with the existing sync TTS tool and
stream that PCM. Fallback stays the last resort when synthesis produces
no audio.

Refs #91997
2026-09-24 19:28:05 -05:00