Commit Graph

276 Commits

Author SHA1 Message Date
teknium1
b80292c9df fix(dashboard): resuming a chat never lands inside a subagent or /branch child row
After a gateway restart (ws_orphan_reap) the dashboard chat resumes the
predecessor's "latest descendant". `_session_latest_descendant` followed
EVERY parent_session_id edge, so when the last turn had spawned a subagent
(tools/delegate_tool.py stamps `_delegate_from` on a child row whose parent
is the primary chat) the resume target became that subagent row. The user's
chat then continued inside a row that `GET /api/sessions` hides
unconditionally (`_delegate_from IS NOT NULL`), the sidebar kept showing the
reaped predecessor with its stale title, and /title in the live chat
disagreed with the list — exactly the forensics in #115092. The same walk
followed /branch forks, /new reset children and tool-owned rows.

`_DESCENDANTS_SQL` now follows continuation edges only, with the predicate
the session list's chain CTE already uses (no `_delegate_from`, no
`_branched_from`, not a reset child, source != 'tool'). Compression
continuations still resolve to their newest leaf.

The reap-continuation writer the issue title blames does not exist on main:
the only `_delegate_from` writer is the delegate spawn path, and this is how
a subagent row became the "successor".
2026-09-20 12:39:58 -07:00
teknium1
4ff5d91ac7 chore: merge origin/main (resolve apps/desktop/src/app/settings/custom-endpoints-settings.tsx, apps/desktop/src/types/hermes.ts, hermes_cli/web_routers/config_env.py, tests/hermes_cli/test_web_server.py) 2026-09-19 10:54:40 -07:00
teknium1
9202bae546 fix(dashboard): custom endpoint validation persists the base URL that served /models (#65488)
A custom OpenAI-compatible endpoint typed without `/v1` in the Desktop
onboarding or Settings > Custom endpoints flow was probed only at
`{base}/models`, while the CLI's probe_api_models falls through to the
`/v1` alternate. Whatever the probe reported, the URL was saved verbatim
and the runtime POSTs `{base_url}/chat/completions` to it, so a server
that only serves `/v1/*` 404'd every chat request.

Both dashboard validators (`/api/providers/validate` OPENAI_BASE_URL
branch and `/api/providers/custom-endpoints/validate`) now share one
probe that tries the URL as entered and then its `/v1` variant (or the
stripped variant), and return `resolved_base_url` — the base that
actually served the model list. The Desktop onboarding persists that
URL, and the Settings "Test" rewrites the form's URL to it so the
following Save stores a URL chat can reach.

Co-authored-by: Jeongseok Kang <jskang@lablup.com>
2026-09-19 10:04:52 -07:00
kshitijk4poor
8c1f8ed646 fix(web): the dashboard's submitted custom endpoint survives the model assignment
A bare-`custom` main-slot pick hands the submitted endpoint to `switch_model` as the current
one, and the applied config used to carry it back verbatim. Once the credential step
re-resolves that target, an env/config endpoint (`CUSTOM_BASE_URL`, a stale `model.base_url`,
the `OPENROUTER_BASE_URL` mirror) could replace what the user typed — and
`_apply_main_model_assignment` then persisted it over the submission.

`_validated_main_model_selection` now restores the submitted endpoint on the result for a bare
`custom`/`local` target, together with the wire protocol that endpoint mandates: `model.base_url`
and `model.api_mode` are persisted as a pair, so a mode derived from the displaced host (an
`anthropic_messages` host behind `CUSTOM_BASE_URL`, say) would have been saved next to the
submitted URL.
2026-09-19 22:05:32 +05:30
teknium1
50638799b4 chore: stack on #115828 to resolve config_env.py conflict
validate_custom_endpoint resolves the base via #115828 _probe_openai_compatible_models first, then runs the
transport probe (_probe_transport_route / _auto_api_mode) against the resolved base and returns models,
model_details, transport_checked and resolved_base_url; TS keeps both resolvedBaseUrl update + transport notify.
2026-09-19 03:36:30 -07:00
teknium1
ab241af386 fix(desktop): custom endpoint Test probes the transport route, not just /v1/models
A Responses-only (or Anthropic-compatible) host lists models on GET /models and
then 404s every POST /chat/completions, so Test passed and the first chat
failed (#93622, item 2). validate_custom_endpoint now keeps the probe client
open and, after the models GET, POSTs a one-token request to the route of the
pinned api_mode — or of the mode the runtime's URL auto-detect resolves to
(_detect_api_mode_for_url, fallback chat_completions, the same order
runtime_provider_custom uses). 404/405/501 fails validation with a message that
names the transport and the route; 200/400/401/422/429 prove the route exists;
a network error or timeout stays inconclusive so a local server still loading
its model is not blocked. The response reports the mode it checked as
transport_checked; the Desktop success toast names it so an auto-detected mode
is visible before Save.

The Providers-pane OPENAI_BASE_URL probe returns model_details next to models
too, so an alias picked there keeps its canonical_model / reasoning_effort.

Co-authored-by: SacrEllfarch <2091538824thx@gmail.com>
Co-authored-by: JackLee992 <124239570+JackLee992@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-19 01:32:20 -07:00
teknium1
363b8a6fcb fix(desktop): custom endpoints pin an API mode and keep /v1/models alias metadata
Settings > Custom Endpoints assumed Chat Completions: the form, its types, the
update payload and _write_custom_endpoint carried no api_mode, so a
Responses-only (or Anthropic-compatible) host validated fine on /models and
then 404'd on every POST /chat/completions. Validation also flattened each
/v1/models row to a bare id, so a reasoning alias like gpt-5.6-sol-high
(canonical_model + reasoning_effort) was saved as a literal upstream model.

- Desktop form: API Mode segmented control (Auto-detect / Chat Completions /
  Responses API / Anthropic Messages — the same set `hermes model` offers);
  threaded through toPayload, hydrated from GET read-back.
- CustomEndpointUpdate.api_mode (Literal) persisted as providers.<id>.api_mode,
  the key the CLI writes and the runtime reads; None (older UI) leaves a
  hand-written mode alone, "" clears it. GET rows report api_mode.
- validate returns model_details (id / canonical_model / reasoning_effort)
  next to the unchanged string[] models; _parse_model_ids is now a projection
  of _parse_model_entries.
- Save keeps the alias metadata in providers.<id>.models and, when the picked
  default is an alias, persists the canonical model and pins its effort under
  agent.reasoning_overrides (the resolve_reasoning_config chokepoint).

Fixes #93622
Supersedes #69824 (@SacrEllfarch), #82148 (@JackLee992), #93693 (@fangliquanflq)
2026-09-19 00:17:19 -07:00
teknium1
859c883ce5 fix(dashboard): custom endpoint validation persists the base URL that served /models (#65488)
A custom OpenAI-compatible endpoint typed without `/v1` in the Desktop
onboarding or Settings > Custom endpoints flow was probed only at
`{base}/models`, while the CLI's probe_api_models falls through to the
`/v1` alternate. Whatever the probe reported, the URL was saved verbatim
and the runtime POSTs `{base_url}/chat/completions` to it, so a server
that only serves `/v1/*` 404'd every chat request.

Both dashboard validators (`/api/providers/validate` OPENAI_BASE_URL
branch and `/api/providers/custom-endpoints/validate`) now share one
probe that tries the URL as entered and then its `/v1` variant (or the
stripped variant), and return `resolved_base_url` — the base that
actually served the model list. The Desktop onboarding persists that
URL, and the Settings "Test" rewrites the form's URL to it so the
following Save stores a URL chat can reach.

Co-authored-by: Jeongseok Kang <jskang@lablup.com>
2026-09-19 00:06:15 -07:00
teknium1
5ae7ec8623 fix(setup): a provider configured after boot unblocks the dashboard chat without a restart
The serve process's free-tier boot record (`free_tier_bootstrap._record`) is
built once; when the boot inventory found nothing (mint failed or the tier is
off) it says `provider_configured: false` for the process lifetime and
`setup.status` answers from it, so the Ink chat (dashboard /chat, `hermes
--tui`) parks every new session on "Setup Required" no matter what the user
configures afterwards — the Models page, a picker key save, or `hermes setup`
from a shell all land on disk and change nothing in the running process.

Reconcile the record on read instead of re-minting on write: `reconcile_record`
re-runs the cheap inventory (`_inventory_other_providers`, the resolver ladder
with the free-tier rung hidden) when a `False` record's config files
(`config.yaml` / `.env` / `auth.json`) moved since the inventory that built it,
replaces the record's inventory half and broadcasts `setup.ready`. The mint
verdict is kept as is: only the boot bootstrap and its retries mint.
`wait_for_record` (what `setup.status` reads) reconciles, which also covers
writes this process never saw (`/setup`'s `hermes setup` handoff, `hermes model`
in docker exec, a hand edit); the two in-process write paths from #114708
(`/api/model/set`, `model.save_key`) call it for the immediate broadcast.

Dropped from #114708: `retry_bootstrap_mint(force=True)` on the write paths — a
forced portal mint (network, cooldown bypassed) inside an HTTP write; the record
only needed its inventory refreshed. Not taken from #108773: OR-ing the record
with `_has_any_provider_configured()` — that first-run guard counts host-wide
credentials (other host-wide agent-CLI logins) and answers True on a blank machine (verified
live on this host), which would open the gate with nothing configured.

The 4001 hint for a sessionless `config.set model` named "Settings -> Models",
a Desktop-only surface; the Ink chat has no Settings and the dashboard has a
Models page. One string now names /setup, the Models page and Settings ->
Models. The TUI's "Setup Required" panel no longer advertises `/model` as the
in-place fix (sessionless picks are refused by design, 75b1b43ec1).

Co-authored-by: chelsealong <chelsealong@126.com>
2026-09-18 12:52:36 -07:00
kokhlo
1979f70429 fix(dashboard): refresh the setup record after a main-slot model assignment
A serve process whose boot-time free-tier mint failed keeps its bootstrap
record at provider_configured: false for the whole process lifetime, so the
web chat stays gated on "need setup" even after a provider lands on disk via
POST /api/model/set or a picker key save. Re-run the bootstrap inventory on
both write paths so setup.status (and the setup.ready broadcast) follow the
new configuration without a restart.

Fixes #114697
2026-09-18 12:52:36 -07:00
teknium1
76c640ff11 fix(web): list, activate and delete legacy custom_providers entries on Custom Endpoints
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.

The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
2026-09-18 12:44:32 -07:00
teknium1
08b791b9a3 fix(config): custom endpoint list marks a mixed-case key as current after activate
`switch_model` stores the active provider as `custom:<lowercased name>`,
so `GET /api/providers/custom-endpoints` compared the literal stored key
(`EXllamav3`) against `custom:exllamav3` and marked nothing current right
after a successful activate. The list now uses the same dual-spelling
match the delete path already applies to the model mirror.
2026-09-18 09:23:16 -07:00
teknium1
66dfecca89 fix(config): custom endpoint save/delete detach the mirror for every key spelling
Follow-up to the cherry-picked #114573: close the rest of the class.

- Save (edit) path: `_write_custom_endpoint` slugified `body.id` before the
  lookup, so editing a listed `local-127.0.0.1:8283` row forked a
  `local-127-0-0-1-8283` twin next to the original (the third site #92019
  named). Resolve the stored key first, slug only for fresh names.
- Delete path: `switch_model` writes the activated provider either as the
  stored key or as `custom:<lowercased name>`; the detach compared the raw
  key case-sensitively, so deleting a mixed-case endpoint (`EXllamav3`)
  removed the entry but left `model.base_url`/`provider` routing the agent
  to the deleted host (#62269 class). Match both spellings.
- Numeric YAML keys (`providers: 2070:`) resolve to an int stored key;
  coerce before it reaches `switch_model`/the response.
- Tests: replace the helper-level unit tests with two route-level
  invariants in tests/hermes_cli/test_web_server.py (list → save → activate
  → delete round-trip for dotted and mixed-case keys; slug fallback + 404
  control).

Co-authored-by: dmf01614-source <244684154+dmf01614-source@users.noreply.github.com>
2026-09-18 09:23:16 -07:00
teknium1
693430d010 test: give bare profile-dir fixtures an identity marker (#113377 red-fix, pass 2)
The roster / bot-relay known-profile set / sudo-user profile resolver now go
through named_profile_is_live, so fixtures that mkdir a bare profiles/<name>
dir and expect it to be served or resolved need a real identity marker.
Each fixture touches profiles/<name>/config.yaml right after the mkdir, the
same way production creates a profile. No production code changed.

Files:
- tests/hermes_cli/test_web_server.py
  (test_messaging_platforms_profile_scopes_gateway_reads,
   test_custom_endpoint_save_scopes_to_the_requested_profile)
- tests/tools/test_bot_retry_policy.py (shared fixture, 4 tests)
- tests/tools/test_bot_turn_lock.py
  (test_relay_deliver_returns_target_busy_error,
   test_relay_deliver_serializes_then_succeeds)
- tests/tui_gateway/test_relay_live_author.py
  (test_live_relay_stamps_the_sender_as_a_delivery_author)

Contract rewrites: none — no assertion pinned the old bare-dir-is-a-profile
contract; every change is fixture-only.

Sweep: every test file referencing _roster / bot_mode_probe / bot_relay /
list_profiles / _resolve_sudo_user_profile_env (52 files, 1687 tests) runs
green with these changes.
2026-09-16 16:48:42 -07:00
KoNit-K
9af62e7ef0 fix(dashboard): skip unreadable plugin manifests 2026-09-15 18:48:59 -07:00
teknium1
bdc7916196 fix(ux): plain-language, actionable user-facing messages (dashboard)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 03:36:22 -07:00
teknium1
9188e708b3 fix(gateway): served_profiles bind to a verified gateway identity, not bare PID existence
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.

One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
2026-09-13 15:41:01 -07:00
teknium1
0885e19a75 fix(dashboard): a rejected model switch is a 400, never a flattened model: block
_denormalize_config_from_web ran the new switch_model validation inside the
pre-existing 'except Exception: pass' disk-read fallback. Any rejection
(offline models.dev, no OpenRouter key, unlisted model) left config['model'] a
flat string, and PUT /api/config's deep-merge wrote that string over the on-disk
model: dict, destroying provider/base_url/api_mode/context_length/model_slots.

Only the load_config() read keeps its fallback; validation propagates as the
HTTPException(400) the caller's http_failure passes through. Invariant test PUTs
a rejected model against a real config.yaml and asserts 400 + byte-identical file.
2026-09-13 05:21:02 -07:00
teknium1
11576390fe refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.

Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.

Sites -> canonical:
  hermes_cli/cli_model_switch_mixin.py::_persist_global_switch          -> deleted; _commit_model_switch calls persist_model_selection
  hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
  gateway/slash_commands_model.py::_persist_model_switch_to_config       -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
  tui_gateway/model_switch.py::_persist_model_switch                     -> deleted; _apply_model_switch calls persist_model_selection
  hermes_cli/web_server_config.py::_apply_main_model_assignment          -> apply_model_selection(result) (+ explicit custom api_key)
  hermes_cli/web_server_config.py::_validated_main_model_selection       -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
  hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
  acp_adapter/server.py::_resolve_model_selection                        -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError

Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.

Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.

Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
2026-09-13 05:21:02 -07:00
Teknium
c9b71ebaf7 fix(dashboard): startup schema reconcile opens state.db read-only first
The dashboard's `_eager_reconcile_own_session_db` did an unconditional
writable `acquire()` at every startup. When the gateway shares that
state.db the dashboard became a second long-lived writable SessionDB
owner: a close-time WAL checkpoint plus a possible FTS rebuild in
`_init_fts`, the two-writer vector behind the corruption reports in
#107688 and #100896 ("5 live SessionDB handles" precursor, gateway +
dashboard both holding the WAL).

Route the startup reconcile through `_open_session_db_at_path(...,
read_only=True)`, which already bootstraps a missing store and heals a
stale/malformed schema through exactly ONE writable open before
reopening read-only. A healthy store now gets zero writable opens from
the dashboard while the #79531/#80037 "bring schema current before the
first poll" contract is kept (existing heal test unchanged).

Live repro (healthy store, count writable SessionDB.__init__ calls from
the startup worker): before=1 after=0.

Reported-by: #107688, #100896 (@kokhlo diagnosis)
Refs #107688 #100896
2026-09-11 06:21:48 -07:00
Teknium
2466684db5 fix(models): never auto-switch to a provider the user has no credentials for
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).

detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.

Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
2026-09-10 03:58:37 -07:00
Teknium
d0c90039a7 fix(messaging): keep profile status truthful without credential inheritance
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
2026-09-07 05:56:33 -07:00
liuhao1024
5dfa2f8374 fix(dashboard): env credentials enable a platform on the profile-scoped messaging status path
The scoped branch of _platform_enablement consulted only config.yaml's
platforms: section, but the `hermes gateway setup` wizard writes .env
credentials and never a platforms: entry. The desktop always sends
?profile=default (normalizeProfileKey maps the primary profile to
`default`), so the Settings - Messaging page showed a working bot as
"Disabled" while /api/status reported it connected (#104614).

Mirror _enable_from_env (gateway/config_env.py): env credentials alone
enable a platform, an explicit enabled: false still wins. Only the
profile's own .env (env_on_disk) is consulted, so the root install's
os.environ credentials still never leak into a profile's state.

Fixes #104614
2026-09-07 05:56:33 -07:00
杨瑾
d347fcaf97 test(desktop): cover live token in SPA index 2026-09-05 16:30:35 +05:30
thinkingsheep
1866921d33 fix(desktop): serve the current SSH session token 2026-09-05 16:30:35 +05:30
Teknium
da96762ef8 fix(compat-fallout): repoint 8 test files off dropped facade names (run_agent._DB_PERSISTED_MARKER/_EPHEMERAL_SCAFFOLDING_FLAGS/_is_ephemeral_scaffolding, cli.AIAgent, main._make_tui_argv/_print_curator_recent_run_notice/_format_time_ago, model_switch._collect_authed_provider_slugs, acp has_provider shim); arg_coercion test imports model_tools to populate registry 2026-09-03 15:50:55 -07:00
Teknium
4fde117f4b simplify(compat): tests — repoint 638 web_server.<name> references (321 attr, 217 monkeypatch/patch.object, 60 from-imports, 40 patch() strings) across 65 test files to the owning modules 2026-09-03 14:21:52 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
3f0e1165b3 refactor(tools_config): extract post-setup hooks into tools_config_post_setup 2026-09-02 16:11:26 -07:00
kshitijk4poor
af964412d9 test(web): drop idx_sessions_effective_activity before emulating a pre-column store
The salvaged #95487 adds an expression index over sessions.last_activity_at.
Two web_server tests emulate a legacy store with ALTER TABLE ... DROP COLUMN
last_activity_at, which SQLite refuses while an index references the
column ('error in index ... after drop column'). Drop the index first —
the same adjustment the PR already made in test_schema_read_probe.py. A
real pre-column store has neither the column nor the index, so the healed
path under test is unchanged.
2026-09-03 04:23:35 +05:30
Brooklyn Nicholson
3a0e7df799 fix(state): a busy session store reads as busy, not as damaged or empty
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.

Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.

Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.

On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.

Fixes #100436

Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
2026-09-01 22:42:19 -05:00
Agi-Asi
54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
chelsealong
6d407ca1a4 fix(desktop): stop model_context_length edits from being dropped or wiped
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.

Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
2026-08-31 10:07:17 -07:00
Alvin T. Veroy
e17fd0a708 fix(state): decode errors now reach the heal path and fail loud in TUI (residual #98924 surfaces)
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:

- web_server._open_session_db_at_path: the one-writable-open heal only
  caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
  to decode SQLite's own error message over corrupt file bytes) bypassed
  it, so the heal documented for malformed schema never fired (#98924
  Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
  vtables whose probe raised UnicodeDecodeError, the same too-narrow
  catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
  could not open, so prompt.submit streamed the turn while persisting
  nothing (#98924 Failure 2). It now returns False and prompt.submit
  fails the RPC with code 5072 so desktop maps it to a toast, mirroring
  the disk-full/5070 convention. session.create stays silent per its
  pinned degraded-mode contract.
2026-08-31 09:56:43 -07:00
Brooklyn Nicholson
2119ed7b4a test(model): cover numeric YAML provider keys in picker and CRUD
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
Teknium
39a5aa91ed fix(serve): serve Desktop token page at / in headless mode (#94227)
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).

Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.

Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
2026-08-26 15:51:22 -07:00
Teknium
847af7301a test: re-pin refusal exit codes and gate patch points to the shared contract
The apt/docker CLI tests pinned exit 1; refusals are now exit 2
(refused-by-contract, distinct from errors). The web_server guards
patched the module-local detect_install_method alias, which the shared
admission gate no longer consults — patch hermes_cli.config directly.
2026-08-26 11:41:04 -07:00
codexbt
3dea11d703 fix(web_server): recheck WEB_DIST existence dynamically in mount_spa
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".

- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.

Closes #82614
2026-08-26 04:48:58 -07:00
fangliquanflq
66186dc58f fix(desktop): keep bot reconciliation off inactive backends 2026-08-26 00:21:29 -07:00
Teknium
1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Mauricio Ruiz
091106092b fix(dashboard): recover update success after restart
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
2026-08-23 00:50:41 -07:00
Bruce Xu
50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
poisdahl
fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
poisdahl
abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Teknium
19113b34d1 test(tools): repin selector/picker tests to the provider-string contract
Update the sibling tests that pinned the old use_gateway-writing
contract: image/video selector and reconfigure rows now assert the
single provider string ('nous' managed / 'fal' BYOK) plus legacy-key
popping, the stt/video picker writes drop the use_gateway expectation,
the web_server managed-browser select asserts the persisted 'nous'
cloud_provider, and explicit-local STT pins no-cloud-fallback against
a stored raw-config selection.
2026-08-19 16:10:01 -07:00
Teknium
c9ce66e25e fix(desktop): one roster row per bot even when a backend is registered under two addresses — /api/status install_id + roster collapse
Backend: /api/status now carries a stable random install_id persisted once
under the root HERMES_HOME, shared by every profile of the install.
Desktop: roster enumeration captures it per connection (TTL-cached probe),
buildAgentRoster collapses same-install rows with a deterministic canonical
pick (active > local > ssh > remote > cloud > earliest), the @name-device
handle rule runs after the collapse, and Settings → Gateways shows a
display-only 'Same backend as' hint. Backends without install_id bypass the
collapse (fully backward compatible).
2026-08-17 19:22:57 -07:00
adybag14-cyber
d685ea4df5 docs(termux): document community pkg distribution 2026-08-17 15:47:17 -07:00
poisdahl
7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
Teknium
1e6717baf2 fix(web_server): discover root user plugins under profile-scoped processes
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.

Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.

Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.

Fixes #87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
2026-08-16 02:07:16 -07:00
Nicolas Formenton
c75d835555 feat(desktop): mark a session as unread/read with a persisted watermark 2026-08-15 01:21:40 -07:00