fd602278c7097574143fc9ad64bb8ce6f0441ab9
980 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fd602278c7 |
fix(tui-gateway): report Fast only where the route accepts the priority tier
A profile-wide agent.service_tier: fast reaches every new session, including local llama.cpp ones. The request builders already drop the fast parameter on routes that do not bill for it, but session.info reported fast: true whenever the tier was priority. The desktop then kept the Fast switch visible so it could be turned off, tagged the model row Fast and appended "· Fast" to the pill for a local model. session.info now asks resolve_fast_mode_overrides, the gate the request builders use, before it reports fast. While a model switch is pending, the new pick decides, because the agent's base URL still belongs to the old route. service_tier is still reported as stored. The TUI status bar reads info.fast alone. |
||
|
|
179124cdff |
fix(gateway): carry customCSS in the change-watcher's resolve_skin
The salvage landed customCSS in server.py's resolve_skin, but that function's runtime home is change_watcher.py: method_ctx.bind_module rebinds the split module's bodies onto server.py's globals at import time, silently overwriting the server.py copy — so skin.changed went out without customCSS and the desktop lost user CSS on every switch. Carry the field in change_watcher.py's resolve_skin (the live source of truth) and drop the duplicated watcher block from server.py, restoring main's single-owner layout. |
||
|
|
c60ba75199 |
fix(desktop): pass skin customCSS through to desktop theme context
The web dashboard supports customCSS in theme YAML (PR #14776) but the desktop app's skin pipeline dropped the field. Users who wanted custom styling had to hack app.asar, which gets overwritten on every update. This adds customCSS passthrough through the full skin pipeline: - HermesSkin (apps/shared) and DesktopTheme types gain customCSS?: string - skinToDesktopTheme() passes skin.customCSS through - applyTheme() injects a scoped <style id="hermes-desktop-custom-css"> tag on theme apply, and removes it when switching to a CSS-less theme - SkinConfig gets custom_css field, read from YAML as "customCSS" - _build_skin_config() caps at 32 KiB (same as the web dashboard) - resolve_skin() emits "customCSS" in the gateway JSON-RPC payload Users now put CSS in ~/.hermes/skins/<name>.yaml under the customCSS key — it persists across updates because the skin dir is outside app.asar. Closes #53013 Refs #53012 |
||
|
|
903a57c540 |
fix(tui_gateway): subtract disabled_toolsets from the client-surface fold-in
_get_platform_tools correctly subtracts agent.disabled_toolsets, but _with_session_toolsets unions the client-surface (and profile-role) toolsets back in afterwards — so disabled_toolsets: [project] was a no-op on desktop/TUI, the only surfaces where the toolset exists, and the agent kept hijacking the word "project" (#54433). The fold-in now applies the same subtraction via _load_disabled_toolsets, keeping desktop_ui (the client own control surface, not a model toolset). Covers both the focus/coding posture path and the configured fallback. Fixes #54433 Co-authored-by: Hermes Agent <agent@hermes.local> |
||
|
|
7131fe3f03 |
fix(desktop): warn when the GUI is older than the backend (reverse contract check)
reportBackendContract() warned only when the backend reported a contract LOWER than the GUI's required value. The reverse skew — a GUI build older than the backend it drives — passed silently. That happens in practice to any long-running desktop app that survives a backend update without relaunching; the stale GUI then drives newer gateway code and fails cryptically downstream. Warn in both directions: - backend contract < required (existing): "Backend out of date" toast with the one-click backend update action, unchanged. - backend contract > required (new): "Hermes app out of date" toast pointing at the client update overlay (openUpdateOverlayFor on the client target). The new toast mirrors the existing skew toast's ergonomics: persistent 24h snooze on dismiss, immediate clear once the two sides align, and snooze reset on alignment so a later skew warns right away. Each direction dismisses the other's toast, so at most one skew warning is ever live. i18n keys added to all shipped locales; DESKTOP_BACKEND_CONTRACT docs in tui_gateway/server.py now describe both directions. Fixes #60542 Co-authored-by: Hermes Upgrade Staging <alex@vencounsel.com> |
||
|
|
3bbe5c593a |
fix(tui): propagate the live agent's provider into the slash worker (#57283)
A Desktop/TUI MoA session pins the live agent to the virtual moa provider, but the persistent slash-worker subprocess was spawned with only --model. HermesCLI re-resolves provider from config.yaml, so /moa dispatched the MoA preset NAME (model='default') to the configured real provider — openrouter 402 max_tokens errors in the original report. Forward the parent agent's resolved provider on the worker argv (--provider) and pass it through to HermesCLI at both spawn sites (first-use in slash.exec and the post-switch _restart_slash_worker). The issue's first layer (active-profile.json persistence) was already fixed on main by d27180ba7f; this is the remaining provider-propagation layer. |
||
|
|
0cf900ea17 |
fix(skills): keep built-in provenance for skills the catalog dropped
Co-authored-by: fangliquanflq <fangliquan@qq.com> |
||
|
|
35ad70c035 |
fix(tui-gateway): measure reap sleep by clock divergence, re-arm via scheduler
Detect the host sleep by the divergence between the wall clock and the monotonic clock since arming (the monotonic clock does not advance while asleep), instead of an absolute monotonic deadline — a timer that fires without elapsed wait time (tests, spurious dispatch) shows zero divergence and reaps normally. Re-arm through _schedule_ws_orphan_reap with the fired timer as the expected one so the fresh closure's identity guard matches the entry it installs. |
||
|
|
eda61f644e |
fix(tui-gateway): grant the WS orphan-reap grace in awake time
threading.Timer's wait elapses in wall-clock time on platforms without a monotonic condvar (macOS lacks pthread_condattr_setclock), so a system sleep made the WS-orphan reap timer fire at the instant of wake — before the Desktop app's WS reconnect or session.resume could re-bind a transport — and every sleep/wake cycle longer than the 20s grace 404'd the open session. Record a monotonic deadline when arming the reap and re-arm for the remaining awake time when the wall-clock timer fires early; a reconnect still cancels the chain as before. Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com> |
||
|
|
788d82fc97 |
fix(tui): honor agent.disabled_toolsets in the desktop/TUI gateway (#44499)
The classic CLI and the messaging gateway forward agent.disabled_toolsets to AIAgent, where get_tool_definitions strips the named toolsets even out of composite defaults like hermes-cli (#17309). The desktop/TUI gateway dropped the list, so disabled_toolsets: [browser] silently had no effect on Desktop — the only consumer was _get_platform_tools' name-level subtraction, which cannot reach inside a composite default toolset. Users configuring an explicit BrowserOS MCP server had no way to suppress the built-in browser_* tools, and the agent then favored the unauthenticated built-in browser. Forward the config list at every gateway agent construction and refresh site: _make_agent, _background_agent_kwargs, the reload.mcp per-session refresh (refresh_agent_mcp_tools disabled_override) and tools.show. Salvaged from #44505 (AIalliAI). Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com> |
||
|
|
0f65d698d0 | Merge remote-tracking branch 'origin/main' into ethie/pm-clean | ||
|
|
1674499d00 |
fix(compression): archive the rows the compressor held, not every row up to the newest
A gap below the newest held id, and turns appended above an unpersisted current turn, were summarized away without being read. Name those held ids and clone the rest. |
||
|
|
ed6e37f9c1 | Merge remote-tracking branch 'origin/main' into ethie/pm-clean | ||
|
|
177f275b77 |
fix(plugins): route inject_message to the TUI session_key queue
Ink TUI and desktop never registered an inject host, and sharing set_gateway_message_injector with a live messaging gateway would let the last writer win. A separate host queues the reported session_key onto that session's prompt queue and leaves other keys for the gateway. Fixes #87412 |
||
|
|
66d54ebf51 |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # .github/workflows/docker.yml # Dockerfile # apps/desktop/src/app/settings/about-settings.tsx # docker/stage2-hook.sh |
||
|
|
e7cc3836e1 |
fix: TUI auth fallback no longer crashes on model: <id> config shorthand
`_resolve_agent_model_runtime` builds the pre-agent fallback notice from
`(_load_cfg().get("model") or {}).get("provider")`. `load_user_config_effective`
keeps the `model: <id>` string shorthand as a string (every other reader —
gateway/run.py, cron/scheduler.py, hermes_cli/main.py — accepts it), so a
primary AuthError on such a config raised `'str' object has no attribute
'get'` inside the fallback branch: `setup.runtime_check` answered ok=False
("'str' object has no attribute 'get'") and `_make_agent` failed, on a
backend whose configured fallback chain was perfectly usable.
Repro: `tests/tui_gateway/test_pre_agent_fallback_notice.py::
test_fallback_build_survives_model_string_shorthand` (red on base).
Surfaced as the cross-file leak `test_protocol.py::test_config_roundtrip`
(writes `model: test/model` into a home left bound on `server._hermes_home`)
→ `test_tui_gateway_server.py::test_setup_runtime_check_agrees_with_
session_fallback_chain`.
|
||
|
|
d75d3fe15c |
Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
|
||
|
|
b066034eaf |
fix(desktop): preserve interim messages across history and tool boundaries (#107386)
* fix(desktop): preserve interim commentary across history and tool boundaries * fix(desktop): preserve interim commentary across history and tool boundaries * fix(desktop): keep public Codex commentary out of hydrated thinking Project profile-scoped, sanitized display commentary and reasoning for REST and gateway history without changing persisted replay items. Preserve canonical finals across stream recovery and keep tool-delimited interim text through settlement. --------- Co-authored-by: Xipong <217837358+Xipong@users.noreply.github.com> |
||
|
|
537ce77f52 |
fix(tui-gateway): exiting mid-tool no longer orphans the foreground command's process tree
Foreground terminal commands run in their own session (start_new_session) so an interrupt can kill the whole tree, which also puts them outside the host's process group. When the tui_gateway left mid-command (client closed stdin, or SIGTERM) nothing killed them: _shutdown_sessions closed the agents, the SIGTERM path hard-exits after a 1s grace, and the `bash -c` + child tree survived, reparented to init. - tools/environments/base.py: execute() records every in-flight foreground command; kill_live_foreground_processes() kills their trees through the backend's own _kill_process (the same kill an interrupt uses). - cleanup_all_environments() (the exit funnel of the CLI, one-shot, messaging gateway and terminal_tool's atexit, so `hermes serve` too) kills them first. - tui_gateway _shutdown_sessions (EOF atexit + SIGTERM handler) and the serve SIGTERM/SIGINT exit-flush handler interrupt running turns, wait up to 0.5s for them to settle so the tool call ends with a result the final persist records (no dangling tool_call in state.db), then kill any foreground command still alive. - ComputeHost.close() kills them too: every caller os._exit()s right after. Covers the case of b9dac83d366c (Desktop quit: serve SIGTERM handler) on every host. |
||
|
|
97de4b4ef8 |
Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups: - utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long double-quoted scalar is never folded after an escaped backslash. pm-clean builds every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there (ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml. - hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every read-only probe (source_git_env) now refuses promisor lazy fetches, and the partial-clone test targets that probe (red without the flag). - .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply. Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's --include-integration invocation. - apps/desktop: package.json has no build block here, so main's macOS locale-marker restore joins the darwin branch of the existing after-pack.mjs, and its test loads the hook from electron-builder.config.cjs and imports PlatformPackager from app-builder-lib's root (electron-builder 27 exports no ./out paths). The win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design. - reconciliation.ts: main's rowId hydration (#119326) was merged into the first of pm-clean's split helpers only; the resolver is now one helper both halves use. - en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted. - Tests main added with `import yaml` use hermes_yaml, like the rest of the tree. |
||
|
|
d956f0ae57 |
fix: key the tools[] pin by code version; never re-add config-excluded tools
Review follow-up on the byte-identical tools[] pin. - The pin records the code identity that built it (checkout/build sha, else the release version). Written by the same code, every pinned tool that is still available keeps its pinned bytes, including tools whose parameters are derived per surface (delegate_task, text_to_speech, memory, patch). The per-tool "parameters differ -> take current" rule replaced those bytes on every surface hop and rewrote the ~44KB pin each time. A pin from other code (`hermes update`, legacy name lists) takes the current definitions once and is re-pinned. - A pinned tool this process did not build is carried forward only while this agent's toolset selection allows it (enabled minus disabled toolsets and role reservations, before check_fn). It must also pass the session schema gates on the merged array, so browser_exec never comes back once terminal is gone. Client-surface toolsets (desktop_ui, project) still carry across hops: no config choice removed them there. - The rotation compaction child inherits the parent's pin in the publish transaction. - `hermes sessions recover` keeps pin rows in its system_prompts sweep and clears dangling pin hashes, as lost-and-found now does too. Profile moves carry the pin like the prompt. A continuing session whose pin is missing or unreadable (a row swept by an older build) pins the tools it sends on that turn, so later hops stay stable. |
||
|
|
01217c5fc2 |
fix(tui_gateway): withdraw an approval that ends before its settle hook attaches
_emit_approval_request sends the frame, then attaches the settle hook that withdraws it when the queue wait ends. A client answering by approval.respond (or another surface resolving it) inside that window ends the wait first; register_gateway_settle returns False and the request was left in open_requests, so every later session.resume replayed a prompt nobody was waiting on. Honour the False: withdraw at once. |
||
|
|
d9e10e0eb1 |
fix(tui): a secondary profile's key save and session model stay in that profile
One Desktop backend serves several profiles. Two session-bound paths read or wrote the LAUNCH profile instead of the session's: - model.save_key was not @_profile_scoped: a key saved from a secondary session (session_id) or for an explicit profile landed in the launch profile's .env, and the handler then exported it into the shared os.environ. It now binds the profile scope like model.options, and the explicit os.environ publish is gone (save_env_value already publishes to the bound scope, and to os.environ only for the launch profile). reconcile_record() already skips a non-launch home, so the profile-param guard around it is dropped. model.disconnect had the same gap (it removed the launch profile's credentials) and gets the same decorator. - session.create's info.model and the first state.db row resolved the default model from the launch profile's config. _session_default_model() resolves it under the session's own profile scope; the same launch-model fallback in the lazy resume info, the fallback session info, the live session identity and the branch row now use it too. Found by the two-tenant Desktop backend canary (tests/e2e/core/tenancy/test_two_tenant_desktop_backend.py): alpha's saved key appeared in default's .env, and alpha's session.create reported default's model. |
||
|
|
327dc2a721 |
refactor(config): move read-failure reporting into hermes_cli/config_read_errors.py
The parse-failure banner, the active-failure record, FailedConfigRead and the write refusals are one topic; config.py had grown past 4,100 lines with this PR. Pure move (no behaviour change): config.py drops to 3,999 lines, below origin/main. Callers outside config.py now import from the new module. |
||
|
|
db2f07c5cf |
fix(config): a transient read error no longer wipes config.yaml
Every fail-open config reader turned a read error on an intact file into
a stand-in ({} from read_raw_config and the TUI's _load_cfg_raw, defaults
or a possibly stale last-known-good from load_config). Writers then
mutated that stand-in and saved it; the writer re-reads the file, finds it
fine, and merges by deletion, so one EMFILE/EIO during a TUI config.set,
a dashboard save, a migration step or any load_config()->save_config()
caller replaced the whole file with the stand-in plus one key.
- Fallbacks from a failed read are now FailedConfigRead (a dict subclass
carrying the error). Readers are unchanged; save_config() and
atomic_config_write() refuse to persist one, so the error reaches the
caller and the file stays byte-identical. The subclass rides the
load->mutate->save round trip, so every writer is covered without
touching each of them.
- A read error's fallback is no longer cached (nor recorded as the next
last-known-good), so the retry reads the file again.
- PUT /api/config merges into a new dict, so it reads strictly instead.
- Gateway /verbose and /footer wrote the effective view back (fail-open
to {} on error, ${VAR} values expanded); they now round-trip the raw
file strictly.
|
||
|
|
c13ea774e6 |
refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0 on source installs, rewritten by release stamping) leaked v0.0.0 into About, /api/health, User-Agents, and plugin compat, and source updates showed "couldn't reach update server" because identity and channel authority disagreed with the checkout. Now: get_version_info() resolves install stamp -> live git -> unknown, never pyproject metadata, never a package constant. Source checkouts derive identity from their reachable release tag; the completion tail of every successful install/update/historical takeover atomically rewrites install-stamp.json with that identity; a stale source stamp whose commit no longer matches HEAD defers to live git. ACP/TUI use derived_version for display and base_version for protocol fields; all ~44 runtime __version__ consumers migrated; hermes_cli.__version__ and generated _version.py are gone; release stamping only touches the native manifests external builders consume (nix/tauri/cargo) and passes release identity straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0. Desktop no longer synthesizes a competing install-stamp.json: the checkout owns its stamp, and desktop-bootstrap classification keys on the bootstrap-complete marker. verify-bootstrap-version-stamp.py now cross-checks the checkout's stamp (baseVersion + commit == HEAD). Validation: 31-file focused suite green (version identity, stamping, adoption, providers, gateway, acp/tui runtime identity, api server via extras env, release graph); desktop tsc + 25 vitest green; real-repo probe: base=unknown derived=git.0635606.dirty source=git on this checkout; clean-env imports resolve entirely from this tree; windows footgun + compat-pointer scans clean. |
||
|
|
f4a38b1ee8 | Merge origin/main into ethie/pm-clean | ||
|
|
1af97a3f31 |
fix(tui_gateway): drop peer-less global broadcasts instead of printing them to serve stdout
_broadcast_global_event fell back to _emit -> write_json -> the module stdio
transport whenever no WS peer was registered. That fallback is only right for
the stdio TUI (tui_gateway.entry), whose stdout IS the JSON-RPC channel. In
`hermes serve` / the dashboard, JSON-RPC runs over WS and Desktop captures
stdout into desktop.log, so once the last WS client left, the change watcher
(sessions.changed every ~15-25s), the free-tier bootstrap (setup.ready, fired
before any client connects) and the orphan reaper (session.reclaimed) printed
raw frames there (~1100 `[hermes] {"jsonrpc": ...}` lines).
entry.main() now marks stdout as the RPC channel; everywhere else a peer-less
global broadcast is dropped with a debug log. Clients re-pull state on
connect (gateway.ready / setup.status), so nothing is lost.
|
||
|
|
9de3d9fdaa |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # tui_gateway/methods_session.py |
||
|
|
33a30fdd81 |
fix(tui_gateway): a warm session.resume reports the chat's own model, not the profile default
Reloading a Desktop session that was still live on the backend took the reattach path, whose `info` said `model: _resolve_model()` (the config default) and carried no provider. Desktop writes info.model/provider straight into the session state the picker renders, so a chat pinned to a plugin provider in the composer flipped to the profile default on reload, and flipped back once the backend had dropped the session and the eager resume restored the stored override. Reported as "the model changes randomly when reloading" with the Claude subscription plugin. `_live_session_identity(session)` answers with the precedence `_session_info` already uses for the eager path (queued switch, metadata mirror, built agent, the deferred record's composer override) and falls back to the profile default only when the chat never made a pick. The reattach `info` now carries `provider` too, like every other resume shape. |
||
|
|
4cb2054801 | Merge remote-tracking branch 'origin/main' into ethie/pm-clean | ||
|
|
07646a7f72 |
fix: a chat switched to Codex no longer sends its model to the previous provider's endpoint
The messaging gateway's /model wrote only model + provider to the session row. A row the Desktop/TUI had persisted with the Nous Portal route kept that base_url and api_mode, so resuming a chat switched to openai-codex built provider=openai-codex on the Portal URL and posted gpt-6-luna-900k to inference-api.nousresearch.com/v1/chat/completions: "Model 'gpt-6-luna-900k' isn't available on ChatGPT or Codex Subscription". - SessionDB.update_session_model writes the whole route (provider, base_url, api_mode) in both shapes resume reads (top-level for TUI/Desktop, gateway_runtime for the CLI) whenever a provider is given; the gateway /model passes the switch result's endpoint. - Resume readers (TUI/Desktop, CLI, gateway rehydrate) drop a persisted base_url that is another built-in provider's canonical endpoint, so rows already written by older builds heal. - A gateway session override whose credentials failed to re-resolve is resolved for its own provider on the turn instead of being layered over the default provider's runtime (which produced openai-codex + Nous key + Nous URL). If it still cannot resolve, that turn runs on the whole default route with the existing one-shot "Provider fallback" notice; the override is kept and retried next turn. |
||
|
|
3f4b8840f1 |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # pyproject.toml # tests/fixtures/resolution_allowlist.json |
||
|
|
7ed6534c7e | Merge origin/main: browser fence composed with the dispatch/retry split (#115184); server registration, i18n, contracts | ||
|
|
358c21c113 | Merge remote-tracking branch 'origin/main' into ethie/pm-clean | ||
|
|
3e00a356a4 |
fix(mcp): one reader for mcp_servers.<name>.enabled (#119567)
The `enabled` key had four parsers. The MCP client (`_parse_boolish`) read `enabled: 0` as on; the toolset resolver and editor (`_parse_enabled_flag`) read it as off. The server list (`summarize_server`, `/api/mcp/servers`) read any non-`False` value as on, so `enabled: "false"` showed on while the agent skipped it. The catalog and `hermes mcp list` accepted only true/1/yes, so `enabled: on` showed off while the server ran. `tools/mcp_tool_common.py::mcp_server_enabled` is now the only reader, and every surface calls it. `_parse_boolish` treats YAML numbers by truthiness (0 off, other numbers on). Everything else keeps the client's semantics: the off words are off, absent / null / junk stay on, with the existing warning for junk. The desktop MCP page mirrors the rule in `serverEnabled` (`apps/desktop/src/lib/mcp-servers.ts`). One case table (`mcp-enabled-cases.json`) drives the Python invariant test and the vitest test, so the page and the runtime cannot drift apart again. |
||
|
|
803165526b | Merge remote-tracking branch 'origin/main' into ethie/pm-clean | ||
|
|
4b803147d3 |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # apps/desktop/src/app/contrib/onboarding-kickoff.ts # hermes_cli/profiles.py |
||
|
|
333898b353 |
feat: setup profile sessions get the setup toolset by profile role; nothing else can grant it (#119491)
The setup toolset (empty until NS-964 registers request_catalog_install) is reserved for the profile whose backend-written profile.yaml carries role: setup. Two points enforce it: - Grant: tui_gateway/server.py::_load_enabled_toolsets folds the in-scope profile's role toolsets into all three return paths (configured CLI toolsets, coding posture, HERMES_TUI_TOOLSETS pin), next to the client-surface set. The pin keeps it too: an operator pin picks configurable toolsets and must not strip the profile's own. - Deny: model_tools._select_tool_names strips toolsets reserved for any other role from every selection, including None/"all", a saved platform_toolsets list, profiles.configure and the env pin. That is the one point every surface's selection passes, so a default session cannot get the tool by any route. The role is read with read_profile_meta on get_hermes_home(), which under a session's home override is that session's profile dir (no directory scan). setup stays out of _HERMES_CORE_TOOLS, CONFIGURABLE_TOOLSETS and the platform-native recovery loop (it has no tools, so the loop skips it). Linear NS-963. |
||
|
|
3c6c132366 |
Setup profile is minted by the backend and found by role, not by name (#119456)
* feat: setup profile is minted by the backend and found by role, not by name The guided onboarding runs in a profile the desktop used to create itself (profiles.create with a soul, "already exists" treated as success) and recognise by the literal "hermes-setup". The upcoming setup toolset grants catalog installs to that profile, so the marker that grants it must be written only by the backend. - profile.yaml carries `role: setup`; read_profile_meta / write_profile_meta / ProfileInfo know it; profiles.list and GET /api/profiles report it. - hermes_cli/setup_profile.py: ensure (find by role, adopt a pre-role hermes-setup dir, else clone default + soul + role) and reset (soul, memories, skills back to the created state, in place). The soul text moves here from the renderer. - tui_gateway/methods_onboarding.py: onboarding.ensure_setup_profile and onboarding.reset_setup_profile. Neither takes a name; profiles.create and profiles.configure already reject `role` (unknown key, 4000). - Copies never inherit the role: --clone-all, profile import, and a distribution that ships profile.yaml drop it. - setup.status for a named profile reports `ready` once the boot bootstrap settled. Since one host backend serves every profile (#118246) the desktop's setup-profile probe lands on this branch, which never set `ready`, and the kickoff waited forever. - Desktop: SETUP_PROFILE, ensureSetupProfile(profiles.create) and composeSetupSoul are gone. store/setup-profile.ts holds the name the backend returned (or the roster's role row after a relaunch); kickoff, handoff and the build card use it. The dev reset calls the reset RPC. * fix: write the setup soul as bytes so Windows keeps \n line endings * refactor(desktop): drop the renderer's setup-profile store; the backend is the only owner Kickoff reads the name straight from onboarding.ensure_setup_profile and records it on $setupSession, which every later step already carries. The handoff recovery check reads the roster row's role. No renderer module holds a setup-profile name or a fallback lookup. |
||
|
|
c13287c915 |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # apps/desktop/electron/main.ts # hermes_cli/backup.py # hermes_cli/config.py # hermes_cli/plugin_catalog.py # hermes_cli/plugins_cmd.py # hermes_cli/plugins_cmd_catalog.py # hermes_cli/plugins_discovery.py # hermes_cli/profiles.py # hermes_cli/update_cmd_deps.py # pyproject.toml # tests/gateway/test_dm_topics.py # tests/hermes_cli/test_config.py # tests/hermes_cli/test_plugins_cmd.py # tests/hermes_cli/test_update_autostash.py # tests/tools/test_lazy_deps.py # tools/lazy_deps.py # tools/skill_ledger.py # utils.py # website/docs/user-guide/security.md |
||
|
|
70f5dc5f46 |
feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
|
||
|
|
0706dffca1 |
fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to `utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the on-disk document so user comments, key order, quoting and blank lines survive every write. Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed (tui_gateway only) but nothing else used it, so each new writer regressed the class. - save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented example blocks are appended only when the file is created. - round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...) are force-quoted at every depth; duplicate keys are tolerated like PyYAML. - direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py, profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write. |
||
|
|
fe1e5c2638 |
tui_gateway: resolve the launch home at call time for launch-profile secrets
model_switch bound launch-profile turns to server._hermes_home, the home frozen at import, and read <that home>/.env through launch_secret_scope. The launch state.db handle already resolves the launch home at call time (#112692: the patched _hermes_home when a test changed it, else the live process home) precisely so a harness that re-homes the process after import is honoured; the secrets path did not, so under a developer shell with a custom HERMES_HOME every turn read the import-time home's .env — a guarded root for the suite — and test_compute_host_turn_protocol ran 2 failed + 1 error under a bare pytest. Share one _launch_home() between both paths. Restore that file's _wait timeout to main's 5 s: the 20 s bump only stretched the wait for a turn.end that the guard error had already made impossible. |
||
|
|
afc3b7c6f3 |
feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash merges ( |
||
|
|
265e68d784 |
fix(auth): one primary_failure_wording() helper labels quota vs auth failure at all three fallback surfaces
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482 sibling: 'Primary auth failed' for a 429 on the gateway surface). - Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the flag is only True when _ensure_runtime_credentials returned False). - Fold the new tests: 2 parametrized production-path tests + 1 gateway test. |
||
|
|
c68d889a35 |
fix(desktop): unmask the slash worker's real error behind the dispatch fallback
The Desktop only surfaced the underlying slash.exec failure when the command.dispatch fallback refused with "not a quick/plugin/skill command", but the gateway's refusal now reads "not a quick/plugin/bundle/skill command" (the bundle branch was added later). The stale regex never matched, so a launch-profile "slash worker start failed: ... UnscopedSecretError" was buried under the routing noise "not a quick/plugin/bundle/skill command: approvals" (#115427). Accept both wordings (older gateways still emit the short one) and drive the test with the current gateway wording. Also trims the spawn-site comment in tui_gateway/server.py to the WHY. |
||
|
|
4c2de106ca |
fix(tui_gateway): launch-profile slash worker spawns under own home after multiplex flip
Once the multiplexed gateway serves a second profile home (first foreign-profile request), _SlashWorker.__init__ spawns served_profile_child_env(target_home=None, inherit_credentials=True) for the launch profile's own worker, which fail-closes with UnscopedSecretError (no target + no bound scope while multiplex is on). Every worker-routed slash command then 5030s on the launch profile. Pass the launch home explicitly when multiplexing is active and no served home is set; the multiplex-off default is unchanged. Regression test: with _MULTIPLEX_ACTIVE, a profile_home=None worker spawns under its own home instead of raising. Fixes #115427 (Python spawn-site half; the Desktop TS unmask regex is a separate Electron-surface change). |
||
|
|
78b5dbfb08 |
fix(sessions): pin the Bot Chat composer-pick reach through _apply_model_switch -> SessionDB -> session.resume
Replace the direct `_stored_session_runtime_overrides` call test with one that drives the production path: `_apply_model_switch` on a follow_profile_config row under profile B writes the composer_override_profile marker into B's real state.db row (`_persist_live_session_runtime`), and `session.resume` on the deferred path restores the pin while B's config.yaml model is unchanged and drops it once the profile model moves. Launch home A carries a different model throughout, so the compare is proven to run under the session profile's scope (A->B->A). The test is red when the persist hunk in server.py or the record/_resume_eager wiring in methods_session.py is reverted; the first-write projection in session_workdir.py is pinned by the extended `_ensure_session_db_row` marker test (a pick before the first send). One `_row_follows_profile(row)` helper replaces the three copies of `follow_profile_config or title == "Bot Chat"`; identity is the persisted marker, the raw title compare stays only inside the helper as the legacy-row fallback. Docs: one sentence in the Bot Mode guide on the composer pick sticking to the chat until the Bot's profile model changes. |
||
|
|
a16d45828f | fix(sessions): preserve explicit Bot Chat model overrides |