Commit Graph

23263 Commits

Author SHA1 Message Date
Teknium
0e4552199a fix(status): keep fatal platform entries visible when gateway startup failed
Follow-up to #80451. /api/status cleared gateway_platforms whenever the
gateway process was down — correct for a clean stop (stale 'connected'
states are noise) but wrong for startup_failed, where the fatal entries
ARE the diagnosis: per-profile credential collisions and auth failures
(multiplex '<profile>:<platform>' keys) that the single exit_reason
string cannot express. #80451's writer-identity and freshness filters
already drop entries from other/older processes, so preserving
fatal-state entries here cannot leak another gateway's live state.

Live-validated shape: a real multiplex gateway (2 secondary profiles,
rejected tokens) persists telegram / alpha:telegram / beta:telegram
fatals in gateway_state.json; /api/status previously reported {} for
platforms while state was startup_failed.
2026-08-16 21:54:31 -07:00
Teknium
3b9a963b8e refactor(xai): lift API-key precedence into resolve_xai_http_credentials behind prefer_api_key
Rework of the #88049 inline early-return per review:

- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
  checks the explicit XAI_API_KEY first and falls back to OAuth. The key
  is read through tools.tool_backend_helpers.resolve_provider_secret
  (config -> profile secret scope -> env/.env -> credential pool) so the
  preferred path enforces the same scope policy as the existing fallback
  branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
  XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
  mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
  prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
  root cause for /v1/tts 403s (#87045, supersedes the inline shape in
  #87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
  the shared resolver; added coverage for the flag's OAuth fallback,
  HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
  profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
  wins (metered billing implication).
2026-08-16 21:20:29 -07:00
liuhao1024
32170dd1a2 fix(x_search): prefer an explicit API key over subscription OAuth
When a paid XAI_API_KEY is configured alongside xAI SuperGrok OAuth,
_resolve_xai_bearer() took the OAuth path unconditionally. The OAuth
credential authorizes /v1/responses but answers in a degraded Grok
explanatory mode with no citations, while the API key returns real
posts - so every x_search query silently degraded (#88040).

Prefer the explicit API key when set (same shape as the TTS fix for
#87045 in #87081), keeping OAuth as the fallback when no API key is
configured. The shared resolver and every other xAI call site are
unchanged.
2026-08-16 21:20:29 -07:00
hermes-seaeye[bot]
1826310f49 fmt(js): npm run fix on merge (#88128)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-17 04:19:24 +00:00
Teknium
55db6187e5 chore(contributors): map skip-agent's commit email for attribution 2026-08-16 21:17:47 -07:00
Teknium
b2cc4ceef0 fix(cron): attribute cloud-path refusals to the cloud-synced script, not a lifecycle command
When check_gateway_lifecycle refuses a cron script that lives on a
FileProvider path, the generic error implied the job contained a dangerous
gateway lifecycle command. Surface the real reason instead — the script
lives on a cloud-synced path whose evicted placeholder could hang the
preflight scan — while staying fail-closed. Regression test asserts the
cron-script scan path blocks without opening the file and that the message
names the cloud-synced path rather than a lifecycle command.
2026-08-16 21:17:47 -07:00
Teknium
42a96e7543 fix(cron): move cloud-placeholder refusal into _read_referenced_script and cover ~/Library/CloudStorage
Widen #88052 per review:
- The walk-level short-circuit only protected _contains_unsafe_gateway_action;
  the sibling caller _read_script_for_scanning still opened cloud-resident cron
  scripts and could hang preflight. Move the check into _read_referenced_script,
  the shared choke point, so every caller fails closed without opening.
- Generalize _is_apple_file_provider_path -> _is_cloud_placeholder_path: detect
  ~/Library/CloudStorage (Dropbox/OneDrive/Google Drive third-party FileProvider
  domains) alongside iCloud's Library/Mobile Documents.
- Regression tests: CloudStorage lexical path blocked without open; the choke
  point itself refuses cloud paths with os.open forbidden.
2026-08-16 21:17:47 -07:00
Agent
e6f59e5b77 fix(terminal): avoid FileProvider reads in lifecycle guard 2026-08-16 21:17:47 -07:00
Teknium
f8f43c9523 fix(desktop): make git worktrees work end-to-end on a remote gateway backend
Cmd/Ctrl+Shift+B worktree flows on a remote gateway route through the
backend's /api/git mirror (hermes_cli/web_git.py), but that mirror had
drifted behind the Electron-local git ops the same UI drives locally, so
the flows broke exactly and only on remote connections:

- Convert-a-branch: the picker offers remote-tracking refs, and the
  Electron op turns "origin/feature" into a local tracking branch. The
  mirror ran `git worktree add <dir> origin/feature` verbatim, which
  either fails or detaches HEAD. It now resolves the ref's remote via
  git (never assuming "origin"), fetches best-effort, and creates the
  worktree with `--track -b <short-name>`.
- branch_list omitted remote-tracking refs entirely and never set the
  `isRemote` flag the renderer's HermesGitBranch contract requires —
  the convert picker on a remote gateway couldn't reach a teammate's
  branch and mislabeled every row's action.
- Branching off an `origin/…` base silently wired the new branch to the
  remote upstream; the mirror now passes `--no-track` like the Electron
  op does.

Renderer side, replace the silent degradation with a capability gate:
when a remote backend predates the /api/git worktree routes, worktree
creation failed with an opaque "Expected JSON … got HTML" toast. The
route-missing shapes now surface a clear "update the Hermes backend"
message (isGitEndpointMissingError, mirroring the sidebar batch-endpoint
detector); real git errors still pass through untouched.

Sibling audit (documented, no code change needed): repo status / review /
file-diff / git-root / default-cwd already route through desktopGit()'s
REST bridge or /api/fs on remote; repo scan is deliberately a no-op there.
Stale comments claiming "empty/false on a remote backend" in projects.ts
and coding-status.ts updated to describe the backend-routed reality.

Fixes #81724
2026-08-16 20:53:44 -07:00
Teknium
a36583e311 fix(desktop): keep cloud bot avatar eye catchlights inside the eyes
The white catchlight dots in BotFace were static circles pinned at the
circle-face eye line (cy 16.5), while the animation clock moves the
pupils to the shape-aware eye line (cy 22 for the cloud). On the cloud
avatar the highlights floated above the eyes instead of inside them.

- Tag the catchlights (data-hb-hl-l/r) and move them with the pupils in
  paintMathFace, offset upper-left of each pupil center.
- Render the initial eyes/catchlights/shut-lids at the shape-aware eye
  line so the first frame matches the animated frames.
2026-08-16 20:51:58 -07:00
Teknium
97b5f66030 docs(state): soften stale SessionDB self-pin wording after #88063
#88048 documented the token-writer self-pin (bound-method thread target +
strong atexit hook) as a permanent contract: "__del__ never runs for
exactly the instances that leak". #88063 then removed both pins (idle
writer retirement + weakref atexit hook), making abandoned handles
eventually collectible.

Reword the __enter__ docstring and the context-manager test module
docstring to describe the pin as historical motivation, note the #88063
behavior, and keep the guidance that owners close deterministically.
No code changes.
2026-08-16 20:43:13 -07:00
Shannon Sands
d66341ab28 fix(status): strict writer-identity ownership for aggregated platform entries (OOF-3)
The freshness window (updated_at >= live process create_time - 2s) had a
P1 boundary hole: a stale failure written by the PREVIOUS process
immediately before a fast restart landed inside the slack and was
aggregated; if that platform was then removed, the new process never
replaces the entry and NAS stays degraded indefinitely.

Replace clock heuristics with persisted writer identity:

- write_runtime_status now stamps every platform entry with the writing
  process's (writer_pid, writer_start_time) — the same PID-reuse
  fingerprint the liveness checks use, so a recycled PID never
  masquerades as the original writer.
- The aggregation ownership filter requires exact equality between an
  entry's stamp and the profile's validated live gateway process
  (get_runtime_status_running_pid + _get_process_start_time). No slack,
  no timestamps. Legacy entries without a stamp fail closed.
- Writer stamps are process recon (same class as the auth-gated
  gateway_pid) and are stripped from all /api/status projections, both
  active-profile and merged cross-profile entries.

Near-boundary regression test: prior-process entry stamped 100ms before
restart is excluded; recycled-pid-different-fingerprint excluded;
legacy no-stamp excluded; current-process entry kept.
2026-08-16 20:32:09 -07:00
Shannon Sands
57279cf2b9 fix(status): freshness-filter aggregated per-profile platform entries (OOF-3)
Gateway startup deliberately preserves plain platform entries in
gateway_state.json across restarts, and the active-profile endpoint
compensates by filtering against current configuration. The cross-profile
aggregation copied raw maps, so a fatal entry for a platform the operator
had since disabled/removed could keep NAS reporting the instance degraded
indefinitely.

The aggregation has no cheap per-profile config context (platform sets
depend on tokens in each profile's .env behind its secret scope), so use
freshness instead: an entry is aggregatable only when its updated_at is
at/after the live gateway process's create time (validated PID via
get_runtime_status_running_pid + psutil create_time; the record's own
start_time field is a PID-reuse fingerprint in clock ticks, not a
timestamp). Config changes require a restart to take effect, so
restart-anchored freshness is exactly the config filter's semantics.
Fail closed: unparseable timestamps or no live process exclude the entry
— a false 'degraded forever' is the worse failure mode.
2026-08-16 20:32:09 -07:00
Shannon Sands
1d46a9fe0f fix(status): aggregate independent per-profile gateway failures; harden key filter (OOF-3)
- /api/status now folds LIVE independent per-profile gateways' platform
  failures (gateway_mode == 'multiple', the OOF-3 deployment mode) into
  gateway_platforms under the validated <profile>:<platform> grammar, so
  NAS fleet health sees them without a schema change. ?profile= requests
  stay unmerged (single-profile view).
- Namespaced-key validation no longer fails open: colon-containing keys
  are grammar-checked even when configured-platform loading throws.
- Platform key segment now accepts hyphens, matching plugin platform IDs
  (plugins/platforms/<dir> names, e.g. foo-bar).
2026-08-16 20:32:09 -07:00
Shannon Sands
f755ed5e90 fix(gateway): surface multiplex profile failures (OOF-3) 2026-08-16 20:32:09 -07:00
Shannon Sands
c127d0c3b1 fix(gateway): attribute scoped credential lock conflicts to the owning profile (OOF-3)
Scoped credential locks (Telegram bot token, Discord bot token, etc.) are
machine-global, but the conflict error only reported the holder's PID:

    Telegram bot token already in use (PID 559). Stop the other gateway first.

On multi-profile hosts (e.g. hosted instances running 13 profiles), a bare
PID gives the operator no way to tell WHICH profile owns the credential —
the exact failure mode observed on zerocool-9781, where the 'default'
profile was misconfigured with the same bot token as 'lead-gen-outreach'
and logged an unattributable conflict every ~5 minutes (4,602 rows).

Fix:
- acquire_scoped_lock() now stamps a 'profile' label on lock records,
  inferred from the process HERMES_HOME (<root>/profiles/<name> layouts,
  'default' for the root home). Omitted when not inferable.
- New scoped_lock_owner_label() resolves the owning profile from a lock
  record: prefers the explicit field, falls back to inferring from the
  persisted hermes_home for locks written before the field existed.
  Labels are validated against the profile-id grammar before use (lock
  files are plain JSON on disk and the label flows into log lines and a
  suggested CLI command).
- _acquire_platform_lock() conflict message now names the owning profile
  and gives the correct remedy:

    Telegram bot token already in use by the 'lead-gen-outreach' profile
    gateway (PID 559). Stop that gateway first
    (hermes --profile lead-gen-outreach gateway stop).

  Records with no attribution signal keep the original PID-only wording.

Testing:
- New TestScopedLockOwnerLabel suite covering label inference (named,
  Docker, root/default, unknown layouts), grammar validation, explicit-
  field preference, hermes_home fallback, and legacy/malformed records.
- acquire_scoped_lock tests for profile stamping and omission.
- Adapter-level tests for profile-attributed, legacy-home-inferred, and
  PID-only conflict messages.
- 76/76 targeted gateway tests pass; broad gateway suite failures are
  baseline-identical (verified via git stash comparison). Ruff clean.
2026-08-16 20:32:09 -07:00
webtecnica
e3c71e052d feat(delegation): record model/provider in live-transcript manifest (#telemetry) 2026-08-16 20:14:51 -07:00
Teknium
2bef9e8f51 chore: map contributor email for tigercraft4 (PR #74468 salvage) 2026-08-16 19:58:49 -07:00
Teknium
b711fd0513 feat(desktop): carry remote gateway headers through the connections registry, test probes, and Settings UI
Completes PR #74468 (remote gateway headers for Cloudflare Access, #74466)
against the v2 multi-connection registry that landed after the PR was
authored, and closes the review blockers:

- connection-registry: additive optional `headers` field on remote/cloud
  entries (normalized through the same forbidden-name filter, secret
  envelopes like `token`); inherited on edit, treated as dial material by
  connectionDialFieldsChanged, preserved by normalizeRegistry, and carried
  through migrateV1ToRegistry. v2 registries without the field load
  unchanged — no version bump.
- main.ts registry paths: connectRegistryBackend dials with the entry's
  headers (readiness probe, ticket mint, descriptor REST via
  getJsonForBackend/fetchJsonForBackend, registry ws-url minting with
  rememberRemoteWsHeaders so renderer upgrades get them injected).
- saveRegistryConnection encrypts incoming plaintext header values with the
  same safeStorage/allowPlainText seam as tokens; sanitizeRegistryConnection
  exposes only header NAMES to the renderer — values never cross IPC.
- Connection tests exercise the leg they validate: both
  hermes:connection-config:test and hermes:connections:test now send the
  configured headers on the HTTP status call, the ws-ticket mint, AND the
  live WebSocket probe (probeGatewayWebSocket grew an injectable `headers`
  option passed as the undici WebSocket constructor's second argument).
- Settings → Connections gains an "Extra gateway headers" editor for
  remote/cloud entries (name + secret value rows, stored values shown as
  saved-but-hidden, clearable), with i18n keys (en + zh; other locales fall
  back through defineLocale).
2026-08-16 19:58:49 -07:00
tigercraft4
fcef62ef72 feat(desktop): support remote gateway headers 2026-08-16 19:58:49 -07:00
Teknium
2b7f49673f fix(desktop): self-heal dropped SSH/HTTP registered remote connections
A dropped registered remote connection (SSH or HTTP) never recovered on
its own: the next boot attempt failed with a transient transport error
("Could not verify the existing SSH backend", ERR_CONNECTION_RESET,
mint timeout), the failure was correctly NOT latched, but nothing ever
re-attempted the boot — the renderer's reconnect machinery only arms
after a completed boot. The app parked on "Desktop boot failed" until
the user manually deleted and re-entered the same connection details,
which merely forced the fresh bootstrap an automatic retry would have
performed (issue 82679, feature ask 80430).

Root causes and fixes:

- electron/backend-start-failure.ts: new isRetryableRemoteBootFailure()
  predicate — a remote, non-reauth boot failure is transient and may be
  retried; local failures and confirmed 401/403 rejections are not
  (a missing capability differs from a transient failure).
- electron/main.ts: the boot-failure progress broadcast now carries
  `retryable` (rides with `error` through updateBootProgress), and a
  failed reuse probe against a cached SSH master tears the stale
  master/tunnel down so the next attempt bootstraps fresh — exactly
  what manual re-entry did.
- use-gateway-boot.ts: bounded self-heal loop for a failed boot whose
  progress is marked retryable — up to 5 re-attempts with the same
  full-jitter backoff as the socket reconnect loop (2s base, 15s cap).
  Exhausted retries end in the real boot-failure recovery overlay,
  never an infinite spinner. Reset on success and on soft switch;
  timer cleared on unmount.
- store/boot.ts: resumeDesktopBootForRetry() re-arms the overlay with a
  retry status while an automatic retry is in flight.

Secondaries already had full-jitter backoff (store/gateway.ts); this
closes the same class for the PRIMARY/registered-connection path.

Tests: predicate matrix (retryable vs reauth-latch mutually exclusive),
plus renderer hook tests proving a transient SSH failure self-heals on
the next attempt, retries are bounded (6 total dials then the recovery
overlay, no further attempts), and non-retryable failures never enter
the loop. Sabotage-verified (disabling either half fails 4 tests).

Fixes #82679
Fixes #80430
2026-08-16 19:58:37 -07:00
fangliquanflq
ace85d63de feat(desktop): add status bar reconnect for offline gateways
Expose the existing profile-aware gateway boot reconnect path through a
single-flight renderer action, and surface a Reconnect button in the
gateway status menu panel whenever the socket is not open. Repeated
clicks share one in-flight reconnect; failures surface through the
existing non-destructive notification UI. Localized copy for all
supported Desktop locales.

Salvaged from PR #80694 (net diff re-applied onto current main; panel
code lives in app/shell/gateway-menu-panel.tsx now).
2026-08-16 19:58:37 -07:00
chelsealong
c0f89d2545 fix(gateway): scope slash.exec's skill-command check to the session's profile
Independent review of the prior commit found the cache-invalidation key
alone doesn't fix the reported #88023 dead path: slash.exec runs as a
_LONG_HANDLER on the pool with a copied context, and no binding of
_HERMES_HOME_OVERRIDE happens between the transport read and the handler
body, so get_skill_commands() there always fell back to the process-level
HERMES_HOME regardless of which profile's session issued the request.

Bind the session's own profile_home around the get_skill_commands() check,
mirroring the same bind/reset-in-finally pattern already used at every
other per-turn HERMES_HOME scoping site (e.g. server.py's prompt-turn and
system-prompt-rebuild paths). This makes the #88023 dead path actually
reachable by the fix instead of only exercising the cache primitive in
isolation.
2026-08-16 19:55:03 -07:00
chelsealong
a9b4ec3126 fix(skills): rescan skill commands cache when active profile changes
Switching Desktop profiles mid-session changes HERMES_HOME but not the
platform scope, so get_skill_commands() kept serving the previous
profile's skill list. A skill only available under the new profile then
looked like a cache miss to callers such as slash.exec, which fall
through to the slash_worker dead path (#88023).
2026-08-16 19:55:03 -07:00
liuhao1024
8dc427949f fix(gateway): accept CJK full-width punctuation as MEDIA path terminators
MEDIA_TAG_CLEANUP_RE (and MEDIA_EXTENSIONLESS_TAG_RE) only recognized
ASCII terminators after a MEDIA:<path> tag. Chinese-language agent
output naturally writes MEDIA:D:\...\zhibao.pdf(782.6 KB)or ...pdf:内容 —
the full-width punctuation failed the trailing lookahead and the
attachment was silently dropped (cron even reported 'delivered') (#88038).

Both lookaheads now accept a CJK full-width terminator set (()〈〉《》:,。;
!?、curly quotes【】) alongside the ASCII set. The #68773 adjacent-tag
splitting guard is covered by a regression test.
2026-08-16 19:54:42 -07:00
Shawn Wang
a40324bb32 fix(gateway): ignore invalid managed Node directories
Signed-off-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
2026-08-16 19:54:37 -07:00
fangliquan
59c7a99084 fix(state): avoid overlapping context manager change 2026-08-16 19:53:07 -07:00
fangliquan
b454e4da76 fix(state): release abandoned session database handles 2026-08-16 19:53:07 -07:00
Jack Lau
ab7f48d4b0 feat(state): support the context-manager protocol on SessionDB
A SessionDB handle cannot be released by dropping the last reference.
Once its background token writer starts, the instance pins ITSELF two
ways: the writer thread's target is a bound method, and
queue_token_counts registers atexit.register(_drain_token_queue_at_exit),
which only close() unregisters. A dropped-but-pinned handle keeps its
state.db/-wal/-shm descriptors for the life of the process, and __del__
never runs for it, so the existing safety net is dead code for exactly
the instances that leak.

That is why owning call sites are expected to close explicitly, in those
words, in the ownership comments in run_agent.py and
tui_gateway/methods_session.py. This adds the ergonomic half of that
contract so an owner can scope a handle and be exception-safe by
construction:

    with SessionDB(path) as db:
        db.append_message(...)

Purely additive. __enter__ returns self, __exit__ closes and returns
False so a caller's exception always propagates, and close() is already
idempotent, so a scope that closes early still exits cleanly. Nothing
changes for callers that already close directly.

Four regressions cover the scope closing the handle, __enter__ returning
the instance itself, the failure path closing while still propagating,
and an early close leaving the exit clean. They assert on the
sqlite_safe_read tracking registry rather than raw descriptor counts,
matching test_session_db_read_conn_pool.py, because SQLite's unix VFS
parks a closed descriptor on a per-inode reuse list and makes raw counts
lag the real connection count.

Refs #88033
2026-08-16 19:52:21 -07:00
Shannon Sands
4822156923 fix(cron): stop retry storms when the gateway is deliberately stopped (OOF-266)
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).

Split the unreachable path by durable operator intent:

- desired_state == "stopped" (written only by the s6 lifecycle commands;
  the same intent signal container-boot reconciliation trusts) -> drop
  the fire with 200 + a structured log line, mirroring NAS's own
  instance_stopped drop. Jobs are not lost: the Chronos provider
  reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
  file without desired_state) -> keep the retryable 503, now stamped
  with Retry-After: 60 so a scheduler that honors it spaces retries
  past the wake/restart window instead of exhausting them inside it.
  The gateway's own pass-through 503s (draining) get the same hint.

The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
2026-08-16 19:50:07 -07:00
hermes-seaeye[bot]
f0ab10455a fmt(js): npm run fix on merge (#88079)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-17 02:34:46 +00:00
Teknium
b82b94811a chore: map contributor email for attribution audit 2026-08-16 19:28:56 -07:00
Jesse Panganiban
307d46c207 fix(desktop): map SSH profile aliases in REST paths 2026-08-16 19:28:56 -07:00
Teknium
bab7be3ca7 feat: raise Codex OAuth context to 900K for gpt-5.6 family and gpt-5.4 (subscription 1M rollout)
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.

gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
2026-08-16 18:31:54 -07:00
Teknium
8236b41771 feat: sync bundled Bot Mode with multi-source roster (Hermes-Bot-Mode#68)
Pulls the multi-source roster into the bundled plugin: profiles.list rows
from the active gateway are merged with the host.agents() union roster
(hermes-agent #86875), so the Bots panel shows agents from every registered
Desktop connection with @name-device handles for duplicates. Feature-detected
and best-effort — an older Desktop build or roster failure leaves the
single-source list untouched.

Adapted for the bundle:
- useRoster queryFn combines the bot_mode_protocol capability read (which
  landed after #68 was cut) with the multi-source merge
- multi-source-roster tests updated for the namespace SDK import harness
- soul-protocol-backfill anchor widened for the new botHandle(name, bot)
  signature

Plugin suite: 143/143.
2026-08-16 18:30:53 -07:00
Teknium
9829064f89 fix: capability fingerprint reads config via the canonical loader
The config-read guard (test_config_read_guard) correctly flagged the
probe's raw yaml.safe_load of config.yaml — raw reads miss the managed
overlay, env expansion, and normalization. Use load_config_readonly()
under a scoped HERMES_HOME override instead. E2E v3/v3b and the guard
both green.
2026-08-16 18:30:53 -07:00
Teknium
4e22d070f4 feat(agent): one-time protocol upgrade for legacy Bot Chat sessions
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.

E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
2026-08-16 18:30:53 -07:00
Teknium
ea4310e76c feat(agent): capability-refresh + timeless prompts for eternal Bot Chat sessions
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.

- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
  capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
  installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
  epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
  epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
  cache so new installs appear), persisted so the next turn reuses the
  new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
  session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
  "Conversation started:" date is dropped (timezone kept); no ticking
  fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
  agent (tool definitions are construction-baked) when the fingerprint
  moves, same session id/history, with a user-visible notice

Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).

Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
2026-08-16 18:30:53 -07:00
Teknium
72bae0e91a fix(agent): Bot Chat gate reads a session-title hint before the DB
Live desktop E2E caught a write-ordering bug the automated E2E missed:
tui_gateway applies pending_title to state.db AFTER the first turn, but
the system prompt builds at turn START — the DB-title gate saw nothing
and the Bot Chat was cached protocol-less forever. The gateway now
hands the agent its intended title at construction and the gate checks
the hint first, DB second (CLI/messaging-gateway paths unchanged).

Live-verified on the running desktop: fresh bot's Bot Chat persisted
with the protocol section, handle, and roster in its system prompt;
regular sessions and SOUL.md untouched.
2026-08-16 18:30:53 -07:00
Teknium
400400e1d9 fix(hermes-bots): composeSoul honors the bot_mode_protocol capability
Found in live desktop E2E: the generated-identity path of composeSoul
still appended the protocol section even when the backend injects it
into the system prompt. New agents now get a clean identity-only SOUL
against capable backends; older gateways keep the append. Covered in
the capability-suppression test.
2026-08-16 18:30:53 -07:00
Teknium
78a4693eef fix(agent): scope the Bot Mode protocol section to canonical Bot Chat sessions
Per review: the protocol belongs only in official Bot Mode interactions,
not every session on a managed install. The prompt builder now injects
the section only when the agent's session row is titled "Bot Chat"
(BOT_CHAT_TITLE, matching the desktop's createCanonicalChat pin and the
`hermes -p <bot> chat -c "Bot Chat"` resume target). Regular sessions
never carry it; the desktop composer middleware owns @mention sends.

Title is read once at first prompt build and the rendered prompt is
cached + DB-restored — cache-safe. E2E against the real AIAgent +
SessionDB: absent in an untitled session, present in Bot Chat,
byte-stable across rebuilds, absent after retitle, absent with the
flag off. Overhead unchanged (~916B, Bot Chat sessions only).
2026-08-16 18:30:53 -07:00
Teknium
d516496cef fix: track bundled plugin.js sources past the tsc-artifact gitignore
apps/desktop/src/**/*.js is gitignored (stale tsc output shadows .tsx),
which silently dropped the hermes-bots plugin.js from the adoption
commit — tests shipped, source didn't, CI ENOENT'd. Negate the pattern
for src/plugins/*/plugin.js: adopted plain-ESM plugins have no .tsx
sibling, so the shadow hazard cannot apply.
2026-08-16 18:30:53 -07:00
Teknium
2b39e92e9e feat(agent): core Bot Mode teammate protocol — stable-tier prompt section
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.

- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
  (process, home), keyed off the agent's OWN home (not ambient
  HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
  agent.bot_mode_protocol (default True), stable tier, byte-stable
  across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
  the bundled plugin gates ALL SOUL protocol writes on it (backfill,
  composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere

Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
2026-08-16 18:30:53 -07:00
Teknium
366d8814b8 feat(desktop): bundle Bot Mode (hermes-bots) as a built-in, default-on plugin
Adopts the Hermes-Bot-Mode desktop plugin (NousResearch/Hermes-Bot-Mode)
into apps/desktop/src/plugins/hermes-bots/, registered by the bundled
vite glob and ON by default. It stays a pure @hermes/plugin-sdk consumer
in plain-ESM plugin.js form; users disable it live in Settings > Plugins.

- contrib/plugins.ts: bundled glob accepts plugin.js entries
- contrib/runtime-loader.ts: a disk/runtime copy of an id that ships
  bundled is skipped (standalone installs predating adoption cannot
  double-register)
- package.json: check:test:plugins runs the plugin's node:test suite in
  CI (138 tests)
- source: Hermes-Bot-Mode @ c19baba, incl. today's #107/#103/#99 merges
2026-08-16 18:30:53 -07:00
hermes-seaeye[bot]
3c108589fb fmt(js): npm run fix on merge (#88016)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-16 23:43:53 +00:00
Brooklyn Nicholson
d8f40d4176 test(desktop): steer suite drives the real redirectPrompt path; hydration fixture carries durable row shape
The live suite previously called appendMidTurnUserMessage directly, leaving
redirectPrompt's appendAfterActiveReply guard — the production decision of
WHERE a correction lands — outside the harness. Both hooks now mount together
sharing one state map, exactly as the desktop wires them, so a regression in
the caller (not just the insert) goes red. Verified by mutation: disabling the
guard fails 2/4.

Also covers the rejected-redirect path: a not_running response discards the
optimistic bubble instead of stranding a correction the model never saw.

The hydration fixture now carries the durable row shape the client actually
receives (row_id, reasoning, provider call_id/response_item_id on tool_calls)
instead of a hand-simplified echo, so the 'mirrors real state.db rows' claim
is honest. The fake-timer steer id counter is gone with the local insert —
ids come from redirectPrompt itself.
2026-08-16 18:38:34 -05:00
Brooklyn Nicholson
f1f694aac1 test(desktop): harden steer-order suite against fake-timer id collisions
Review follow-ups: steer ids now come from a monotonic counter instead of
Date.now() (frozen under fake timers — two steers without a clock advance
would have collided), and the settle-above assertion documents its
load-bearing sealed-bubble assumption.
2026-08-16 18:38:34 -05:00
Brooklyn Nicholson
b2f345f540 test(desktop): pin steered-turn transcript order end-to-end
A steered turn's contract — pre-steer output above the correction bubble,
post-steer output and the settled reply below it — was fixed across
several PRs (#73793/#83151 class, settle fixes) but only covered piecewise:
the mid-turn insert as a unit, the settle math as a unit. Nothing drove the
real stream reducer through a whole steered turn, and nothing asserted the
durable-row hydration renders the same order after reload.

Two suites close that:
- steer-arrival-order: full event sequences through useMessageStream's real
  handler + the real optimistic insert — single steer with tool activity,
  steer racing message.complete, double steer in one turn.
- steered-turn-hydration-order: toChatMessages over persisted row shapes
  copied from a real state.db steered turn, including a tool result that
  lands after the correction row.
2026-08-16 18:38:34 -05:00
hermes-seaeye[bot]
d378a25e85 fmt(js): npm run fix on merge (#88014)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-16 23:33:14 +00:00
Teknium
4ee87f76ee fix(desktop): read cron run-history from the owning gateway
When Hermes Desktop works against a REGISTERED gateway connection, cron
jobs execute on that gateway and persist their run sessions in the
gateway's state.db. But every REST call in the app — the cron surface
included — carried only `profile`, so `hermes:api` routed it through the
local profile pool and `_list_cron_job_runs_sync` read a local state.db
with zero `source='cron'` rows. Every job showed "No runs yet" while the
same endpoint on the gateway returned the real runs (#87882).

Fix at the routing seam:

- HermesApiRequest gains an optional `connectionId`. The renderer's cron
  helpers (list/get/runs/delivery-targets/create/update/pause/resume/
  trigger/delete/blueprints) now tag the active registry connection via a
  new connectionScoped() twin of profileScoped(), fed from the same
  setApiRequestConnection seam store/gateway already maintains for the
  plugin socket.
- The hermes:api main-process handler resolves a tagged request through
  ensureRegistryBackend — the SAME pool the job list and WS traffic use —
  instead of the legacy profile route. Shared remote/cloud hosts (one
  gateway, many profiles) get the path scoped with ?profile= via the new
  pathWithProfileScope helper, factored out of pathWithGlobalRemoteProfile.
- '' / 'local' / absent connectionId keep the byte-identical v1 route, so
  single-source and connection-config-remote users are unaffected.

This covers the run-history panel, the sidebar cron peek, and every other
cron surface in one place, since they all funnel through the same helpers.

Fixes #87882
2026-08-16 16:27:16 -07:00