A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
_render_peers_map_view called _sibling_resolutions per account, and each call
ran _all_profile_host_configs(), which parses every profile's config.yaml via
list_profiles() and re-reads its honcho.json; show() re-runs after every
workspace switch. cmd_peers_map now scans once and threads the rows through.
_seen_gateway_accounts turned a locked or corrupt state.db into the same [] that
means "no gateway traffic yet"; it now notes the error on stderr before returning
the empty list. The missing-table fallback for old schemas is unchanged.
Iterating a honcho SyncPage walks every following page, so _api_workspace_peers
pulled the whole workspace on its first call, never hit the `< 50` break, re-walked
pages 2..N and returned duplicates; the 200-peer cap that exists so a public bot's
workspace cannot stall the CLI was defeated and the `p<N>` picks pointed at the
wrong peers. It now reads `.items` per page and stops at the cap; `_api_workspaces`
reads `.items` too.
The wizard's `new_host` probe looked only at the host block, so an install that
keeps peerName/enabled/workspace at the root with no hosts.hermes block read as
fresh and Enter defaulted to pinning every account onto one peer. The probe
includes the root.
`_seen_gateway_accounts` dropped rows whose origin said `is_bot`, but
SessionSource.to_dict never serializes that field, so the filter was dead; removed
with its test row. `_sanitize_peer_id` was a copy of session_peers.sanitize_peer_id.
A confirmed workspace repoint is persisted, so the exit line no longer says
"Nothing changed" after one.
The identity step treated any host block without a mapping key as a new install and defaulted the choice to single peer. An install with enabled, workspace and peerName then had Enter write pinUserPeer: true and merge every gateway account onto the operator. cmd_setup now decides new-ness before its prompts populate the block, and only a block with none of the mapping, peerName, workspace or enabled keys defaults to single.
The docstring said grouping session rows by (source, user_id) enumerates every account the gateway handled. record_gateway_session_peer overwrites a row's user_id, so a shared thread keeps only its last author. The docstring now says that, and a test pins it.
clone_honcho_for_profile copied sessionPeerPrefix but not sessionAiPeerPrefix. A new profile cloned from a default block with the AI prefix on fell back to unprefixed session names and collided with the default profile's gateway sessions.
peers map showed sanitize(prefix + user_id) for prefixed accounts. The runtime appends a sha256 suffix when sanitizing changed the id or it collides with an explicit peer, so the preview named a peer the gateway never writes to. The preview now builds a HonchoSessionManager with no client and asks it.
Clearing the last alias in peers map popped userPeerAliases from the host block, and the host then inherited the root aliases again. The host block now keeps an empty map, which both config readers treat as an explicit override. The summary says when that empty map hides root aliases.
_api_workspace_peers returns peer IDs instead of dicts. The created date
it carried was never printed. Callers index the list directly.
_preview_peer_resolution uses _sanitize_peer_id instead of a local copy.
_sibling_resolutions strips the display suffix itself; its one caller did.
_seen_gateway_accounts opens the database under contextlib.closing and
parses origin_json once into a dict. _save_alias_map picks the target
block first and writes it once. cmd_peers_map renders through one local
show() and computes the previous resolution on one path for both account
numbers and typed runtime IDs.
peers map read seen accounts from state.db but required a non-empty
session_key. every telegram row on a long-lived install written before the
gateway stamped that column was dropped, so the picker showed no accounts
while the same rows grouped fine by (source, user_id). the key was never
used by the grouping.
The picker-cap comment repeated the constant's name; the
sessionAiPeerPrefix field comment restated the field name and its
symmetry with sessionPeerPrefix before reaching the collision it
prevents. Both now state only the why.
Extends the read-only 'hermes honcho peers' view with a 'map' action
that joins two sources: workspace peers fetched from the Honcho API,
labeled from local config (your peer, each profile's AI peer, alias
targets, runtime peers of seen accounts, user-* fallback peers,
'unrecognized' otherwise), and the gateway accounts recorded in
state.db with what each currently resolves to. Targets are picked
from the workspace list so a typo cannot silently create a peer;
every assignment states its consequence (aliases move future
messages only; a runtime peer left behind keeps its history).
'w' lists every workspace the key can reach — the wrong-workspace
fallback — and can repoint the profile's workspace on explicit
confirmation. With multiple profiles, saving a root-cascading map
asks whether to write root or fork the host block, root writes warn
when sibling profiles sit on other workspaces, and the accounts
table marks siblings that resolve an account differently. Offline
the command degrades to typed targets over the local account list.
The setup wizard's gateway step closes by pointing at the command.
The step declares its scope up front: human mapping only, with each
Hermes profile bringing its own AI peer. A note explains aliases as
the join between platform accounts and named peers. Each shape now
says when it fits. Fresh configs default the choice to [1] single
peer — the common personal setup — instead of [3], which silently
fragmented a solo operator's gateway account away from their
peerName history. Configured setups keep their detected shape as
the default.
force_refresh_token gated the adopt-from-disk paths behind the failure cooldown.
After one of our exchanges failed transiently, a 401 within the next 30s returned
None even when a sibling process had already rotated and written a valid grant,
so the operation raised HonchoAuthError with a good token sitting on disk.
Adopting is a disk read, not an exchange; the cooldown exists to stop replaying a
single-use refresh token, so the two adopt checks now run before the gates.
_write_config parsed honcho.json twice under the lock and only wrapped the first
read into ConfigWriteRefused; _refuse_unparseable now returns the parsed dict and
the branches use it. The getattr/isinstance duck-typing collapses to one guard.
cli._read_config reuses oauth's tolerant reader (BOM-tolerant, like the strict
reader the write side uses) instead of its own utf-8 copy.
The CLI's _write_config took only the best-effort file lock, so an in-process
refresh thread could still interleave with a command's read-modify-write. The
dashboard's Honcho save (_write_provider_honcho) still seeded its whole-file
rewrite from a tolerant reader, so a honcho.json that exists but does not parse
was replaced by the active host's block alone from the UI - the same bug class
this PR closes on the CLI and refresh paths. `hermes profile create --clone`
swallowed the new ConfigWriteRefused as "plugin not installed".
One parametrized test covers the web writer for the corrupt and parseable cases.
_write_config applied a command's edits relative to the snapshot the read took, but never moved that snapshot after a successful write. A second write on the same object therefore compared A -> B -> A against A, saw no change, and left disk at B. After a write the snapshot and path now follow the caller's dict, so the next write applies only the edits made since.
_read_config() falls back to ~/.honcho/config.json or the default profile when honcho.json does not exist, so cfg.path never matched the write path and _write_config() wrote the whole dict. That overwrote a refresh rotation a serve child had written onto the honcho.json the setup login created moments earlier. When the local file exists at write time, the seed snapshot is now overlaid with the local file and only the command's edits are applied onto it.
install_grant writes the login grant to disk before the wizard's later prompts. _apply_grant_to_host wrote it only into the live cfg, so _write_config saw the grant as an edit and copied it over whatever a serve child rotated onto disk meanwhile. The snapshot now takes the grant too, so the final save leaves the newer on-disk grant alone.
_ReadConfig builds its snapshot in __init__; _read_config no longer
assigns the attributes after the fact. cmd_enable prints its hint
inline, since _no_credential_hint had one caller. The split
force_refresh_token signature and call fit on one line. Docstrings and
comments keep the what and the one non-obvious why. No behavior
changes; _apply_edits keeps its original body.
clone and enable accepted HONCHO_API_KEY from the environment as proof the
block could authenticate and wrote enabled: true. the variable can be absent
from the next process, leaving an enabled block with nothing behind it, which
is the cohort the previous commit set out to remove.
_resolve_api_key takes env=False at both write sites, so a block is enabled
only when honcho.json itself holds a key, an oauth grant, or a self-hosted
baseUrl. status and setup keep the environment fallback for display.
every hermes honcho command read honcho.json, changed a field, and wrote the
whole dict back with no lock. a token refresh in another process that landed
between the read and the write was overwritten with the old access and
refresh tokens, and the next refresh replayed a rotated single-use token.
_write_config now holds the same cross-process file lock the refresh path
holds, re-reads disk under it, and applies only the keys the command changed
since its _read_config(). untouched keys keep their on-disk value, so a
rotation survives; a credential the command set on purpose still wins. a
write with no prior read keeps today's whole-file behavior.
a named profile cloned from a default profile that signed in with oauth
got a host block with enabled: true and nothing to authenticate with.
hosts.hermes holds the grant, its apiKey is not inherited (#66125), and
copying the oauth block would make two blocks replay one single-use
refresh token. 'hermes honcho enable' on a fresh profile wrote the same
shape. status then showed the profile as enabled while every honcho
call ran without memory and the plugin quietly stayed inactive.
clone_honcho_for_profile and cmd_enable now resolve a credential for the
target block (its own apiKey, the root apiKey, HONCHO_API_KEY, or a base
url) before writing enabled: true. _resolve_api_key takes the block to
check so both share one definition of "can authenticate".
cohorts:
clone from an oauth default block, no root key: block written without
enabled; the client's auto-enable rule turns it on once a credential
appears (setup apikey writes the root key, or a per-profile login)
clone from a default block with a host-level static key only: same
clone with a root apiKey, an env key, or a base url: enabled as before
enable on an empty or fresh block with no credential: refused, one
message names the profile's setup command and the hosts.<name> key,
nothing is written
legacy blocks already on disk as enabled with no credential: nothing
rewrites them; the plugin already treats them as unusable and stays
inactive; enable now prints the same message instead of "already
enabled"
env-only key: counted as a credential at write time, as the client
does at run time; if the variable later disappears the client still
refuses to initialize the block
after the token endpoint revoked a grant (invalid_grant), running
'hermes honcho setup', choosing apikey and pasting a valid key changed
nothing. the wizard wrote the key to the root apiKey only. hosts.<name>
still held the dead access token under apiKey and the grant under oauth,
and the host block wins the lookup, so status kept reporting the revoked
grant and every call kept failing (#97990).
the apikey branch now drops the host's oauth block and writes the chosen
key onto the host block as well as the root. a dead access token is no
longer shown as the current key, so a blank answer with no other key
aborts instead of keeping the grant. a static host key without a grant
is kept as before.
a truncated or hand-edited honcho.json reads as {} on the tolerant read
path. the next write then replaced the file with only the current host
block: an oauth refresh, a login, or any 'hermes honcho' command that
saves a setting wiped every other host and the root keys.
_read_config_strict now raises on a parse error the same way it raises
on a read error, and logs one sentence naming the file. _rotate_and_persist
treats that as a refresh failure and enters the cooldown without spending
the refresh token. install_grant raises into the setup flow. in cli.py
every write goes through _write_config, which runs the same check first
and raises ConfigWriteRefused; the honcho router, the setup wizard and the
profile sync print the sentence instead of a traceback. the wizard checks
before asking its questions. read paths keep the {} fallback.
follows #95860, which added the strict reader for unreadable files and
kept a .corrupt copy on a parse error. leaving the original file in place
keeps the same bytes without a second copy of the tokens on disk.
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.
Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.
holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.
openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.
Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
_all_profile_host_configs() built per-profile host keys inline as
f"{HOST}.{profile}" ("hermes.work") while profile_host_key() — used by
honcho status/enable/sync and the runtime memory plugin — produces the
underscore form ("hermes_work"). The lookup always missed, so
'hermes honcho peers' showed "(not set)" / leaked the raw malformed key
into the AI-peer column for every non-default profile. Profile names
needing sanitization (dots/spaces) were doubly broken.
Verified live: with hosts["hermes_work"] populated, cmd_peers showed
'work ... hermes.work' before the fix and 'work ... hermes' after.
Tests: host keys match the writer form, sanitized profile names resolve,
peers output shows populated identities with no key leak, and clean
fallback for profiles without a block.
_authed_call checks the dead-grant marker before calling, retries a
confirmed auth failure once after a forced refresh, and records the
failure for the one-time notice. Operations re-resolve their peer and
session objects inside the call, so a retry after a client rebuild no
longer reuses objects bound to the old transport. Tool handlers now
return an explicit auth error instead of an empty result, and non-auth
failures keep their fail-open behavior.
Installing a memory provider (Honcho, mem0, hindsight, ...) from the
dashboard Plugins page failed on hosted deployments with a permission
error: the setup endpoint shelled out to
`uv pip install --python sys.executable`, which targets the sealed
read-only venv under /opt/hermes (immutable hosted image, NS-579/#49113).
The correct mechanism already exists: tools/lazy_deps.py redirects
installs to the writable durable target on the data volume
(HERMES_LAZY_INSTALL_TARGET=/opt/data/lazy-packages) when the venv is
sealed (HERMES_DISABLE_LAZY_INSTALLS=1), appends the target to the END
of sys.path (core venv always wins collisions), and constrains shared
deps to core-venv versions. The dashboard installer simply never used
it.
Fix:
- tools/lazy_deps.py: new public install_specs() — installs arbitrary
manifest-declared pip specs through the same environment routing as
ensure(): venv-scoped by default, durable-target on sealed images,
refused with an actionable reason when gated off (config kill switch
or sealed venv without a target — never surfaces raw EROFS/EACCES).
Specs are validated with _spec_is_safe(); post-install it invalidates
import/metadata caches so availability rechecks in the same process
see the new packages without a restart. Never raises.
- hermes_cli/web_server.py: _install_memory_provider_pip_dependencies
now calls install_specs() instead of building its own uv/pip
subprocess. Blocked installs surface the gate reason in the setup
results; the response's status block reflects post-install
availability (stale 'missing deps' state clears immediately).
- hermes_cli/memory_setup.py, plugins/memory/honcho/cli.py,
plugins/memory/mem0/_setup.py: CLI setup wizards routed through
install_specs() too — same sealed-venv failure mode, same fix.
No hosted setup path writes to /opt/hermes anymore; provider discovery
and installation now use the same environment (sys.path activation is
shared with the lazy-install bootstrap in hermes_bootstrap).
Tests:
- tests/tools/test_lazy_deps.py: TestInstallSpecs — gating matrix
(sealed+no-target blocked with immutable-deployment reason, config
kill switch, sealed+target proceeds), spec-safety rejection before
any subprocess, venv-scoped vs --target command display, failure
stderr passthrough, never-raises contract.
- tests/hermes_cli/test_web_server.py: setup endpoint routes pip
through lazy_deps (regression guard asserts no direct 'pip install'
subprocess), blocked-reason surfacing, same-response availability
recheck clears stale missing state.
Fixes NS-605 (Plain T-1111).
Adds a device authorization grant flow alongside the existing loopback
OAuth flow, so `hermes setup` can connect to Honcho cloud from SSH and
other no-browser environments.
- oauth.py: new HTTP seams — _http_post_form_status (non-raising, since
RFC 8628 polling reads the OAuth error off a 400) and _http_get_json
for the RFC 8414 metadata probe
- oauth_flow.py: DeviceCode, request_device_code, poll_for_token with
slow_down backoff (+5s, capped at 60s) bounded by expires_in, typed
errors (AccessDenied, DeviceCodeExpired, AuthorizationTimeout), and
supports_device_login (fail-closed metadata gate); device flow ends in
the same install_grant tail as loopback so refresh/status work
unchanged
- oauth_flow.py: loopback callback now serves a "sign-in was not
completed" page on consent cancel instead of the success page
- cli.py: cloud menu offers oauth / device / apikey; the device option
only appears when the host advertises the grant, and becomes the
default when no browser is detected
- 18 new tests covering the full flow against a local fake AS, backoff
schedule, error mapping, deadline bound, metadata gate, and wizard
branches
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-ups for the consolidated salvage:
- Memoize the config.yaml-derived timeout on the file's mtime_ns so the
rebuild-on-timeout-change check from PR #57437 costs one stat() per
get_honcho_client() call instead of a full YAML load on the hot path.
- hermes honcho status now displays the host-block-resolved
dialecticCadence (remnant from PR #63776, whose runtime fix landed
in #62290).
- AUTHOR_MAP entries for the salvaged contributor emails.
Replace the hand-rolled ensurepip bootstrap (and five other one-off
pip-install code paths) with hermes_cli.tools_config._pip_install, which
prefers the bundled uv (fast, needs no pip in the venv), falls back to
python -m pip, and bootstraps pip via ensurepip only when missing.
Sites unified:
- hermes_cli/setup.py: _install_neutts_deps, _install_kittentts_deps,
modal SDK install, daytona SDK install
- hermes_cli/memory_setup.py: memory-plugin pip deps (previously dead-ended
when uv AND pip binaries were both absent)
- hermes_cli/dingtalk_auth.py: qrcode auto-install (previously invoked
'python -m uv' which is not how uv ships)
- agent/lsp/install.py: --target LSP server installs
- plugins/google_meet/cli.py, plugins/platforms/matrix/adapter.py,
plugins/platforms/google_chat/oauth.py, plugins/memory/honcho/cli.py
Tests updated to assert the ladder behavior (uv-first, pip fallback,
ensurepip bootstrap) instead of the removed bespoke branches.
ruff check --fix --select F541 . on current main. Pure prefix removals;
adjacent-string concatenations keep the f only on interpolating fragments.
No string content or live placeholder altered.
* feat(memory): OAuth token storage and refresh for the Honcho provider
* feat(memory): refresh the Honcho OAuth token in the client and session
* feat(memory): zero-CLI loopback OAuth authorization flow
* feat(memory): generic memory-provider OAuth connect endpoints
* feat(desktop): memory-provider OAuth connect link
* feat(memory): CLI OAuth sign-in with source-tagged authorize links
* fix(memory): IP-literal loopback redirect and consent config_path on the authorize link
* fix(memory): profile-scope the memory-provider OAuth endpoints
* refactor(desktop): generic memory-provider OAuth client functions
* docs(memory): trim OAuth module docstrings to the invariants
* docs(memory): document OAuth connect as an optional auth method
* fix(memory): send home-relative display path to consent, not the absolute path
* perf(memory): cache OAuth token expiry in memory to skip the hot-path disk read
* fix(memory): log OAuth refresh failures at warning, not debug
* feat(memory): fall back to an OS-assigned loopback port when 8765 is taken
* test(memory): cover the desktop Connect launcher, status, and provider dispatch
* fix(desktop): keep the memory-provider dropdown one size regardless of connect state
* fix(desktop): move the memory connect link to the description line, leaving the dropdown untouched
* refactor(memory): move OAuth connect routes out of web_server into a memory-layer router
* refactor(desktop): import MemoryConnect directly, drop the single-export barrel
* fix(memory): launch CLI OAuth sign-in right after the auth choice, not after the wizard
* fix(desktop): auto-clear the OAuth error state instead of leaving it sticky
* test(honcho): isolate auth-method prompt from deployment-shape wizard tests
main's wizard suite scripts the cloud prompts without the OAuth auth-method step; auto-answer it in the shared helper so the answer lists stay shape-only.
* docs(honcho): document query-adaptive reasoning level (reasoningHeuristic)
README never mentioned reasoningHeuristic and listed reasoningLevelCap as an orphaned cap with the wrong default (— vs "high"). Add the query-adaptive scaling note + the reasoningHeuristic/reasoningLevelCap rows (grouped under Dialectic & Reasoning), matching the wording already on the hosted honcho.md page, and add a pointer from the memory-providers overview.
* fix(honcho): default the CLI peer prompt to the OAuth consent name
The CLI runs the grant with apply_config=False, so the peerName the user just entered at consent was dropped and the wizard's 'Your name' prompt fell back to $USER. Surface it as a transient OAuthCredential.consent_peer_name (set even when config isn't merged) and seed the prompt default from it.
* feat(honcho): split OAuth client_id by surface (cli=hermes-agent, desktop=hermes-desktop)
resolve_endpoints now picks the client_id from the initiating surface and
threads it through authorize -> token exchange -> persisted grant -> refresh,
so the CLI and desktop register as distinct OAuth clients. Surface-specific
env overrides (HONCHO_OAUTH_CLIENT_ID_CLI/_DESKTOP) win over the generic
HONCHO_OAUTH_CLIENT_ID, which still overrides every surface.
* feat(honcho): show OAuth vs API key in status; detect existing OAuth in setup
status now prints 'Auth: OAuth (clientId, token valid Xm/expired)' instead of
masking the OAuth access token as a generic API key; setup notes an existing
OAuth grant when re-run.
* docs(honcho): drop 'shared pool' wording from unified observation mode help
* fix(honcho): cross-process lock around OAuth refresh to prevent grant revocation
The in-process threading lock can't stop a sibling process (another profile or
the desktop app sharing honcho.json) from replaying the single-use refresh
token and tripping reuse-detection, which revokes the whole grant. Guard the
read-refresh-persist section with an OS file lock on <config>.lock so only one
process rotates at a time; the others re-read the freshly-persisted token.
Best-effort: platforms without flock degrade to in-process serialization.
* refactor(honcho): one OAuth client (hermes-agent) for all surfaces
Collapse the per-surface client_id split. CLI and desktop now use a single
client_id (hermes-agent); consent branding/UI still adapt via the source query
param. One grant identity means no clientId-vs-refresh-token desync that could
get the grant revoked. HONCHO_OAUTH_CLIENT_ID still overrides for self-hosting.
* fix(honcho): per-session resolves to session_id, never remapped by title
Reorder resolve_session_name so stable identifiers win over labels: gateway
per-chat key first, then the per-session session_id, then the cwd map / title.
A (possibly auto-generated) title can no longer remap a live per-session
conversation onto a second Honcho session mid-stream — fixes the desktop, which
is per-conversation via session_id. Consequence: a gateway's per-chat key now
also wins over a title (titles never remap a stable id).
'everyone collapses to your peer' read as a promise about all traffic.
pinUserPeer pins the user-side peer and is checked before userPeerAliases
(session.py:335), so a pin overrides every alias — including agent peers.
For a multi-agent operator that silently pools distinct agents onto one
peer, the opposite of intent.
Scopes the wording to 'every non-agent gateway user', notes the pin
overrides aliases, and points agent-mesh operators at pinUserPeer:false +
userPeerAliases instead. Same correction in the wizard menu/echo text,
the plugin README, and the website Honcho page.
The single/multi/hybrid 'deployment shape' was a misnomer: these keys only
affect the gateway (the one entrypoint supplying a runtime user ID), and the
three preset names stamped a lossy taxonomy onto three orthogonal knobs while
hiding which keys got written.
Replace it with an intent-led tree gated on gateway detection:
- _gateway_platforms() lazily inspects the gateway config (best-effort, no
hard dependency); the step auto-skips when no platform is connected.
- 'who talks to this?' → just me / me+others (pooled?) / only others, deriving
pinUserPeer + userPeerAliases + runtimePeerPrefix and echoing the result.
- [e] drops to a raw-knob editor for power users.
- The single→multi orphan guard survives as a pooling steer.
The setup wizard wrote the legacy pinPeerName even though pinUserPeer is
the canonical key that outranks it in the resolver — so it had to scrub
the canonical key afterward to stop it winning. Write pinUserPeer directly
and migrate any legacy pinPeerName onto it on touch (setup load + clone),
which removes the precedence-fighting entirely.
Resolver still reads pinPeerName as a back-compat alias; that's deferred.
When Hermes runs in TUI mode, the gateway child process communicates with
the Node.js parent over a JSON-RPC protocol on stdin. Subprocess calls that
inherit this stdin fd can trigger a race condition where the child's stdin
read returns EOF, causing the gateway to exit cleanly (exit code 0) mid-tool-
execution.
This is the same root cause as issue #14036 (byterover plugin) and PR #39257
(SSH environment backend). This commit applies the fix — stdin=subprocess.DEVNULL
— to all 85 subprocess.run() and subprocess.Popen() calls that execute inside
the TUI gateway child process.
Scope: TUI-context code only (agent/, tools/, plugins/, tui_gateway/server.py).
CLI code (cli.py, hermes_cli/), tests, scripts, and gateway process management
are excluded — they don't run inside the TUI child and inherit the terminal's
stdin, not the JSON-RPC pipe.
85 call sites across 28 files. All files pass syntax check.
hermes doctor and hermes honcho status warned 'Honcho config not found'
whenever ~/.honcho/config.json was absent, even though HONCHO_API_KEY in
.env resolves a working config via HonchoClientConfig.from_global_config()
-> from_env(). Both now check hcfg.api_key/base_url before warning.
Co-authored-by: oxngon <98992931+oxngon@users.noreply.github.com>
Self-hosted Honcho setup had four sharp edges:
- local/cloud URLs ending in /vN double-prefixed by the SDK (/v3/v3/... 404)
- authenticated local servers had no setup prompt for a JWT/bearer token
- profile-derived host keys could be dot-containing workspace IDs Honcho rejects
- memory-provider config files with API keys written world-readable per umask
This keeps existing behavior but makes those paths safer:
- strip a trailing /vN version segment from any configured baseUrl before SDK
init (the SDK's route builders always prepend their own version prefix);
auth-skipping stays loopback-only
- add an optional local JWT/bearer prompt in honcho setup, stored under
hosts.<host>.apiKey
- derive new profile host keys with underscores, still reading legacy
hermes.<profile> blocks
- write memory-provider config files atomically with 0600 via a shared
utils.atomic_json_write(mode=) arg (honcho/hindsight/mem0/supermemory)
- skip honcho.json parsing in gateway cache-busting unless Honcho is the active
memory provider; memoize by honcho.json mtime when active
- bust the gateway agent cache on memory.provider change
- add a hermes memory setup <provider> one-liner so fresh installs can configure
a named provider without the picker (the per-provider hermes <provider>
subcommand only registers once that provider is active)
Closes#20688, #29885, #26459, #30246, #33382, #32244.
Co-authored-by: BROCCOLO1D
Three related regressions stemming from the pinUserPeer alias landing:
- Setup wizard read host-only fields when detecting current shape but the
parser supports root-level config and gives host pinUserPeer higher
precedence than pinPeerName. Re-running setup could mis-detect shape
and silently flip routing. Detection now uses the same resolver order
as HonchoClientConfig, and each shape branch scrubs every peer-mapping
key before writing so a stale pinUserPeer=false can't outrank a freshly
written pinPeerName=true. Multi no longer auto-writes
userPeerAliases={} (was silently masking root-level baselines).
- clone_honcho_for_profile inherited pinPeerName but not pinUserPeer, so
a default profile configured with the newer key produced cloned
profiles without the pin.
- Gateway cache-busting signature fingerprinted Honcho user-peer fields
but not ai_peer. Since HonchoSessionManager freezes cfg.ai_peer at
init, mid-flight aiPeer edits kept assistant writes on the old peer
until an unrelated cache eviction. ai_peer is now part of the
signature.