Commit Graph

554 Commits

Author SHA1 Message Date
John Paul Soliva
2bade5daa7 fix(memory/hindsight): pin the isolation-vs-shaping split for scoped reads
`langfuse._secret` and `azure_identity_adapter._scoped_env` were changed to
raise rather than fall back, because swallowing `UnscopedSecretError` hides the
spawn-site bug the exception exists to surface. `_scoped_setting` looked like it
contradicted that, so make the split explicit and pin it.

Hindsight already follows the contract for everything that decides WHERE data
goes: `mode`, `apiKey` and the `bankId` partition read through bare
`get_secret`, so a scopeless multiplexed read raises. In `_load_config` that
raise happens on `HINDSIGHT_MODE` before any shaping value is reached, so the
swallow below cannot mask an isolation failure.

Presentation shaping is deliberately not in that class. `MemoryManager._each_provider`
logs an `initialize` failure at WARNING and drops the provider for the session,
so raising there would cost the whole memory provider because a speaker prefix
could not be resolved. It degrades to the provider's own default instead —
never to `os.environ`, which under multiplex is the default profile's.

The test names the offending key rather than asserting that something raised:
routing `mode` through the shaping helper shifts the failure to
`HINDSIGHT_API_KEY`, which a bare `pytest.raises` would still accept.
2026-09-13 14:41:26 -07:00
John Paul Soliva
a48dde5316 fix(memory/hindsight): resolve retain shaping through the profile scope, never os.environ
_load_config() reads the Hindsight bank, mode and retain tags through the profile
secret scope, but _apply_retain_settings() then discarded that answer and re-read
os.environ whenever the config value was falsy:

    return cfg.get(key) or os.environ.get(env_var, default)

Under gateway.multiplex_profiles os.environ holds the DEFAULT profile's .env, so a
secondary profile's scoped miss came back as the default profile's retain tags,
observation scopes, source and speaker prefixes — the fallback-after-miss shape
gateway/AGENTS.md forbids. Tags are Hindsight's retrieval partition and
metadata.source is opt-in by design, so the secondary's memories were both
mislabelled and selectable by the default profile's tag filters.

Both halves now go through _scoped_setting(), which resolves the value with
get_secret() and falls back to the provider's OWN default — a miss is a miss, the
same rule embedded.py already applies to the daemon's key and base URL. The three
raw reads left inside _load_config() (retain_source, retain_user_prefix,
retain_assistant_prefix), directly under the comment declaring them per-profile,
go through it too.

Single-profile deployments are unchanged: with no scope installed get_secret()
still reads the process env, where the value IS this profile's own.

Fixes #108865
2026-09-13 14:41:26 -07:00
kshitijk4poor
476dbfed3a perf(honcho): peers map scans the profile directory once, not once per account row
_render_peers_map_view called _sibling_resolutions per account, and each call
ran _all_profile_host_configs(), which parses every profile's config.yaml via
list_profiles() and re-reads its honcho.json; show() re-runs after every
workspace switch. cmd_peers_map now scans once and threads the rows through.

_seen_gateway_accounts turned a locked or corrupt state.db into the same [] that
means "no gateway traffic yet"; it now notes the error on stderr before returning
the empty list. The missing-table fallback for old schemas is unchanged.
2026-09-13 19:05:39 +05:30
kshitijk4poor
71f3516df2 fix(honcho): peers map reads one SDK page at a time; legacy root-level installs keep their shape
Iterating a honcho SyncPage walks every following page, so _api_workspace_peers
pulled the whole workspace on its first call, never hit the `< 50` break, re-walked
pages 2..N and returned duplicates; the 200-peer cap that exists so a public bot's
workspace cannot stall the CLI was defeated and the `p<N>` picks pointed at the
wrong peers. It now reads `.items` per page and stops at the cap; `_api_workspaces`
reads `.items` too.

The wizard's `new_host` probe looked only at the host block, so an install that
keeps peerName/enabled/workspace at the root with no hosts.hermes block read as
fresh and Enter defaulted to pinning every account onto one peer. The probe
includes the root.

`_seen_gateway_accounts` dropped rows whose origin said `is_bot`, but
SessionSource.to_dict never serializes that field, so the filter was dead; removed
with its test row. `_sanitize_peer_id` was a copy of session_peers.sanitize_peer_id.
A confirmed workspace repoint is persisted, so the exit line no longer says
"Nothing changed" after one.
2026-09-13 19:05:39 +05:30
Erosika
e05f6d82ce fix(honcho): default the wizard to the detected shape on every existing install
The identity step treated any host block without a mapping key as a new install and defaulted the choice to single peer. An install with enabled, workspace and peerName then had Enter write pinUserPeer: true and merge every gateway account onto the operator. cmd_setup now decides new-ness before its prompts populate the block, and only a block with none of the mapping, peerName, workspace or enabled keys defaults to single.
2026-09-13 19:05:39 +05:30
Erosika
7dacc2ed35 fix(honcho): describe what _seen_gateway_accounts can list
The docstring said grouping session rows by (source, user_id) enumerates every account the gateway handled. record_gateway_session_peer overwrites a row's user_id, so a shared thread keeps only its last author. The docstring now says that, and a test pins it.
2026-09-13 19:05:39 +05:30
Erosika
0ede483479 fix(honcho): clone sessionAiPeerPrefix into new profile host blocks
clone_honcho_for_profile copied sessionPeerPrefix but not sessionAiPeerPrefix. A new profile cloned from a default block with the AI prefix on fell back to unprefixed session names and collided with the default profile's gateway sessions.
2026-09-13 19:05:39 +05:30
Erosika
e467c0ef2c fix(honcho): preview peer resolution through the runtime resolver
peers map showed sanitize(prefix + user_id) for prefixed accounts. The runtime appends a sha256 suffix when sanitizing changed the id or it collides with an explicit peer, so the preview named a peer the gateway never writes to. The preview now builds a HonchoSessionManager with no client and asks it.
2026-09-13 19:05:39 +05:30
Erosika
841b768cd7 fix(honcho): keep an empty host alias map when the last alias is cleared
Clearing the last alias in peers map popped userPeerAliases from the host block, and the host then inherited the root aliases again. The host block now keeps an empty map, which both config readers treat as an explicit override. The summary says when that empty map hides root aliases.
2026-09-13 19:05:39 +05:30
Erosika
7e10f06a5a refactor(honcho): shorten the peers-map helpers
_api_workspace_peers returns peer IDs instead of dicts. The created date
it carried was never printed. Callers index the list directly.

_preview_peer_resolution uses _sanitize_peer_id instead of a local copy.
_sibling_resolutions strips the display suffix itself; its one caller did.

_seen_gateway_accounts opens the database under contextlib.closing and
parses origin_json once into a dict. _save_alias_map picks the target
block first and writes it once. cmd_peers_map renders through one local
show() and computes the previous resolution on one path for both account
numbers and typed runtime IDs.
2026-09-13 19:05:39 +05:30
Erosika
506dcacd9f fix(honcho): list gateway accounts whose session rows predate session_key
peers map read seen accounts from state.db but required a non-empty
session_key. every telegram row on a long-lived install written before the
gateway stamped that column was dropped, so the picker showed no accounts
while the same rows grouped fine by (source, user_id). the key was never
used by the grouping.
2026-09-13 19:05:39 +05:30
Erosika
783d4bdffc docs(honcho): prune wizard and peers-map comments
The picker-cap comment repeated the constant's name; the
sessionAiPeerPrefix field comment restated the field name and its
symmetry with sessionPeerPrefix before reaching the collision it
prevents. Both now state only the why.
2026-09-13 19:05:39 +05:30
Erosika
dc7c8673d1 feat(honcho): add 'hermes honcho peers map' for interactive account-to-peer mapping
Extends the read-only 'hermes honcho peers' view with a 'map' action
that joins two sources: workspace peers fetched from the Honcho API,
labeled from local config (your peer, each profile's AI peer, alias
targets, runtime peers of seen accounts, user-* fallback peers,
'unrecognized' otherwise), and the gateway accounts recorded in
state.db with what each currently resolves to. Targets are picked
from the workspace list so a typo cannot silently create a peer;
every assignment states its consequence (aliases move future
messages only; a runtime peer left behind keeps its history).

'w' lists every workspace the key can reach — the wrong-workspace
fallback — and can repoint the profile's workspace on explicit
confirmation. With multiple profiles, saving a root-cascading map
asks whether to write root or fork the host block, root writes warn
when sibling profiles sit on other workspaces, and the accounts
table marks siblings that resolve an account differently. Offline
the command degrades to typed targets over the local account list.
The setup wizard's gateway step closes by pointing at the command.
2026-09-13 19:05:39 +05:30
Eli Robbins
55dfb517e4 fix(honcho): bust gateway agent cache on session-prefix flips
Address review feedback on #39130: the new sessionAiPeerPrefix setting
affects the resolved Honcho session key, which HonchoMemoryProvider freezes
at construction (self._session_key). Because it wasn't part of the gateway's
cached-agent signature, a live config flip left an existing gateway session
bound to its old, AI-peer-agnostic Honcho session until an unrelated eviction
or restart.

Add honcho.session_ai_peer_prefix to _HONCHO_CACHE_BUSTING_KEYS and the
_extract_honcho_cache_busting_config values so a flip rebuilds the cached
agent on the next turn, mirroring the existing aiPeer / pin_peer_name /
runtime_peer_prefix contracts.

Also close the symmetric gap for the pre-existing user-side sessionPeerPrefix:
it feeds the same resolve_session_name output (per-session/title/per-repo/
per-directory strategies) into the same frozen _session_key, so it had the
identical live-flip staleness bug and was likewise absent from the cache
signature. Fixing both keeps the two prefixes consistent.

Add one config-flip regression test covering both keys, alongside the
existing Honcho cache-signature test.
2026-09-13 19:05:39 +05:30
Eli Robbins
0a7c17159c feat(honcho): add sessionAiPeerPrefix to isolate sessions per AI peer
The gateway_session_key branch of resolve_session_name() returns an
AI-peer-agnostic name, so multiple AI peers sharing one workspace +
peerName + gateway chat key collide on a single Honcho session.

Add sessionAiPeerPrefix (symmetric counterpart to sessionPeerPrefix):
when set, the resolved session name is prefixed with {ai_peer}- on every
resolution path. The prefixed name is re-run through the session-id
length cap so the prefix can never exceed Honcho's limit.

- config field + host/root parsing in client.py
- public resolve_session_name() wraps a new _resolve_session_name_base()
- tests covering parsing, the gateway-key case, cross-peer disjointness,
  the length cap, and a disabled-by-default regression guard
- README: config table + resolution notes
2026-09-13 19:05:39 +05:30
Erosika
21180ae7e8 feat(honcho): rework the setup wizard's gateway mapping step around Honcho's peer model
The step declares its scope up front: human mapping only, with each
Hermes profile bringing its own AI peer. A note explains aliases as
the join between platform accounts and named peers. Each shape now
says when it fits. Fresh configs default the choice to [1] single
peer — the common personal setup — instead of [3], which silently
fragmented a solo operator's gateway account away from their
peerName history. Configured setups keep their detected shape as
the default.
2026-09-13 19:05:39 +05:30
kshitijk4poor
d14734df1f fix(honcho): thread registry rides agent.memory_provider.spawn_context_thread
main folded the plugin's contextvars-inheriting thread wrapper into
agent.memory_provider.spawn_context_thread. The registry this stack adds for
join_plugin_threads keeps a thin client.spawn_context_thread that calls the core
spawner and records the thread under its owner; the four spawn sites import it
from client. main's inherit-profile test builds a real provider now that
_spawn_write is an instance method (the owner is what gets joined).
2026-09-13 19:05:39 +05:30
kshitijk4poor
f1273ed704 refactor(honcho): one reclaim-key helper, one cache-size constant, no dead writer starter
_SESSION_CACHE_MAX_SIZE was assigned twice with two comments describing one
constant. _retain_for_retry and _keep_until_flushed shared the has-unsynced /
current-owner / reclaim body and differed only in what to do when a newer object
owns the key; _reclaim_key_locked returns that owner and each caller keeps its
tail. _ensure_async_writer had no production caller (save() uses the _locked
form under _async_thread_lock); removed, tests retargeted, constructor comment
fixed. shutdown() sets _shutting_down under _async_thread_lock like the other
site so save()'s flag check and the enqueue cannot interleave with it.

tests/test_honcho_session_cache_bounds.py hoists its mid-file imports.
2026-09-13 19:05:39 +05:30
kshitijk4poor
b2ee58d24a fix(honcho): prune orphaned observation flags; share the local-platform set; drop test-shape defenses
A flush that rebuilds an evicted session's SDK session stores its observation
flags again, and the cap pass pruned every per-session dict but that one, so the
dict grew one entry per evicted-then-flushed session. The cap pass now prunes it
with the rest.

The peer-failure notice classified platforms with its own {"cli","tui","desktop"}
set, so an ACP session read the gateway wording ("do not suggest peerName"); it
now uses agent.coding_context.INTERACTIVE_CODING_PLATFORMS, which includes acp.

Three getattr/try-except guards existed only for test doubles (a bare __new__
provider, a SimpleNamespace config); the real types always carry the attribute.
Removed, and the two tests build real objects. An unset timeout resolves to the
client's 30s default, so the join-budget test expects 30, not the 5s floor.

Timing tests keep the >= 2s wall-clock bound the testing rules ask for.
2026-09-13 19:05:39 +05:30
kshitijk4poor
eed52456f6 fix(honcho): an author peer joins with its session's synced observation flags
main's _join_observation_flags returned the manager-wide booleans with a comment
saying #103889 would plug the per-session flags in here; this is that plug. The
two main-side tests that stubbed the old two-tuple _get_or_create_honcho_session
return move to the three-tuple.
2026-09-13 19:05:39 +05:30
Erosika
c04722102f fix(honcho): declare the injection block in config_schema so the desktop panel can pin sessionStart
The generic panel writes every field flat, and the plugin reads `injection` as one object, so a dotted `injection.sessionStart` field would never be read. Declare `injection` as a JSON field instead. Blank clears the pin.
2026-09-13 19:05:39 +05:30
Erosika
29113b24d8 fix(honcho): give the writer join and its drain the shutdown deadline
stop_async_writer bounded only the join and then drained the queue with no deadline, so a shutdown whose budget was already spent could still start uploads, and honcho-ai's add_messages has no per-call timeout. The join and the drain now share the shutdown deadline, no upload starts once it has passed, the writer skips its 2s retry once shutdown began, and shutdown logs one warning with the count left unsynced. The docstrings now say what the budget can do: stop new uploads and bound lock waits, while an upload already in flight runs to the client's HTTP timeout.
2026-09-13 19:05:39 +05:30
Erosika
3861167d63 fix(honcho): keep a failed late save reachable for flush_all after an eviction
A clean session can be evicted while its caller still holds it, and the caller's next save in turn mode flushed inline and never put the object back, so a failed upload left the batch nowhere flush_all() looks. Main reinserted after every flush. A failed save-time flush now reinserts the session when its key is free, or holds it in a retry list when a newer object owns the key, and flush_all() covers both; the collision path in _keep_until_flushed honors the flush result the same way.
2026-09-13 19:05:39 +05:30
Erosika
8523402db0 fix(honcho): bound the shutdown flush by the shutdown deadline
Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
2026-09-13 19:05:39 +05:30
Erosika
4b916022ab fix(honcho): surface the peer notice and audit the injection on the recall sync path
With recallSync on, prefetch popped only the auth notice and returned without writing the injection log. A session whose init failed for a missing user peer never told the model that memory was off, and the audit file stayed empty for every turn. The recall sync branch now pops the peer notice the way the async branch does and records each turn as injected or recall-sync-empty.
2026-09-13 19:05:39 +05:30
Erosika
d532d83eca refactor(honcho): trim duplicated tests and long docstrings
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer
and deferred-save tests that differed only in their inputs. Share the blocking
remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the
added docstrings to the what and the one non-obvious why.
2026-09-13 19:05:39 +05:30
Erosika
e9396ed8f4 fix(honcho): register the recall sync thread with its owner so shutdown joins it
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
2026-09-13 19:05:39 +05:30
Erosika
beab8b6f27 fix(honcho): keep observation flags across a flush rebuild, never orphan an evicted session, namespace dashboard logins
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the
cached path returned none, so recall fell back to the config snapshot. Both paths now return and
store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or
flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the
writer lock, and the writer drains its queue after the join, so a put that raced shutdown is
written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining
budget instead of a fixed ten seconds.

The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop
passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a
gateway platform supplies no user id, the peer notice and tool error no longer recommend
peerName, which would merge every user of that gateway onto one peer. README documents
`injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution.
2026-09-13 19:05:39 +05:30
Erosika
3da6a80d50 fix(honcho): refuse to mint a user peer when no identity or peerName exists
a desktop or cli session with no peerName in honcho.json and no gateway
user id landed on a peer derived from the session key: user-default-<dir>
for per-directory sessions, user-<channel>-<chat> for keyed ones. every
directory got its own phantom peer, so the operator's turns and memory
never reached their real peer and the injected representation went stale
(#93326).

_resolve_user_peer_id now raises HonchoPeerUnresolvedError instead of
deriving a name. a peer is either the declared peerName or an identity
the transport supplied. the provider records the failure, tells the model
once that memory is off and which key to set, returns the same detail
from tool calls, and stops retrying init because a missing config key
does not heal mid-session. the memory-file migration gate loses its
"no owner and no runtime identity" branch: that cohort no longer has a
session to migrate into. whitespace-only peerName is treated as unset
rather than sanitized to "--".
2026-09-13 19:05:39 +05:30
Erosika
cc3bfc2120 fix(honcho): read the sessionStart pin defensively in the first-turn formatter
tests and callers can build the provider without initialize(); the
formatter now treats a missing pin as unpinned instead of raising.
2026-09-13 19:05:39 +05:30
Erosika
c3a5649aac fix(honcho): join every plugin thread on shutdown within one budget
provider shutdown joined the dialectic, sync and memwrite threads for 5s
each and then the async writer, but the session-init thread,
honcho-base-first and honcho-context-prefetch were never joined. any of
them still blocked in httpx when the interpreter finalized aborted the
process with SIGABRT 134 (#37632, #60616, #33485). the 5s join was also
shorter than the 30s http timeout a blocked call can hold (#33485).

spawn_context_thread now registers each thread under its owner (the
provider or the manager) in a weak registry, and shutdown joins every live
thread of both owners inside one deadline: at least 5s, or the configured
http timeout when longer. the manager refuses new prefetch threads once
shutdown began and flushes a late save() inline instead of respawning the
writer. threads that outlive the budget are named in a warning.

the sdk client exposes close() on its http pool. one client is shared by
every manager with the same identity in a gateway process, so a per-agent
shutdown cannot close it; close_honcho_clients() closes all pools and is
registered with atexit when the first client is built, the pattern the
hindsight and mem0 plugins use. follows #69070, #33701, #7627.
2026-09-13 19:05:39 +05:30
Erosika
0cb2977e84 fix(honcho): cap the manager caches and keep unsynced sessions out of eviction
the idle sweep from #71463 left three growth paths open. _peers_cache had no
bound at all, _session_observation (from #98941) grew one entry per session
id and kept orphans when an init failed after add_peers, and a burst of
distinct sessions inside one ttl window was not bounded. the sweep also
evicted sessions whose messages had not reached honcho yet, which in
"session" write mode drops the only copy, and _flush_session re-inserted an
evicted session into the cache with no observation flags, so recall for it
routed from the config snapshot instead of its server config.

_cache and _sessions_cache now cap at 128 entries and _peers_cache at 512,
evicting least recently used first (dict order, refreshed on every hit).
a session with unsynced messages is never evicted by the sweep or the cap.
_configure_session_peers returns the synced flags and get_or_create stores
them under _cache_lock next to the cache entry, so the observation dict can
hold no id the cache does not; eviction drops both. recall reads go through
_cached_session, which stamps updated_at, so a read-only session survives
the idle sweep. follows #71461, #71463, #98936.
2026-09-13 19:05:39 +05:30
Erosika
ff29c27003 fix(honcho): hold a per-session lock across select, send and the _synced flip
_flush_session read the unsynced messages, posted them, and only then set
_synced. in a short-lived run the async writer draining save()'s queue and
the exit-time flush_all() both saw the same batch unsynced, so every turn
of a one-shot run landed in honcho twice (#92458).

each HonchoSession now carries an RLock that _flush_session holds around
the whole select, send, mark sequence. the second flusher enters after the
first marked the batch and sends nothing; different sessions still flush
in parallel. the lock lives on the session object instead of a manager
dict keyed by session id (#86094, #92787): it needs no eviction, and two
flushers of one message list can never hold different locks. tests
adapted from #86094 and #92787.
2026-09-13 19:05:39 +05:30
Aleksei Ivanov
280ac22663 fix(honcho): bound local session cache growth
HonchoSession.messages grew forever -- add_message() appended every
turn and _flush_session() only marked entries synced, never trimmed
them. HonchoSessionManager's four caches (_cache, _peers_cache,
_sessions_cache, _context_cache) had no eviction path besides an
explicit /new reset. A long-lived channel that is never manually
reset accumulates both for the gateway's entire uptime.

Honcho is the durable source of truth (get_or_create() already
re-fetches history from Honcho on a cache miss), so both bounds are
safe: evicting an idle local entry only costs one extra Honcho
round-trip next time that key is used.

- Trim already-synced messages beyond a retention window right after
  a successful flush; unsynced messages are never touched.
- Add a rate-limited idle-TTL sweep triggered opportunistically from
  get_or_create(), so no new background task/watcher wiring is
  needed.

Fixes #71461
2026-09-13 19:05:39 +05:30
686f6c61
2614b900b8 fix(honcho): omit reasoning-contaminated session summaries
Strip think blocks and drop planning-only Honcho summaries at fetch,
format, and cache so they cannot be reinjected as trusted memory.
2026-09-13 19:05:39 +05:30
liuhao1024
a4939af48b fix(memory): scope honcho observation flags per session (#98936)
The four observation booleans on HonchoSessionManager were manager-wide
mutable state initialized from per-session server configs: every session
setup overwrote them, so the last session to initialize retuned recall
routing for all other sessions the manager serves. Store the server-synced
flags under each session's own id instead; the manager-level fields stay
as the config snapshot and sessions that never synced fall back to them.
2026-09-13 19:05:39 +05:30
Erosika
5af0bf4111 test(honcho): pin sessionStart filtering and the injection audit; expose logging in setup
injection.sessionStart: unset renders every component in table order, an
empty list renders nothing, a pinned list renders only those names in table
order (not config order), a host block pin beats root, a non-list value is
treated as unset, and initialize() reads the pin.

logging: off by default, the logging key or HONCHO_LOGGING turns it on, a
host block can turn it back off, HONCHO_INJECTION_LOG overrides the path,
each record carries reason/turn/session_key/bytes/payload, an unwritable
path never raises, and a tools-mode prefetch logs its reason. the record
holds the user's representation verbatim, which is why the default stays off.

the two session-context tests that asserted exact context() kwargs now expect
tokens= as well (adopted from #92964). config_schema declares the logging
switch so hermes memory setup shows it.
2026-09-13 19:05:39 +05:30
Hector Suzanne
3b5d9116a7 fix(honcho): honour contextTokens cap on summary/peer context calls
Salvage of #70951 (Willkons / Alice-Willk-bot). Two of three
session.context() sites never forwarded the configured cap, so Honcho
always returned honcho_chat_summary_long.

Adds the missing get_session_context tokens= regression the sweeper
asked for on that PR.
2026-09-13 19:05:39 +05:30
Eugene Eisenstein
8744d7f7c8 repair injection.sessionStart 2026-09-13 19:05:39 +05:30
Eugene Eisenstein
61438268c6 fix(logging): repair the logging config key and HONCHO_LOGGING
The logging is also improved a little by saving reasons
2026-09-13 19:05:39 +05:30
kshitijk4poor
8900fb2cf8 fix(honcho): adopt a sibling's on-disk rotation even inside our exchange cooldown
force_refresh_token gated the adopt-from-disk paths behind the failure cooldown.
After one of our exchanges failed transiently, a 401 within the next 30s returned
None even when a sibling process had already rotated and written a valid grant,
so the operation raised HonchoAuthError with a good token sitting on disk.
Adopting is a disk read, not an exchange; the cooldown exists to stop replaying a
single-use refresh token, so the two adopt checks now run before the gates.

_write_config parsed honcho.json twice under the lock and only wrapped the first
read into ConfigWriteRefused; _refuse_unparseable now returns the parsed dict and
the branches use it. The getattr/isinstance duck-typing collapses to one guard.
cli._read_config reuses oauth's tolerant reader (BOM-tolerant, like the strict
reader the write side uses) instead of its own utf-8 copy.
2026-09-13 19:05:28 +05:30
kshitijk4poor
24b33d42e0 fix(honcho): every honcho.json writer holds _refresh_lock and reads strictly
The CLI's _write_config took only the best-effort file lock, so an in-process
refresh thread could still interleave with a command's read-modify-write. The
dashboard's Honcho save (_write_provider_honcho) still seeded its whole-file
rewrite from a tolerant reader, so a honcho.json that exists but does not parse
was replaced by the active host's block alone from the UI - the same bug class
this PR closes on the CLI and refresh paths. `hermes profile create --clone`
swallowed the new ConfigWriteRefused as "plugin not installed".

One parametrized test covers the web writer for the corrupt and parseable cases.
2026-09-13 19:05:28 +05:30
Erosika
56231e51ad fix(honcho): advance the read baseline after each write so a revert reaches disk
_write_config applied a command's edits relative to the snapshot the read took, but never moved that snapshot after a successful write. A second write on the same object therefore compared A -> B -> A against A, saw no change, and left disk at B. After a write the snapshot and path now follow the caller's dict, so the next write applies only the edits made since.
2026-09-13 19:05:28 +05:30
Erosika
2f81f6a831 fix(honcho): keep a rotation on honcho.json when the read was seeded from another file
_read_config() falls back to ~/.honcho/config.json or the default profile when honcho.json does not exist, so cfg.path never matched the write path and _write_config() wrote the whole dict. That overwrote a refresh rotation a serve child had written onto the honcho.json the setup login created moments earlier. When the local file exists at write time, the seed snapshot is now overlaid with the local file and only the command's edits are applied onto it.
2026-09-13 19:05:28 +05:30
Erosika
36c98cb5eb fix(honcho): hold the refresh locks while save_config rewrites honcho.json
save_config read the file and wrote it back without the locks the token refresh holds around its own read, exchange and write. A refresh that landed between the two steps was overwritten and its consumed refresh token was gone. The read and the write now run under both locks.
2026-09-13 19:05:28 +05:30
Erosika
be8ce602d0 fix(honcho): keep a rotation that lands while the setup wizard is still asking questions
install_grant writes the login grant to disk before the wizard's later prompts. _apply_grant_to_host wrote it only into the live cfg, so _write_config saw the grant as an edit and copied it over whatever a serve child rotated onto disk meanwhile. The snapshot now takes the grant too, so the final save leaves the newer on-disk grant alone.
2026-09-13 19:05:28 +05:30
Erosika
498545bd23 refactor(honcho): trim the oauth persistence docstrings and helpers
_ReadConfig builds its snapshot in __init__; _read_config no longer
assigns the attributes after the fact. cmd_enable prints its hint
inline, since _no_credential_hint had one caller. The split
force_refresh_token signature and call fit on one line. Docstrings and
comments keep the what and the one non-obvious why. No behavior
changes; _apply_edits keeps its original body.
2026-09-13 19:05:28 +05:30
Erosika
dd6b183c02 fix(honcho): only on-disk credentials count when a write enables a host block
clone and enable accepted HONCHO_API_KEY from the environment as proof the
block could authenticate and wrote enabled: true. the variable can be absent
from the next process, leaving an enabled block with nothing behind it, which
is the cohort the previous commit set out to remove.

_resolve_api_key takes env=False at both write sites, so a block is enabled
only when honcho.json itself holds a key, an oauth grant, or a self-hosted
baseUrl. status and setup keep the environment fallback for display.
2026-09-13 19:05:28 +05:30
Erosika
6966705751 fix(honcho): cli writes hold the refresh lock and merge only their edits onto disk
every hermes honcho command read honcho.json, changed a field, and wrote the
whole dict back with no lock. a token refresh in another process that landed
between the read and the write was overwritten with the old access and
refresh tokens, and the next refresh replayed a rotated single-use token.

_write_config now holds the same cross-process file lock the refresh path
holds, re-reads disk under it, and applies only the keys the command changed
since its _read_config(). untouched keys keep their on-disk value, so a
rotation survives; a credential the command set on purpose still wins. a
write with no prior read keeps today's whole-file behavior.
2026-09-13 19:05:28 +05:30
Erosika
d3b599f374 docs(honcho): prune comments in the oauth persist commits
issue numbers move to the commit bodies; docstrings keep what and the one
non-obvious why.
2026-09-13 19:05:28 +05:30