Commit Graph

574 Commits

Author SHA1 Message Date
teknium1
0a1bc07f15 docs(honcho): say that initOnSessionStart blocks tools-mode startup until Honcho answers
What: the `initOnSessionStart` setup-schema description and the Honcho docs page
now state that in `recallMode: tools` the eager init runs synchronously during
agent construction, that it should stay false for Desktop or a local Honcho that
may be down, and that `timeout` in honcho.json caps each SDK call (both keys
added to the Full Config Reference table).

Why: #51492 reports Desktop `session.resume` / `prompt.submit` timeouts when a
local Honcho is down with `initOnSessionStart: true`. The synchronous eager path
is intentional design (#51562 was closed for that reason:
"ready-before-first-tool-call semantics"), so the fix for users is knowing the
trade-off and the two knobs that bound it.
2026-09-19 10:39:32 -07:00
teknium1
80f4c6d6bf fix(acp): warm numpy with the memory provider so hindsight is covered too (#58083)
The warm-up imported only the provider module. holographic / mnemosyne import
numpy at module top, but hindsight defers the ML stack to is_available() ->
_check_local_runtime() (importlib of hindsight / sentence_transformers), which
ran later on a to_thread worker racing acp-mcp-discovery — the reported hindsight
stack was still reachable. The deadlock partner is numpy's lazy _core init in
every reporter's dump, and a plain `import numpy` up front was every reporter's
workaround, so import_memory_provider_module now also imports numpy (best-effort)
once the provider module is in.

Also: import_memory_provider_module() defaults to the configured memory.provider,
so entry.py drops its duplicate config resolver and outer try; the "ONLY thread"
comment is reworded — hermes_cli's plugin-discovery thread is already running
when hermes acp dispatches.
2026-09-19 02:43:28 -07:00
teknium1
6a93afa0a8 fix(acp): pre-import the memory provider on the main thread on Windows (#58083)
Every faulthandler dump in the thread shows session/new stuck in numpy's
create_module on the main thread while another thread (MCP discovery / ACP
stdin reader) sits in the same lazy import chain — a first-time native
extension import racing another thread deadlocks on Windows (holographic,
mnemosyne and hindsight all reproduce it; a sitecustomize `import numpy`
before any thread exists resolves it every time).

hermes acp now imports the configured memory.provider's module on the main
thread before the MCP-discovery thread and asyncio.run() start (Windows only —
the deadlock is Windows-specific and the import is paid once either way).
plugins.memory.import_memory_provider_module imports the module without
constructing a provider or running register(); the agent build later finds it
in sys.modules.

Trimmed from #91775 (@tigercraft4): same placement and gating; reuses the
existing plugin loader instead of a second module-import routine.

Co-authored-by: tigercraft4 <tigercraft4@tigercraft4.com>
2026-09-19 02:43:28 -07:00
teknium1
396013b17e test: one invariant per bounded cache (LSP docs/baselines, tool-call logger, fuzzy roots, hindsight turns)
Each test is red on origin/main and green on the fix; each also asserts the
control case (evicted LSP file re-opens with diagnostics, every log line
written once, unexpired fuzzy root kept, every hindsight turn shipped once).
Trims the hindsight comment to the why.
2026-09-19 01:55:38 -07:00
Tikkanaditya Siddartha Jyothi
6e1de4850e fix(hindsight): bound the append-mode session turn buffer
_session_turns accumulated every turn's text for the whole session, reset
only at session boundaries, so a never-ending session grew without bound
independent of context compaction.

In append mode each retain ships only the delta since the last watermark
(_session_turns[_last_retained_turn_count:]), and a retained turn is never
read again — sync_turn always slices from the watermark and flush-on-switch
flushes what's left. So after an append retain the buffer drops the retained
prefix (clear() + reset the watermark), bounding it to the un-retained tail.

Overwrite mode is deliberately untouched: legacy/overwrite APIs resend the
whole session each retain because each retain replaces the document, so that
path must keep every turn.

Tests: append trims the retained prefix while shipping every turn exactly
once (no loss, no duplication), stays bounded across 100+ turns, and
overwrite mode keeps the full buffer.

(cherry picked from commit cb77e006754dd4ad5c880f22310d3ac7b3472739)
2026-09-19 01:55:38 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
Sora-bluesky
9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
outpoints
94efd22177 fix(honcho): preserve title provenance through upstream peer routing 2026-09-15 22:30:11 -07:00
outpoints
44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints
231e1c1204 fix(honcho): resolve sessions against agent cwd, not process cwd
The Honcho provider resolved per-repo/per-directory session names from
os.getcwd(), which is the backend process launch directory on
Desktop/gateway hosts (typically $HOME), not the user's workspace. With
a manual sessions map entry for the home directory, every Desktop
conversation landed in that fallback bucket instead of the project's
per-repo session.

Use agent.runtime_cwd.resolve_agent_cwd() — the same single source of
truth already used for system-prompt and context-file discovery — so
Honcho session routing agrees with everything else about where the
agent logically lives. It honors the pinned session cwd, then
TERMINAL_CWD, then the launch directory; CLI sessions launched inside a
project resolve identically either way.

Adds a regression covering the Desktop-style case: backend launched in
$HOME, workspace elsewhere, home-directory manual map present.

Refs #24740

(cherry picked from commit cfb32757de2509dd404acf98311e357d73fa741e)
2026-09-15 22:30:11 -07:00
outpoints
3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
teknium1
e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
Teknium
5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00
kshitijk4poor
d793f7b9fb fix(supermemory): explicit failure sentinel and bounded pending-turn buffer
Two review findings on the salvage stack:

- _write_turns' _quietly consolidation keyed failure on a None result,
  implicitly assuming add_memory never legitimately returns None. A None-
  returning stub (the most common mock idiom) would mark every successful
  write failed and re-append the batch forever. Module-level _FAILED
  sentinel: only a raised exception re-queues.
- _pending_turns had no bound: a persistently failing service accumulated
  one entry per turn for the process lifetime (gateway runs never
  re-initialize), and every retry re-sent the whole accumulated payload —
  O(n^2) upload bytes, and a size-rejected batch could never shrink. Cap
  the buffer (50 turns / 256 KiB, drop oldest, warn once per trim).
2026-09-15 11:55:10 +05:30
kshitijk4poor
bc1d4776db docs(supermemory): purge stale session-ingest / /v4/conversations references
The per-turn capture rewrite removed the raw urllib /v4/conversations
ingest, but stale references survived outside the diff hunks:

- README "Behavior" still carried the "written once via the conversations
  endpoint" paragraph contradicted by the new bullets right above it.
- website memory-providers (EN + zh-Hans) still listed full-session ingest,
  session-end /v4/conversations ingest, and ingest in the base-url and
  api_timeout rows — the PR had updated one line per file but missed the
  rest of the section.

Now every surface describes per-turn documents.add capture with retry.
2026-09-15 11:55:10 +05:30
kshitijk4poor
3cad3c6db8 fix(supermemory): route capture-write logging through _quietly; drop dead default in custom id
Two cleanups on the salvaged per-turn capture path (review of #109359):

- _write_turns hand-rolled the try/except/log that the file's own _quietly
  helper already provides (same exception class, level, exc_info shape);
  the deleted _ingest path used _quietly for the identical call. Consolidate:
  add_memory always returns a dict, so a None result is the failure sentinel.
- _capture_custom_id's `or 'hermes'` never fires: _sanitize_tag already
  returns _DEFAULT_CONTAINER_TAG on empty input, and the literal duplicated
  the constant the helper owns. Verified behavior-preserving (tests green
  with the fallback artificially restored).
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
c6da9e0788 fix(supermemory): serialize capture writes and document at-least-once retry
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.

Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
1b76cfff83 fix(supermemory): keep pending turns across session switch until written
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
03627dbf95 fix(supermemory): write turns via documents API instead of session-end conversations ingest
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).

Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.

Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.

Metadata stays type/session_id/timestamp plus the existing sm_source.
2026-09-15 11:55:10 +05:30
teknium1
732504c8d6 fix(memory/byterover): brv child carries the served profile's cloud key, never the launch profile's
Under gateway.multiplex_profiles os.environ holds the default profile's .env, so `_run_brv`
building the child env from raw os.environ curated a secondary profile's turns into the DEFAULT
profile's ByteRover cloud account (and prefetched the default's memories into the secondary's
context). The local half was already profile-scoped (`_get_brv_cwd`).

The child env now comes from `build_subprocess_env` and, under multiplex, strips the launch
profile's residue and sets BRV_API_KEY only from the served profile's secret scope — a miss means
no cloud key. Single-profile installs pass the process env through unchanged.

Closes #108993 (report and fix direction by @jonpol01).
2026-09-13 14:41:26 -07:00
John Paul Soliva
2bade5daa7 fix(memory/hindsight): pin the isolation-vs-shaping split for scoped reads
`langfuse._secret` and `azure_identity_adapter._scoped_env` were changed to
raise rather than fall back, because swallowing `UnscopedSecretError` hides the
spawn-site bug the exception exists to surface. `_scoped_setting` looked like it
contradicted that, so make the split explicit and pin it.

Hindsight already follows the contract for everything that decides WHERE data
goes: `mode`, `apiKey` and the `bankId` partition read through bare
`get_secret`, so a scopeless multiplexed read raises. In `_load_config` that
raise happens on `HINDSIGHT_MODE` before any shaping value is reached, so the
swallow below cannot mask an isolation failure.

Presentation shaping is deliberately not in that class. `MemoryManager._each_provider`
logs an `initialize` failure at WARNING and drops the provider for the session,
so raising there would cost the whole memory provider because a speaker prefix
could not be resolved. It degrades to the provider's own default instead —
never to `os.environ`, which under multiplex is the default profile's.

The test names the offending key rather than asserting that something raised:
routing `mode` through the shaping helper shifts the failure to
`HINDSIGHT_API_KEY`, which a bare `pytest.raises` would still accept.
2026-09-13 14:41:26 -07:00
John Paul Soliva
a48dde5316 fix(memory/hindsight): resolve retain shaping through the profile scope, never os.environ
_load_config() reads the Hindsight bank, mode and retain tags through the profile
secret scope, but _apply_retain_settings() then discarded that answer and re-read
os.environ whenever the config value was falsy:

    return cfg.get(key) or os.environ.get(env_var, default)

Under gateway.multiplex_profiles os.environ holds the DEFAULT profile's .env, so a
secondary profile's scoped miss came back as the default profile's retain tags,
observation scopes, source and speaker prefixes — the fallback-after-miss shape
gateway/AGENTS.md forbids. Tags are Hindsight's retrieval partition and
metadata.source is opt-in by design, so the secondary's memories were both
mislabelled and selectable by the default profile's tag filters.

Both halves now go through _scoped_setting(), which resolves the value with
get_secret() and falls back to the provider's OWN default — a miss is a miss, the
same rule embedded.py already applies to the daemon's key and base URL. The three
raw reads left inside _load_config() (retain_source, retain_user_prefix,
retain_assistant_prefix), directly under the comment declaring them per-profile,
go through it too.

Single-profile deployments are unchanged: with no scope installed get_secret()
still reads the process env, where the value IS this profile's own.

Fixes #108865
2026-09-13 14:41:26 -07:00
kshitijk4poor
476dbfed3a perf(honcho): peers map scans the profile directory once, not once per account row
_render_peers_map_view called _sibling_resolutions per account, and each call
ran _all_profile_host_configs(), which parses every profile's config.yaml via
list_profiles() and re-reads its honcho.json; show() re-runs after every
workspace switch. cmd_peers_map now scans once and threads the rows through.

_seen_gateway_accounts turned a locked or corrupt state.db into the same [] that
means "no gateway traffic yet"; it now notes the error on stderr before returning
the empty list. The missing-table fallback for old schemas is unchanged.
2026-09-13 19:05:39 +05:30
kshitijk4poor
71f3516df2 fix(honcho): peers map reads one SDK page at a time; legacy root-level installs keep their shape
Iterating a honcho SyncPage walks every following page, so _api_workspace_peers
pulled the whole workspace on its first call, never hit the `< 50` break, re-walked
pages 2..N and returned duplicates; the 200-peer cap that exists so a public bot's
workspace cannot stall the CLI was defeated and the `p<N>` picks pointed at the
wrong peers. It now reads `.items` per page and stops at the cap; `_api_workspaces`
reads `.items` too.

The wizard's `new_host` probe looked only at the host block, so an install that
keeps peerName/enabled/workspace at the root with no hosts.hermes block read as
fresh and Enter defaulted to pinning every account onto one peer. The probe
includes the root.

`_seen_gateway_accounts` dropped rows whose origin said `is_bot`, but
SessionSource.to_dict never serializes that field, so the filter was dead; removed
with its test row. `_sanitize_peer_id` was a copy of session_peers.sanitize_peer_id.
A confirmed workspace repoint is persisted, so the exit line no longer says
"Nothing changed" after one.
2026-09-13 19:05:39 +05:30
Erosika
e05f6d82ce fix(honcho): default the wizard to the detected shape on every existing install
The identity step treated any host block without a mapping key as a new install and defaulted the choice to single peer. An install with enabled, workspace and peerName then had Enter write pinUserPeer: true and merge every gateway account onto the operator. cmd_setup now decides new-ness before its prompts populate the block, and only a block with none of the mapping, peerName, workspace or enabled keys defaults to single.
2026-09-13 19:05:39 +05:30
Erosika
7dacc2ed35 fix(honcho): describe what _seen_gateway_accounts can list
The docstring said grouping session rows by (source, user_id) enumerates every account the gateway handled. record_gateway_session_peer overwrites a row's user_id, so a shared thread keeps only its last author. The docstring now says that, and a test pins it.
2026-09-13 19:05:39 +05:30
Erosika
0ede483479 fix(honcho): clone sessionAiPeerPrefix into new profile host blocks
clone_honcho_for_profile copied sessionPeerPrefix but not sessionAiPeerPrefix. A new profile cloned from a default block with the AI prefix on fell back to unprefixed session names and collided with the default profile's gateway sessions.
2026-09-13 19:05:39 +05:30
Erosika
e467c0ef2c fix(honcho): preview peer resolution through the runtime resolver
peers map showed sanitize(prefix + user_id) for prefixed accounts. The runtime appends a sha256 suffix when sanitizing changed the id or it collides with an explicit peer, so the preview named a peer the gateway never writes to. The preview now builds a HonchoSessionManager with no client and asks it.
2026-09-13 19:05:39 +05:30
Erosika
841b768cd7 fix(honcho): keep an empty host alias map when the last alias is cleared
Clearing the last alias in peers map popped userPeerAliases from the host block, and the host then inherited the root aliases again. The host block now keeps an empty map, which both config readers treat as an explicit override. The summary says when that empty map hides root aliases.
2026-09-13 19:05:39 +05:30
Erosika
7e10f06a5a refactor(honcho): shorten the peers-map helpers
_api_workspace_peers returns peer IDs instead of dicts. The created date
it carried was never printed. Callers index the list directly.

_preview_peer_resolution uses _sanitize_peer_id instead of a local copy.
_sibling_resolutions strips the display suffix itself; its one caller did.

_seen_gateway_accounts opens the database under contextlib.closing and
parses origin_json once into a dict. _save_alias_map picks the target
block first and writes it once. cmd_peers_map renders through one local
show() and computes the previous resolution on one path for both account
numbers and typed runtime IDs.
2026-09-13 19:05:39 +05:30
Erosika
506dcacd9f fix(honcho): list gateway accounts whose session rows predate session_key
peers map read seen accounts from state.db but required a non-empty
session_key. every telegram row on a long-lived install written before the
gateway stamped that column was dropped, so the picker showed no accounts
while the same rows grouped fine by (source, user_id). the key was never
used by the grouping.
2026-09-13 19:05:39 +05:30
Erosika
783d4bdffc docs(honcho): prune wizard and peers-map comments
The picker-cap comment repeated the constant's name; the
sessionAiPeerPrefix field comment restated the field name and its
symmetry with sessionPeerPrefix before reaching the collision it
prevents. Both now state only the why.
2026-09-13 19:05:39 +05:30
Erosika
dc7c8673d1 feat(honcho): add 'hermes honcho peers map' for interactive account-to-peer mapping
Extends the read-only 'hermes honcho peers' view with a 'map' action
that joins two sources: workspace peers fetched from the Honcho API,
labeled from local config (your peer, each profile's AI peer, alias
targets, runtime peers of seen accounts, user-* fallback peers,
'unrecognized' otherwise), and the gateway accounts recorded in
state.db with what each currently resolves to. Targets are picked
from the workspace list so a typo cannot silently create a peer;
every assignment states its consequence (aliases move future
messages only; a runtime peer left behind keeps its history).

'w' lists every workspace the key can reach — the wrong-workspace
fallback — and can repoint the profile's workspace on explicit
confirmation. With multiple profiles, saving a root-cascading map
asks whether to write root or fork the host block, root writes warn
when sibling profiles sit on other workspaces, and the accounts
table marks siblings that resolve an account differently. Offline
the command degrades to typed targets over the local account list.
The setup wizard's gateway step closes by pointing at the command.
2026-09-13 19:05:39 +05:30
Eli Robbins
55dfb517e4 fix(honcho): bust gateway agent cache on session-prefix flips
Address review feedback on #39130: the new sessionAiPeerPrefix setting
affects the resolved Honcho session key, which HonchoMemoryProvider freezes
at construction (self._session_key). Because it wasn't part of the gateway's
cached-agent signature, a live config flip left an existing gateway session
bound to its old, AI-peer-agnostic Honcho session until an unrelated eviction
or restart.

Add honcho.session_ai_peer_prefix to _HONCHO_CACHE_BUSTING_KEYS and the
_extract_honcho_cache_busting_config values so a flip rebuilds the cached
agent on the next turn, mirroring the existing aiPeer / pin_peer_name /
runtime_peer_prefix contracts.

Also close the symmetric gap for the pre-existing user-side sessionPeerPrefix:
it feeds the same resolve_session_name output (per-session/title/per-repo/
per-directory strategies) into the same frozen _session_key, so it had the
identical live-flip staleness bug and was likewise absent from the cache
signature. Fixing both keeps the two prefixes consistent.

Add one config-flip regression test covering both keys, alongside the
existing Honcho cache-signature test.
2026-09-13 19:05:39 +05:30
Eli Robbins
0a7c17159c feat(honcho): add sessionAiPeerPrefix to isolate sessions per AI peer
The gateway_session_key branch of resolve_session_name() returns an
AI-peer-agnostic name, so multiple AI peers sharing one workspace +
peerName + gateway chat key collide on a single Honcho session.

Add sessionAiPeerPrefix (symmetric counterpart to sessionPeerPrefix):
when set, the resolved session name is prefixed with {ai_peer}- on every
resolution path. The prefixed name is re-run through the session-id
length cap so the prefix can never exceed Honcho's limit.

- config field + host/root parsing in client.py
- public resolve_session_name() wraps a new _resolve_session_name_base()
- tests covering parsing, the gateway-key case, cross-peer disjointness,
  the length cap, and a disabled-by-default regression guard
- README: config table + resolution notes
2026-09-13 19:05:39 +05:30
Erosika
21180ae7e8 feat(honcho): rework the setup wizard's gateway mapping step around Honcho's peer model
The step declares its scope up front: human mapping only, with each
Hermes profile bringing its own AI peer. A note explains aliases as
the join between platform accounts and named peers. Each shape now
says when it fits. Fresh configs default the choice to [1] single
peer — the common personal setup — instead of [3], which silently
fragmented a solo operator's gateway account away from their
peerName history. Configured setups keep their detected shape as
the default.
2026-09-13 19:05:39 +05:30
kshitijk4poor
d14734df1f fix(honcho): thread registry rides agent.memory_provider.spawn_context_thread
main folded the plugin's contextvars-inheriting thread wrapper into
agent.memory_provider.spawn_context_thread. The registry this stack adds for
join_plugin_threads keeps a thin client.spawn_context_thread that calls the core
spawner and records the thread under its owner; the four spawn sites import it
from client. main's inherit-profile test builds a real provider now that
_spawn_write is an instance method (the owner is what gets joined).
2026-09-13 19:05:39 +05:30
kshitijk4poor
f1273ed704 refactor(honcho): one reclaim-key helper, one cache-size constant, no dead writer starter
_SESSION_CACHE_MAX_SIZE was assigned twice with two comments describing one
constant. _retain_for_retry and _keep_until_flushed shared the has-unsynced /
current-owner / reclaim body and differed only in what to do when a newer object
owns the key; _reclaim_key_locked returns that owner and each caller keeps its
tail. _ensure_async_writer had no production caller (save() uses the _locked
form under _async_thread_lock); removed, tests retargeted, constructor comment
fixed. shutdown() sets _shutting_down under _async_thread_lock like the other
site so save()'s flag check and the enqueue cannot interleave with it.

tests/test_honcho_session_cache_bounds.py hoists its mid-file imports.
2026-09-13 19:05:39 +05:30
kshitijk4poor
b2ee58d24a fix(honcho): prune orphaned observation flags; share the local-platform set; drop test-shape defenses
A flush that rebuilds an evicted session's SDK session stores its observation
flags again, and the cap pass pruned every per-session dict but that one, so the
dict grew one entry per evicted-then-flushed session. The cap pass now prunes it
with the rest.

The peer-failure notice classified platforms with its own {"cli","tui","desktop"}
set, so an ACP session read the gateway wording ("do not suggest peerName"); it
now uses agent.coding_context.INTERACTIVE_CODING_PLATFORMS, which includes acp.

Three getattr/try-except guards existed only for test doubles (a bare __new__
provider, a SimpleNamespace config); the real types always carry the attribute.
Removed, and the two tests build real objects. An unset timeout resolves to the
client's 30s default, so the join-budget test expects 30, not the 5s floor.

Timing tests keep the >= 2s wall-clock bound the testing rules ask for.
2026-09-13 19:05:39 +05:30
kshitijk4poor
eed52456f6 fix(honcho): an author peer joins with its session's synced observation flags
main's _join_observation_flags returned the manager-wide booleans with a comment
saying #103889 would plug the per-session flags in here; this is that plug. The
two main-side tests that stubbed the old two-tuple _get_or_create_honcho_session
return move to the three-tuple.
2026-09-13 19:05:39 +05:30
Erosika
c04722102f fix(honcho): declare the injection block in config_schema so the desktop panel can pin sessionStart
The generic panel writes every field flat, and the plugin reads `injection` as one object, so a dotted `injection.sessionStart` field would never be read. Declare `injection` as a JSON field instead. Blank clears the pin.
2026-09-13 19:05:39 +05:30
Erosika
29113b24d8 fix(honcho): give the writer join and its drain the shutdown deadline
stop_async_writer bounded only the join and then drained the queue with no deadline, so a shutdown whose budget was already spent could still start uploads, and honcho-ai's add_messages has no per-call timeout. The join and the drain now share the shutdown deadline, no upload starts once it has passed, the writer skips its 2s retry once shutdown began, and shutdown logs one warning with the count left unsynced. The docstrings now say what the budget can do: stop new uploads and bound lock waits, while an upload already in flight runs to the client's HTTP timeout.
2026-09-13 19:05:39 +05:30
Erosika
3861167d63 fix(honcho): keep a failed late save reachable for flush_all after an eviction
A clean session can be evicted while its caller still holds it, and the caller's next save in turn mode flushed inline and never put the object back, so a failed upload left the batch nowhere flush_all() looks. Main reinserted after every flush. A failed save-time flush now reinserts the session when its key is free, or holds it in a retry list when a newer object owns the key, and flush_all() covers both; the collision path in _keep_until_flushed honors the flush result the same way.
2026-09-13 19:05:39 +05:30
Erosika
8523402db0 fix(honcho): bound the shutdown flush by the shutdown deadline
Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
2026-09-13 19:05:39 +05:30
Erosika
4b916022ab fix(honcho): surface the peer notice and audit the injection on the recall sync path
With recallSync on, prefetch popped only the auth notice and returned without writing the injection log. A session whose init failed for a missing user peer never told the model that memory was off, and the audit file stayed empty for every turn. The recall sync branch now pops the peer notice the way the async branch does and records each turn as injected or recall-sync-empty.
2026-09-13 19:05:39 +05:30
Erosika
d532d83eca refactor(honcho): trim duplicated tests and long docstrings
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer
and deferred-save tests that differed only in their inputs. Share the blocking
remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the
added docstrings to the what and the one non-obvious why.
2026-09-13 19:05:39 +05:30
Erosika
e9396ed8f4 fix(honcho): register the recall sync thread with its owner so shutdown joins it
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
2026-09-13 19:05:39 +05:30
Erosika
beab8b6f27 fix(honcho): keep observation flags across a flush rebuild, never orphan an evicted session, namespace dashboard logins
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the
cached path returned none, so recall fell back to the config snapshot. Both paths now return and
store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or
flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the
writer lock, and the writer drains its queue after the join, so a put that raced shutdown is
written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining
budget instead of a fixed ten seconds.

The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop
passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a
gateway platform supplies no user id, the peer notice and tool error no longer recommend
peerName, which would merge every user of that gateway onto one peer. README documents
`injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution.
2026-09-13 19:05:39 +05:30
Erosika
3da6a80d50 fix(honcho): refuse to mint a user peer when no identity or peerName exists
a desktop or cli session with no peerName in honcho.json and no gateway
user id landed on a peer derived from the session key: user-default-<dir>
for per-directory sessions, user-<channel>-<chat> for keyed ones. every
directory got its own phantom peer, so the operator's turns and memory
never reached their real peer and the injected representation went stale
(#93326).

_resolve_user_peer_id now raises HonchoPeerUnresolvedError instead of
deriving a name. a peer is either the declared peerName or an identity
the transport supplied. the provider records the failure, tells the model
once that memory is off and which key to set, returns the same detail
from tool calls, and stops retrying init because a missing config key
does not heal mid-session. the memory-file migration gate loses its
"no owner and no runtime identity" branch: that cohort no longer has a
session to migrate into. whitespace-only peerName is treated as unset
rather than sanitized to "--".
2026-09-13 19:05:39 +05:30
Erosika
cc3bfc2120 fix(honcho): read the sessionStart pin defensively in the first-turn formatter
tests and callers can build the provider without initialize(); the
formatter now treats a missing pin as unpinned instead of raising.
2026-09-13 19:05:39 +05:30