4 Commits

Author SHA1 Message Date
kshitijk4poor
b2ee58d24a fix(honcho): prune orphaned observation flags; share the local-platform set; drop test-shape defenses
A flush that rebuilds an evicted session's SDK session stores its observation
flags again, and the cap pass pruned every per-session dict but that one, so the
dict grew one entry per evicted-then-flushed session. The cap pass now prunes it
with the rest.

The peer-failure notice classified platforms with its own {"cli","tui","desktop"}
set, so an ACP session read the gateway wording ("do not suggest peerName"); it
now uses agent.coding_context.INTERACTIVE_CODING_PLATFORMS, which includes acp.

Three getattr/try-except guards existed only for test doubles (a bare __new__
provider, a SimpleNamespace config); the real types always carry the attribute.
Removed, and the two tests build real objects. An unset timeout resolves to the
client's 30s default, so the join-budget test expects 30, not the 5s floor.

Timing tests keep the >= 2s wall-clock bound the testing rules ask for.
2026-09-13 19:05:39 +05:30
Erosika
8523402db0 fix(honcho): bound the shutdown flush by the shutdown deadline
Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
2026-09-13 19:05:39 +05:30
Erosika
e9396ed8f4 fix(honcho): register the recall sync thread with its owner so shutdown joins it
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
2026-09-13 19:05:39 +05:30
Erosika
c3a5649aac fix(honcho): join every plugin thread on shutdown within one budget
provider shutdown joined the dialectic, sync and memwrite threads for 5s
each and then the async writer, but the session-init thread,
honcho-base-first and honcho-context-prefetch were never joined. any of
them still blocked in httpx when the interpreter finalized aborted the
process with SIGABRT 134 (#37632, #60616, #33485). the 5s join was also
shorter than the 30s http timeout a blocked call can hold (#33485).

spawn_context_thread now registers each thread under its owner (the
provider or the manager) in a weak registry, and shutdown joins every live
thread of both owners inside one deadline: at least 5s, or the configured
http timeout when longer. the manager refuses new prefetch threads once
shutdown began and flushes a late save() inline instead of respawning the
writer. threads that outlive the budget are named in a warning.

the sdk client exposes close() on its http pool. one client is shared by
every manager with the same identity in a gateway process, so a per-agent
shutdown cannot close it; close_honcho_clients() closes all pools and is
registered with atexit when the first client is built, the pattern the
hindsight and mem0 plugins use. follows #69070, #33701, #7627.
2026-09-13 19:05:39 +05:30