The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.
Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.
- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
With gateway.multiplex_profiles enabled, the primary Telegram message
handler is the closure returned by _make_default_profile_message_handler(),
so its __self__ is absent. The early intake filter
(_is_user_authorized_from_message) recovered the GatewayRunner via
self._message_handler.__self__ and, finding none, fell back to env-only
authorization — never evaluating the configured chat allowlist through
GatewayRunner._is_user_authorized(). Every non-global sender was then
default-denied in an explicitly allowlisted group.
Prefer the platform-bound authorization callback registered via
set_authorization_check(): it routes through the runner's full auth chain
(platform + group allowlists, pairing store, allow-all) and survives the
closure wrapping, whereas the bound-handler lookup does not. The bound
handler remains the fallback for setups without a registered callback, and
the pairing-passthrough guard for unknown DMs is preserved.
Fixes#87132
The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".
The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.
Changes:
- plugins/observability/nemo_relay: the wrapper always passes
timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
costs one telemetry span, never the agent; it also sets scope_errored
(so close_session skips the ATIF export for the wedged session) and
warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
reaches the host (fails on the pre-fix code — sabotage-verified),
a TimeoutError flags the session and disables its export, and the
generic error path keeps its scope_errored contract.
Session-finalize hooks ran synchronously on the gateway event loop from
three call sites (shutdown drain, session-expiry watcher, /new reset).
A plugin hook doing heavy blocking work froze the whole loop: adapter
heartbeats stopped, the drain machinery could not run, and systemd
eventually SIGKILLed the process mid-export. Observed live on a
multi-day 4.7G session where the nemo_relay observability plugin
serialized a full-session ATIF trace inside on_session_finalize.
Changes:
- gateway/run.py: new GatewayRunner._finalize_session_off_loop()
dispatches hermes_cli.lifecycle.finalize_session via the gateway
executor under asyncio.wait_for (10s budget), mirroring
_cleanup_agent_resources_off_loop (#53175). Shutdown finalize and
the session-expiry watcher now use it.
- gateway/slash_commands.py: /new reset path uses the same helper.
- plugins/observability/nemo_relay: ATIF export is now bounded
(HERMES_NEMO_RELAY_ATIF_EXPORT_TIMEOUT_S, default 30s) and skipped
entirely for sessions whose Relay scope operations already errored
(their exporter state is unreliable and the export can be
pathologically slow).
- tests/gateway/test_finalize_session_off_loop.py: regression tests
proving the loop stays live under a wedged hook and the budget is
enforced.
Migrates 8 of 12 native dialog call sites in the kanban dashboard plugin
to the SDK's ConfirmDialog primitive (added in PR #50550):
- moveTask, moveSelected, applyBulk, deleteTask, deleteSelected,
archiveBoard, removeAttachment, doPatch
The 4 remaining carve-outs (window.prompt for completion summary,
window.alert for missing summary, cli_hint clipboard fallback) are
documented inline — the host's ConfirmDialog hardcodes onClick → unmount,
preventing the keep-open-across-validation behavior the completion-summary
form needs. Followup: upstream a `disabled` prop to ConfirmDialog and
rebuild the completion body using host Dialog components.
New architecture:
- useKanbanDialogs(t) — Promise-based dialog state machine at
KanbanPage scope. request({kind, ...}) returns {confirmed, summary?}.
- KanbanDialog component — renders ConfirmDialog from SDK for kind=confirm.
- performMoveTask(taskId, newStatus, count, summary) — extracted shared
dispatch path for single + bulk moves (optimistic UI + PATCH/POST +
error recovery).
- requestDialog prop threading — KanbanPage → BoardSwitcher,
TaskDrawer → TaskDetail → doPatch/AttachmentsSection. Every call
site has a defensive fallback to window.confirm if the prop is
missing (verified by test_dashboard_done_actions_prompt_for_completion_summary
counting the cancel guards + destructive:true markers in the bundle).
New host i18n keys (web/src/i18n/en.ts + types.ts):
- kanban.confirmDoneMany / confirmArchiveMany / confirmBlockedMany
- kanban.trash.confirmTitle / confirmManyTitle
Tests:
- Replaced bundle-string-only completion-summary test with behavioral
coverage: bundle cancel-guard count + destructive marker count, plus
backend tests that confirm cancel preserves old status and confirm
dispatches the expected PATCH/DELETE body.
- Removed the SDK_CONTRACT_VERSION snapshot test from
web/src/plugins/registry.test.ts (forbidden by AGENTS.md
"Don't write change-detector tests"; the two remaining tests in that
file already cover the new SDK surface behaviorally).
Closes#50547 (consumers of #50550).
Cross-vendor re-review: Gemini 3.5 Flash + GPT-OSS 120B (both SHOULD-FIX,
no remaining BLOCKERs after these fixes).
Closes#77833
stream_events() only detected client disconnect via send_json()
raising WebSocketDisconnect. When no events were pending (idle
board), send_json was never called, so the poll loop ran forever
even after the client disconnected — leaking one poll task per
disconnect.
Fix: race ws.receive() with a timeout matching the poll interval.
On timeout, poll the DB as before. On disconnect, exit cleanly.
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.
Re-anchored accordingly:
- Discovery-time pre-registration, module reuse, and the `provides_tools`
opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
since those tools registered before `registration_start` and the slice
cannot see them.
- A failed materialization no longer carries attribution across. The
failure path now sweeps the whole ownership ledger for the plugin key,
not just the `registration_start:` slice, so the pre-registered tools
are disposed along with the adapter. Attribution and the registry now
agree at zero instead of reporting tools the process is not serving.
tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
On macOS (256 soft fd limit), routing the weixin/email pollers through a
local HTTP proxy leaked one TCP socket per failed poll/connect cycle
until the gateway hit `[Errno 24] Too many open files` and crashed
(launchd respawn loop). Live capture showed 216 of 256 fds pinned on
connections to the proxy, ~214 of them abandoned.
Code-side gaps fixed:
- email adapter, `connect()`: no try/finally around the IMAP test
connection — a failure in login/ID/select/search abandoned the
connected socket with no owner. Every reconnect-watcher retry builds
a fresh adapter, so each retry against an unreachable/proxied host
leaked another fd. Teardown now runs in `finally`.
- email adapter, IMAP teardown: `imaplib.IMAP4.logout()` only swallows
`OSError` internally; on a broken connection `LOGOUT` raises
`IMAP4.abort` before the internal `shutdown()`, leaving the socket
open. New `_close_imap()` helper chases a failed `logout()` with an
unconditional `shutdown()`; used in `connect()` and
`_fetch_new_messages()`.
- weixin adapter: repeated poll failures through a proxy strand
sockets in the aiohttp connector where the tight keepalive reaper
never sees them. The poll loop now recycles its ClientSession
(swap-then-close, safe for concurrent `_process_message` tasks)
after each MAX_CONSECUTIVE_FAILURES streak, tearing down the
connector and every socket it holds.
Targeted tests: tests/gateway/test_poller_fd_lifecycle.py (9 tests).
Reported by @EthanHunter1229 with measured fd captures.
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
Pre-existing inconsistency: _flush_media_group_event used return
in CancelledError while _flush_text_batch and _flush_photo_batch
used raise. Changed to raise for consistency and to properly
propagate task cancellation. Made more visible by the hold-queue
changes in #83878.
Address review on #83878:
- Permanent fatal fences all hold producers and discards pending maps on
teardown instead of re-populating a queue that can never drain.
- Any hold created while connected schedules a tracked redispatch (cancel-
after-pop no longer orphans until a future reconnect).
- Redispatch failures re-hold current + remainder without tight-looping.
Regression coverage for the three residual paths, plus the interaction
with OOF-156's connect-failure classification: the retryable network
path (telegram_connect_error) must NOT clear the hold queue — reconnect
is precisely what drains it; only non-retryable fatals discard.
The disconnect drop-guard (#55971) correctly prevents dispatch into a
torn-down session. Destroying the event was wrong: by enqueue/flush time
python-telegram-bot has already acked the update and advanced the polling
offset, so Telegram never redelivers. Result: silent permanent loss, no
log, no error.
Hold inbound events (text/photo/media-group) when the drop-guard fires,
salvage pending batch maps on teardown, cancel+await the redispatch task
in the delivery cancel map (lifecycle-tracked), and redispatch from
_mark_connected after reconnect. Cap the hold queue (default 64), dedupe
by object identity, discard on non-retryable fatal. Cancel-after-pop in
flush paths also holds.
Distinct from #72037 (cancel-after-pop during follow-up supersession) and
#81528 (boundary discard). Tests use delay=0 and entered/release Events —
no wall-clock races; includes production terminal-step coverage.
ActualProfile.fetch_models() overrides ProviderProfile's default
implementation with its own Actual-specific base_url resolution
(ACTUAL_BASE_URL env var, hosted-vs-local normalization), but called raw
urllib.request.urlopen(req, timeout=timeout) directly instead of the base
class's open_credentialed_url(). Every other provider either uses the
base class default or forwards to it via super() and gets
SafeCredentialRedirectHandler for free — Actual is the only provider that
attaches a Bearer token to its own Request object and opens it with the
stdlib's default redirect handling, which forwards every header,
including Authorization, across a cross-origin redirect.
Actual's own feature surface makes the trigger realistic: ACTUAL_BASE_URL
is a first-class, documented way to point this provider at a self-hosted
or local-offline endpoint (see the local-loopback no-auth path already
handled elsewhere in this provider), so a misconfigured or compromised
endpoint 302-ing to another host leaks ACTUAL_API_KEY to it.
Fix: import and call the same open_credentialed_url() the base class
uses, keeping Actual's own URL-resolution logic unchanged.
Adds an end-to-end regression test using two real local HTTP servers (no
mocking of the security module itself) — one redirects, the other
records the Authorization header it receives — mirroring
test_urllib_security.py's own redirect tests. Also repoints the existing
fetch_models test's mock from urllib.request.urlopen to
hermes_cli.urllib_security.open_credentialed_url, since fetch_models no
longer calls the former. Mutation-verified: the new redirect test fails
on pre-fix code with the Authorization header observed at the redirect
target.
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.
Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
Builds on the three salvaged commits: adds the sources and integration points
they leave out, so a pip-installed memory provider is not a second-class
citizen next to a directory install.
Discovery
- Project-local providers (./.hermes/plugins/<name>/), gated on
HERMES_ENABLE_PROJECT_PLUGINS exactly as PluginManager gates its own project
scan. Completes the four sources CONTRIBUTING.md and AGENTS.md already
promised; memory was the only discovery system missing two of them.
- find_provider_dir() now resolves a package entry point to its directory.
This is load-bearing: config_schema.py (the dashboard panel) and cli.py (the
`hermes <provider>` subcommands) are read from disk rather than imported, so
without a directory a pip-installed provider silently lost both.
- list_memory_provider_names() includes entry-point providers, so they appear
in the dashboard's memory.provider dropdown.
Resolution stays import-free. hermes_cli.plugins.resolve_module_origin() is
extracted from _resolve_module_source() (added by the salvaged #76567) and
shared, so discovery walks a module's file layout instead of importing it.
find_provider_dir() is called from the dashboard and from argparse setup, long
before the operator has chosen a provider — importing every installed candidate
would execute third-party code on the strength of a package being present.
A test asserts the resolution leaves no side effects and no sys.modules entry.
Registration
- PluginContext gains register_memory_provider(). Memory was the only provider
category without one; context engine, image gen, video gen, web search,
browser, TTS, transcription, secret source, dashboard auth and platform all
have one.
- _ProviderCollector delegates unknown register_* calls to a real
PluginContext instead of carrying three hand-written no-ops. It silently
dropped register_tool/register_hook, and had no register_auxiliary_task at
all — despite PluginContext.register_auxiliary_task documenting a memory
provider (hindsight's pre-retain dedup) as its worked example. It can no
longer drift behind PluginContext.
- A raise after register_memory_provider() no longer costs the provider. The
loader caught it into a debug log, discarded the registered instance, and
fell through to "instantiate any MemoryProvider subclass" — returning a
different, unconfigured provider. A silent downgrade that looked like
success, and the exact outcome of calling register_auxiliary_task.
Activation is unchanged: still gated on memory.provider naming the plugin, and
covered by a test so the real PluginContext cannot start requiring
plugins.enabled — that would break every existing user-installed provider.
Verified end to end against a real third-party provider (kainappsinc/elephant)
installed by pip alone, with no directory copy: it appears in the dropdown,
resolves its directory, loads with its tools, and renders its dashboard panel.
Closes#40101.
Slack session keys include the workspace id since #70190, but the kanban
notifier rebuilds the wake source from a subscription row that has no scope
column, so every terminal-event wake keyed without the workspace.
The legacy-key adoption shipped in the same change (`_legacy_slack_session_key`,
`_recovered_row_matches_source_scope`) resolves that unscoped key onto the same
session_id, so the wake passes the busy guards that are keyed by routing key
(`_active_sessions`, `_running_agents`) and only collides afterwards, on session
id, under the per-session turn lease (#64934) — which serializes it behind the
live turn's flush. On a live Slack gateway that shows up as a duplicate run on
one task plus 400+s of waiting before the woken turn starts.
Same failure mode as #56580 / #72191 (chat_type), one field over, and it needs
no schema change: `_thread_metadata_for_source()` already stamps
`slack_team_id`, the notify subscription persists that dict as
`delivery_metadata`, and the notifier already unpacks it. Rows written by
`kanban_tools._maybe_auto_subscribe` carry no workspace, so fall back to the
adapter's channel → workspace map via `scope_id_for_chat()`, read with getattr
so adapters opt in and unscoped platforms' keys stay byte-identical. Slack
answers it from `_remember_channel_team`, which drops channels claimed by two
workspaces, so an unknown or ambiguous channel degrades to today's behavior
instead of guessing wrong.
Also adds the contributor email mapping the attribution check requires.
Co-authored-by: Junie <junie@jetbrains.com>
Swaps the Google flash entry in the OpenRouter and Nous Portal curated
lists to the newly released gemini-3.7-flash (half the price of
3.6-flash: $0.375/M in, $1.875/M out per OpenRouter live metadata;
served on both endpoints, verified live). Also updates the OpenRouter
plugin fallback_models mirror and regenerates model-catalog.json.
Scoped to the two named providers: vertex/gemini/gmi curated lists and
aux defaults still carry 3.6-flash.
on_memory_write spawns a fire-and-forget daemon thread that was never
stored on self, so shutdown() couldn't join it — the exact problem the
PR fixes for the async writer thread. Store as self._memwrite_thread
and include it in the shutdown join loop.
Review follow-up for salvaged PR #83500.
The previous gate compared session.user_peer_id against a fresh
_resolve_user_peer_id() call on the same manager. Both values come from
the same resolver with the same inputs, so a non-owner triggering a new
session in a shared channel passed the check and received the owner's
MEMORY.md/USER.md under their peer.
The owner is now a config fact: _declared_owner_peer_id() returns the
sanitized peerName, and migration runs only when the session's user peer
is that peer. Without a declared peerName, migration runs only when no
runtime gateway identity is present (the single-operator CLI path).
Aliases still work: a platform ID mapped onto peerName resolves to the
owner peer before the comparison.
Tests now derive each session's user peer from the real resolver instead
of hand-picking mismatched ids, so the non-owner test fails against the
old gate.
The owner gate from #82038 compared against config.peer_name directly,
which is None for most single-user setups — sanitizing None would raise
and the gate never accounted for pinned/runtime/aliased identities.
Resolve the owner the same way sessions do, and add the non-owner skip
regression test the original PR shipped without.
Co-authored-by: menhguin <menhguin@users.noreply.github.com>
migrate_memory_files() uploads USER.md/MEMORY.md with peer=user_peer — the
session's runtime user. In shared channels, a non-owner's new thread uploads
the owner's full profile under the NON-OWNER's peer; Honcho's deriver then
attributes the owner's psychometrics/medical/biography to that person. This
was the root contamination vector (55/70 contaminated sessions carried the
payload). Skip migration unless the session user is the configured owner.
SOUL.md unaffected (uploads under assistant peer).
Co-authored-by: Minh Nguyen <menhguin@users.noreply.github.com>
sync_turn called manager._flush_session() directly, which flushes
synchronously every turn no matter what writeFrequency says — the
"async", "session", and every-N-turns modes were dead configuration
on the main turn path. Route through save(), the dispatcher that
actually implements those modes.
Same bug class reported in #19650 (starship-s) and #72708 (Diaspar4u);
this takes the minimal one-line routing fix without their broader
lifecycle refactors.
Co-authored-by: starship-s <45587122+starship-s@users.noreply.github.com>
Provider shutdown() only called manager.flush_all(), which drains the
queue but never joins the async-writer thread — manager.shutdown()
exists and nothing called it. The writer thread could still be blocked
in httpx I/O at interpreter exit (the #37632 crash class). Now
shutdown() calls manager.shutdown() (flush + join) when persistence is
enabled, and a new manager.stop_async_writer() (join only, no flush)
when saveMessages is false, so containment and clean teardown compose.
The containment commit skipped the whole turn when either side was
empty, which would drop a real user message on interrupted or
tool-only turns. Keep the guard for fully-empty turns only and skip
empty sides individually inside the sync loop.
Salvages #67559 — original gated sync_turn/on_memory_write/on_session_end but missed shutdown(), whose flush_all() still persisted on exit. hermes-sweeper review (salvageability=high) flagged this as the one gap.
Guard sits after the worker-thread joins, not at the top: cleanup is independent of persistence, and a top-of-method return would leak _prefetch_thread/_sync_thread. Adds TestShutdown and clarifies the saveMessages=false README row.
Credit @Matroskin86 (original PR author).
The saveMessages knob has been parsed by HonchoClientConfig since its
introduction but was never consumed: sync_turn, on_memory_write and
on_session_end persisted to Honcho regardless. With saveMessages=false the
provider now never writes automatically (raw turns, memory-write conclusion
mirroring, session-end flush) while read/tools paths stay fully functional.
Guard uses getattr with a True default so legacy/injected configs keep the
old behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018giroL5zeMPnPxERxAxXHY
dialectic_query collapsed every backend failure to an empty string, so
the explicit honcho_reasoning tool rendered timeouts, server errors,
and genuinely-empty answers identically as 'No result from Honcho.'
(#36098 issue 4). Operators debugging 'search works but reasoning does
not' were sent down representation/observation rabbit holes when the
real cause was a 30s timeout on a medium-reasoning dialectic call.
Add raise_errors to dialectic_query (default false — automatic
injection keeps its fail-quiet behavior and cadence backoff) and pass
it from the explicit tool call, returning a tool error that names the
failure and points at the timeout knob. Auth errors keep their
dedicated handler.
Two silent-auth-failure paths from #36098 (also #66125):
- the local-URL guard only escaped the 'local' placeholder when the
HOST BLOCK had apiKey. A top-level apiKey in honcho.json — explicit
user intent, and what 'hermes honcho setup' writes for single-host
configs — was dropped on the floor, so AUTH_USE_AUTH self-hosts
401'd on every request. Now any explicit key in honcho.json (host
block or top level) is honored; only env-sourced keys are still
treated as likely-cloud and skipped for local URLs.
- named-profile host blocks do not inherit the default host's apiKey
(credential isolation is by design), but the failure was silent:
the profile ran unauthenticated and every tool said 'no context'.
Affirm isolation and warn loudly at config-resolution time instead,
the outcome #66125 proposed if inheritance is rejected.
_all_profile_host_configs() built per-profile host keys inline as
f"{HOST}.{profile}" ("hermes.work") while profile_host_key() — used by
honcho status/enable/sync and the runtime memory plugin — produces the
underscore form ("hermes_work"). The lookup always missed, so
'hermes honcho peers' showed "(not set)" / leaked the raw malformed key
into the AI-peer column for every non-default profile. Profile names
needing sanitization (dots/spaces) were doubly broken.
Verified live: with hosts["hermes_work"] populated, cmd_peers showed
'work ... hermes.work' before the fix and 'work ... hermes' after.
Tests: host keys match the writer form, sanitized profile names resolve,
peers output shows populated identities with no key leak, and clean
fallback for profiles without a block.
Salvage of #2757 by @teyrebaz33 — rebased onto current Honcho plugin layout.
Stray control characters (e.g. terminal escapes pasted into HONCHO_BASE_URL
or config baseUrl) are dropped with a warning so SDK construction cannot
crash startup on Invalid non-printable ASCII character errors.
_resolve_or_create_client() used a plain dict.get(config.host) that
fails for dot-form profile host keys (e.g. "hermes.profile_a") even
though the _host_block() helper defined nearby handles the legacy
dot-form → underscore-form fallback correctly. The result:
_host_has_key evaluates to False for every authenticating user,
so effective_api_key is set to "local" and every Honcho API call
returns 401 Invalid JWT — cascade failure into silent data loss
for cross-peer queries and message sync.
Fixes by calling the existing _host_block() helper instead of
reimplementing the direct lookup. Local variable renamed from
_host_block → _host_block_local to avoid shadowing the function.
Closes#37436
HonchoClientConfig.from_global_config() only consulted top-level
baseUrl / base_url / HONCHO_BASE_URL in ~/.honcho/config.json. The
Honcho SDK's native config format — and what Claude Desktop writes —
nests the URL at endpoint.baseUrl. Users with that config format had
their self-hosted Honcho container silently ignored: every honcho_*
call routed to https://api.honcho.dev with a workspace_id that does not
exist there, so tools returned empty data with no error anywhere.
Resolution order in from_global_config(), highest first:
1. endpoint.baseUrl (SDK-native, what Claude Desktop writes)
2. baseUrl / base_url (root-level, existing behavior)
3. HONCHO_BASE_URL (existing env var)
4. HONCHO_URL (the SDK's own env var, honcho/client.py:234)
HONCHO_URL is also read in from_env(). from_global_config() delegates to
from_env() whenever the config file is missing or unreadable, so an env
fallback wired into only one of the two would silently do nothing for
users with no config file.
A non-dict endpoint value falls through cleanly rather than raising.
Existing users are unaffected — the new sources are consulted only when
the existing ones resolve to None.
The INFO log for the base_url-unset case now says so explicitly instead
of printing only the host. The SDK resolves that case from its own
ENVIRONMENTS map (honcho/client.py:36-39), which for environment=
production means the public cloud; a self-hosted user whose config was
not picked up otherwise sees a healthy-looking startup line.
Closes#43800.
- _submit_background and _prefetch_provider: replace unreadable
(lambda inner: (lambda: ctx.run(inner)))(fn) with functools.partial(ctx.run, fn)
- from_env(): set config_path=resolve_config_path() so bound_config_path()
doesn't re-resolve from ContextVar on daemon threads (the exact bug
the PR fixes for from_global_config)
Review follow-ups for salvaged PR #83525.
The dict was written on every build and popped/cleared on eviction and
reset, but no read site remained — timeout staleness detection moved
into the cache key itself (a timeout change produces a new identity and
_slot_for evicts the old slot), which the isolation tests already pin.
Flagged in review by @spfcraze.
Profile isolation is a ContextVar; plain threading.Thread targets start
with an empty context, so the plugin's nine daemon threads (session
init, prewarm, first-turn base/prefetch, prefetch, sync, memwrite,
async writer, context prefetch) resolved ambient state — config path,
active host, hermes home, oauth token paths — against the DEFAULT
profile whenever they ran under a routed profile's turn.
Adds spawn_context_thread(), which copies the caller's context at spawn
time so the thread sees the profile scope it was created under, and
routes every plugin thread spawn through it. Defense-in-depth under the
bound-config work: even ambient resolution on these threads now lands
on the right profile.
The copy_context approach follows the gateway's own
_run_in_executor_with_context pattern; #81401 applied it to the init
thread, this extends it to all nine spawns.
Co-authored-by: angel12 <angel12@users.noreply.github.com>
Replaces the process-wide first-config-wins client singleton with a
per-identity slot map. The singleton baked the first profile's
workspace_id and bearer into one shared client, so in multi-profile
processes (gateway multiplexer, dashboard, cron) every profile's
memory landed in whichever workspace initialized first — cross-tenant
bleed with no error (#69123, #74065).
cache key: (host, workspace, base_url, environment, provenance paths,
effective timeout, credential fingerprint). the fingerprint hashes the
OAuth REFRESH token (stable across in-place access-token rotation,
changes on re-auth/account switch) or the static api key — so
re-running 'hermes honcho setup' to switch accounts produces a new
identity instead of silently reusing the old account's client and
writing tenant B's data with tenant A's bearer, a hole per-path keys
alone cannot close.
same-identity slots with a different fingerprint or timeout are
EVICTED on replacement, so credential churn can't accumulate pinned
clients — the replaced client's pools close when its last holder
drops. timeout changes rebuild via the key (the old explicit staleness
check is subsumed). failed in-place OAuth rotation resets only the
client's own slot. reset_honcho_client() clears everything, preserving
test and oauth-flow re-login semantics.
per-config-identity caching was first proposed in #69142; the
provenance-key shape follows #81401. this implementation adds the
credential fingerprint and eviction they lacked.
Co-authored-by: NaMinhyeok <NaMinhyeok@users.noreply.github.com>
Co-authored-by: angel12 <angel12@users.noreply.github.com>
Profile isolation in every multi-profile process (gateway multiplexer,
dashboard, cron) is a ContextVar (set_hermes_home_override) that
threading.Thread targets cannot see. The plugin's daemon threads —
async writer, prefetch, sync, first-turn, init — all funnel through
HonchoSessionManager.honcho, which called get_honcho_client() with NO
config, re-resolving resolve_config_path()/resolve_active_host() from
the ContextVar-blind thread context: every background memory access
landed on the DEFAULT profile. Worse, the OAuth paths did the same, so
a token refresh on a daemon thread could persist the rotated token
into the wrong profile's honcho.json, and a 401 recovery could burn
the wrong profile's single-use refresh token.
- HonchoClientConfig gains provenance (config_path, hermes_home)
captured at resolution time inside the caller's profile scope, with
bound_config_path() for consumers
- manager.honcho passes the bound config instead of re-resolving
- OAuth paths (_apply_fresh_oauth_token, _refresh_cached_oauth,
_reauth_required, _force_reauth) use the bound path
- the honcho.json timeout memo becomes path-keyed instead of
single-slot, so multi-profile processes stop thrashing it and
returning profile A's timeout for profile B
Groundwork for per-identity client caching (#69123, #74065); the
provenance-field shape follows #81401.
Co-authored-by: angel12 <angel12@users.noreply.github.com>
Adds platforms.slack.extra.native_task_cards: when enabled, live tool
calls render as Slack-native plan/task cards via chat.startStream /
chat.appendStream (task_display_mode: plan, task_update chunks) instead
of text/edit progress bubbles. ID-bearing tool_start/tool_complete
callbacks correlate concurrent same-name tool calls correctly; any
native API failure falls back to one continuously edited text update.
The stream is stopped exactly once when the turn finalizes.
Salvaged from PR #29496 onto current main (TurnRunner/TurnContext seam);
closes#29483.
Slack's Agents & AI Apps feature ships a native streaming surface that
renders a live-typing message instead of the edit-based progressive
updates the adapter used until now.
The adapter now implements the existing draft-streaming interface:
- supports_draft_streaming() opts in whenever the app is connected and
native streaming hasn't been detected as unavailable.
- send_draft() starts a stream on the first frame (chat.startStream,
anchored to the resolved thread_ts, with recipient_team_id/user_id
for channel streams) and appends only the delta on subsequent frames
(chat.appendStream is append-only). The consumer's trailing cursor
glyph is stripped before delta computation.
- Unlike Telegram drafts (ephemeral, replaced by a real sendMessage),
a Slack stream IS the final message. send() therefore intercepts the
turn-final delivery for a chat with an active stream whose streamed
text is a prefix of the final content, and seals it via
chat.stopStream with the remaining delta instead of posting a
duplicate. Rich Block Kit (when enabled) is applied to the sealed
message via chat_update, mirroring the finalize path in edit_message.
- Feature-gate errors from chat.startStream (not_allowed,
missing_scope, unknown_method, ...) are cached on the adapter so
subsequent runs skip straight to edit-based streaming with a single
warning naming the fix (enable Agents & AI Apps for the app);
transient errors only disable drafts for the current run via the
consumer's existing send_draft failure handling.
- Segment breaks (new draft_id) and disconnect() seal any open stream
so chats are never left with a dangling live-typing indicator.
No consumer or config changes: streaming.transport auto/draft now
lights up native streaming on Slack through the same interface
Telegram drafts use, and the edit-based path remains the fallback.
_resolve_thread_id() falls back to _last_inbound_thread[chat_id] when no
explicit thread is present. That fallback exists for interactive DMs, where
Google Chat spawns a fresh thread per top-level user message and the adapter
drops thread_id to keep the session key stable. It also fired for cron
deliveries, which carry job_id in their metadata but no thread: the output
landed as a reply inside the last inbound thread instead of starting a new
top-level message.
Bypass the _last_inbound_thread fallback when metadata has a job_id (i.e. the
message is an automated cron delivery), so cron output posts at top level
unless an explicit thread is requested.
1. Use open_credentialed_url() instead of bare urlopen() in
templates.py apply_template() and probe_existing_customization().
Both send Authorization: Bearer headers; bare urlopen forwards
credentials on cross-origin redirects. The codebase has
open_credentialed_url() in hermes_cli/urllib_security.py that
strips credentials on cross-origin redirects — used by 4 other
modules.
2. Guard unavailable_reason() with the dedup set check before
calling it. The gateway builds a fresh AIAgent per message, so
without this guard unavailable_reason() (which calls _load_config()
→ stat + file read + JSON parse, and _check_local_runtime() →
importlib probes) runs on every gateway turn for an unavailable
provider, even though the warning is deduped after the first.
3. Move INDICATOR_GLYPH from Hindsight's eye emoji to a generic
brain (🧠) in core (agent/memory_provider.py). Hindsight overrides
with its own _HINDSIGHT_GLYPH (👁️) in recall_status() and
_emit_saving_indicator(). Other memory providers no longer inherit
Hindsight's brand mark as the default glyph.
Bundles previously-separate Hindsight/memory PRs into a single review surface:
- opt-in synchronous recall (recall_sync) — recall the injected memory in-turn instead of next-turn prefetch (#5820)
- actionable error when local_embedded runtime is missing — tells the user which package to install (#7718)
- default retain_source to 'hermes' so every stored memory self-identifies its provenance
- offer a starter memory template during hermes memory setup, plus warn before overwriting an already-configured bank
- warn when a configured memory provider reports unavailable (#2765)
- deterministic 'recalled N memories' recall indicator — Hermes itself emits a status line when auto-recall injects memory
- 'saving to memory' retain indicator — emitted the moment a turn is dispatched to the writer
Authored by @benfrank241 (ben.bartholomew@vectorize.io).
Salvaged from PR #74379.