a74e0155b6 made attachments[].blocks[] reach the agent through
_append_link_unfurls, but rendered each attachment's blocks with no ceiling.
Slack allows 20 attachments per message, so one alert could project 20x what
a single attachment does (measured: 3,247 chars for 1 -> 64,855 for 20 with
8x400-char rich_text sections each), while the top-level blocks path caps
once at 6000.
Share one budget (_SLACK_UNFURL_BLOCKS_MAX_CHARS, the same 6000 the top-level
path uses) across the array: the first attachment keeps its body, later ones
are truncated against the remainder, and a spent budget still leaves every
header visible. After: 3,247 -> 6,659 chars at 20 attachments.
(cherry picked from commit 3efc1e4c532c38a220d4c6ea53745d59d334aa3e)
Telegram's per-(chat_id, status_key) status-message cache grew without
bound; give it the same _STATUS_MESSAGE_IDS_MAX=2000 FIFO half-trim the
Slack adapter already has. In both adapters, guard the post-await write-back
after a successful edit with a compare-before-write (only re-store the id if
the cached entry is still the one we edited) so an eviction or replacement
that happened during the await is not undone.
Partial salvage of #87480: kept the Telegram bound and both compare-before-write
guards (re-applied by hand, 17586 behind), defined the max as a class attr
like Slack instead of an instance attr, dropped the 4 new tests.
(cherry picked from commit 43ae95e98e)
_handle_events called _save_cursors inline whenever a batch moved the
cursor. It ends in atomic_json_write (mkstemp + fsync + os.replace), and
on the WebSocket transport _handle_events runs once per inbound EVENT
frame, so every message paid an fsync on the gateway's event loop,
stalling every other adapter and in-flight turn for its duration.
Split the snapshot from the write: the payload is still built on the
loop (_channel_state is loop-owned), and the write goes through
asyncio.to_thread. The snapshot is taken under an asyncio.Lock so a
slower, older write can never land after a newer one and regress the
durable cursor. connect() keeps the synchronous _save_cursors.
(cherry picked from commit 299929741429262f497b86a67940ffb37faa1694)
utils.fast_safe_load already exists, is pinned by tests/test_fast_safe_load.py, and
its comment names exactly these payers: 'startup parses config.yaml and every plugin
manifest, so the slow path cost ~0.9 s of cold start'. The migration was started —
hermes_cli/config.py uses it eight times, hermes_cli/main.py and hermes_cli/plugins.py
too — but the file-level loaders it was written for were never converted.
The cost is config SIZE, and the size is the installer's doing: it seeds config.yaml
by copying cli-config.yaml.example, 120,897 bytes of mostly comments. Nothing caches
load_gateway_config() and it has 238 production call sites.
Profiled before assuming a cause — reader.forward 34 ms, scanner.scan_to_next_token
28 ms, reader.peek 13 ms: the pure-Python PyYAML scanner, nothing else.
Measured A/B on one realistic pass (gateway config + every bundled plugin
description), medians of five runs, __pycache__ cleared between arms:
48.8 ms -> 2.46 ms. Per path: load_gateway_config 48.3 -> 1.84 ms, managed
config.yaml 44.4 -> 1.03 ms, 105 plugin.yaml manifests 59.3 -> 5.9 ms.
Same parse, same restricted tag set, same result — only the loader changes. Drops the
three 'import yaml' statements the swap orphaned.
(cherry picked from commit a4efd6a4506043159310eba02c32b673c455b88b)
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
Sibling sites of #119986's class: the openai and meta-ai image backends
resolved their API key through get_secret but the base URL through
os.environ, so on a multiplexed gateway a routed profile's key was sent to
the launch profile's endpoint. Both fields now come from the same scope
(get_secret_str, like the DeepInfra video backend after #119986).
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:
- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
the credential pool, not the environment. Chat and image_gen/openrouter find
it through resolve_runtime_provider. video_gen reported OpenRouter
unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
routed profile's video jobs were submitted, polled and downloaded with the
launch profile's key and billed to that account. A profile whose key lived
only in its own .env could not use the backend at all.
The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.
OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
cached_sdk_client returned the client cached on tools.web_tools before it
read the key, so the Exa, Parallel and AsyncParallel clients kept the key
they were first built with for the life of the process. A key fixed in .env
and applied with /reload still sent the old key (401s until a restart), and
on a gateway serving multiplexed profiles every profile's Exa and Parallel
calls went out on whichever profile's key built the client first, billed to
that account. A key removed from the environment also kept being used.
Resolve the key on every call and reuse the cached client only when it was
built with that key; the slot now holds (key, client) as one value so two
builds racing under different keys cannot record one key beside the other
key's client. Firecrawl already compares its credential before reusing its
client; this brings the two SDK-backed providers in line.
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.
The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.
Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
decline) must authenticate its From:, open access or not: the
pairing code or refusal is mailed back to that address. A granted
sender still needs it short of open access, since a pairing grant
keys on From: just as the allowlist does. Open access follows the
gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
longer exempts a listed address from From: authentication either
(on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
list grants a stranger nothing there, so the env flag alone no
longer exempts one from From: authentication (that path mailed a
pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
dropped. The gateway's check also matches an address by its bare
local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
(a chat username) would admit or pair alice@<any domain>. The lists
are parsed as the gateway parses them, JSON list literals included,
or '["alice"]' would slip past this guard.
_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.
Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
went from 0 events reaching the gateway to 1. pair mails a pairing
code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
stranger@<domain> through, under ignore or pair. Without the
local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
with a forged From: under allow-all beside an EMAIL_ or
GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.
Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.
Refs #68502
Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>
The setup schema offered db_path as f"{display_hermes_home()}/memory_store.db",
so `hermes memory setup` (Enter on the field) and the dashboard form wrote the
active profile's concrete path into plugins.hermes-memory-store. initialize()
only expands a literal $HERMES_HOME, and neither `profile create --clone` nor
`profile rename` rewrites config.yaml, so:
- a cloned profile opened the source profile's memory_store.db: facts stored
in one profile were recalled into the other's prompts, both ways;
- a renamed profile opened profiles/<old>/memory_store.db, which MemoryStore
re-created as an empty DB in a ghost directory under the old name, while the
real facts sat orphaned in the renamed directory.
The schema default is now "$HERMES_HOME/memory_store.db", the value the docs
already give, which initialize() resolves against whichever profile opens it.
save_config also stores a db_path equal to this profile's own DB as the
placeholder, so a config written by an older setup is repaired on its next save
from the CLI or the dashboard. A path anywhere else is kept as given.
- Desktop drawer -> Linear-style modal (smaller than Settings): main column
holds diagnostics, description, result/summary, dependencies, comments,
activity, runs, and the worker log tail; a right property sidebar holds the
inline editors (assignee, model override) plus priority/tenant/workspace/
created rows, estimate, and attachments. Backdrop click or Esc closes.
- GET /tasks/:id gains 'link_tasks' ({id,title,status} per linked task) so
Blocks/Blocked By chips render titles instead of raw ids; older backends
fall back to short ids. Additive; 'links' shape unchanged.
- New backend tests (test_kanban_link_tasks.py) + drawer tests for title
chips and the id fallback.
Follow-up on the two salvaged commits:
- One _sidecar_payload_text() helper decides the /send text for both the adapter
path and _standalone_send (cron / send_message), which had the same raw-markdown
leak on URL-bearing messages and never stripped with PHOTON_MARKDOWN=false.
- BlueBubbles (same iMessage surface) keeps [label](url) targets as bare URLs.
- Tests trimmed to two invariants.
The gateway's tool-progress path emits terminal commands as fenced code
blocks on any adapter whose supports_code_blocks is True. The Photon
adapter set that flag from PHOTON_MARKDOWN, but the sidecar's /send
router (send-format.mjs) silently routes every URL-bearing message
through the plain-text builder — where fences survive as literal backtick
characters — and even the markdown path renders a fence as inline
monospace text, not a block. Net effect: raw fenced terminal commands
(and raw markdown around them) surfaced in iMessage bubbles no matter
what the user's prompt-level style rules said.
Fix the whole class: the adapter never claims code-block support, so the
gateway emits its compact one-line tool preview instead. Prose markdown
passthrough (bold/italic/headings) is unchanged, and no config change is
required — tool_progress can stay at its user's preferred mode.
Tests: new capability + E2E tests pin that the gateway cannot emit a
fence for Photon with real adapter + real display resolution; the old
supports_code_blocks-mirrors-env expectation (which pinned the buggy
behavior) now asserts the flag is always False.
chooseSendFormat() routes markdown containing a raw http(s) URL through
spectrum-ts' text() builder, because the markdown builder's iMessage data
detection 500s on those messages (#73615). text() ships the payload
verbatim, and nothing strips the markers on the way down, so any reply
that mentions a link arrives in iMessage as literal markdown source:
**Release 1.2.0** is out
- **EUR 5** off this month
https://example.com/releases/1.2.0
Every ** is visible in the bubble. Remove the URL from that same reply
and it renders correctly, which is what makes the URL the trigger rather
than the content.
spectrum-ts documents markdown() as degrading to readable plain text on
platforms without native support "instead of surfacing raw ** markers".
Selecting text() ourselves opts out of that guarantee, so the adapter has
to honour it instead.
Strip in _sidecar_send() when the payload is markdown and the sidecar will
downgrade it. The format key is deliberately preserved: the sidecar owns
the builder choice (test_rich_links.py pins that contract), and an older
sidecar without chooseSendFormat must keep rendering natively.
_send_plain_fallback() selects the text builder explicitly via
markdown=False and had the same leak, so it strips too.
Reuses the shared strip_markdown() helper rather than adding a second
implementation, so the PHOTON_MARKDOWN=false path and this path produce
identical output and both inherit any future fix to that helper.
Diagnosis previously reported in #85733, which was closed unmerged.
Stripping must not take the URL with it, though. The shared helper
collapsed [label](url) to label alone, which on this path is worse than
raw markdown: iMessage auto-links bare URLs and nothing else, so the
reply arrives with a description and no way to reach the link.
Open the itinerary on Google Flights <- URL gone entirely
So strip_markdown() takes keep_link_targets, which rewrites
[label](https://url) as "label\nurl" (own line, because these URLs are
often long) and leaves non-http targets such as mailto: or relative
paths label-only, since iMessage won't linkify those either. The default
is unchanged, so the SMS, IRC, Feishu and QQ callers keep dropping the
target as before.
plugins/platforms/line/adapter.py already carries a private
strip_markdown_preserving_urls() for exactly this reason ("LINE
auto-links bare URLs only"). This moves the behaviour behind the shared
helper instead, so Photon's three plain-text paths -- the downgrade,
_send_plain_fallback(), and PHOTON_MARKDOWN=false -- all agree.
gateway/platforms/bluebubbles.py is the same iMessage surface with the
same loss; left alone here to keep this change to one platform.
The Photon hint told the model "Markdown is rendered (bold, italics, lists,
code)", so replies came back with headers, code fences and backticks. That
is wrong on the real delivery path: the sidecar sends any message containing
a URL as raw text (every *, #, ``` and | shows literally), and even on the
markdown path spectrum-ts flattens headings to bold, tables to "a | b"
rows, and turns code into Unicode math-monospace glyphs that break when
copied. The hint also claimed attachments are metadata-only, which stopped
being true when native media send landed.
Both iMessage hints (Photon plugin, BlueBubbles built-in) now ask for a
texting register: short, answer first, no headers/tables/fences/backticks,
commands on their own plain line so they copy, bare URLs. BlueBubbles also
notes that strip_markdown drops the URL of [text](url) links.
Hindsight now ships from its maintainer's repo (vectorize-io/hindsight,
hindsight-integrations/hermes) through plugin-catalog/hindsight.yaml, so the
in-tree copy under plugins/memory/hindsight goes away. Homes that still name
`memory.provider: hindsight` are migrated by hermes_cli/memory_provider_migration.py
(the `hermes update` hook and the first agent start install the catalog plugin);
that path is untouched here.
Core-side special cases that only made sense with the bundled copy go with it:
the `memory.hindsight` LAZY_DEPS feature, `_provider_pip_dependencies`'s
`hindsight-all` expansion in `hermes memory setup` (the plugin's own
`_ensure_local_runtime()` installs the embedded runtime now), the compat-manifest
pointers of the deleted module, and prose that listed hindsight among the in-tree
providers. Generic provider-name lists, the HINDSIGHT_* env hints and the API-key
redaction pattern stay — a catalog-installed hindsight still uses them.
Revert this commit alone to restore the bundled provider.
A memory provider installed from the catalog under $HERMES_HOME/plugins/
(the honcho handoff package) loads under the loader's synthetic user
namespace, but the host-side surfaces still imported the bundled path
`plugins.memory.honcho.*` by name: the Desktop GET/PUT
/api/memory/providers/honcho/config 500'd, /oauth/start|status 404'd
("does not support OAuth connect"), doctor reported "honcho-ai not
installed", and profile clone / post-update sync silently skipped.
Add `plugins.memory.import_provider_module(name, submodule=None)`: the
provider package (or one of its submodules) of whichever copy
`find_provider_dir` resolves — bundled today, user dir after the removal
PR — under the module name the loader already owns, so both copies
behave identically. Route every host-side absolute import through it:
the dashboard host-block resolvers and honcho.json writer, the OAuth
route resolver, doctor's honcho/mem0 checks, `profile create --clone`,
the update hook's profile sync and the holographic store release.
In-process, no lock, no child process.
Two invariant tests (user-dir honcho with the bundled copy gone:
/config?surface=declared is 200 with the schema; the oauth_flow module
resolves from the user directory), red on base. The honcho write test's
`_honcho_resolvers` stub gains the provider-name argument the seam now
carries.
Shape follows the host-module resolution slice of #116566 by @erosika;
the contract module, installer lock and child-process OAuth runner from
that PR are not needed once the modules resolve in-process.
Co-authored-by: Erosika <eri@plasticlabs.ai>
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:
1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
with a per-plugin activation summary (hermes_cli/plugins_activation.py):
activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
plugins.manage install/toggle/update, dashboard REST install, tool-triggered
force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
discover_plugins() short-circuits on _discovered, which is why reload.mcp after
a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
factories not yet wired on the live native client (keyed (plugin, qualname);
a force reload hands back new function objects). Telegram hoists late handlers
ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
the first match per group) and re-wires on the transient-init rebuild; Slack
dedupes register_slack_action_handler per AsyncApp. The gateway runner
subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
hint and plugins.manage results (activation, gateway_reloaded,
restart_required only when no gateway answered) say exactly that.
Parse memory-provider manifests through the declared hermes_yaml layer so manifest-name deny-list entries cannot fail open when PyYAML is absent. Exercise enabled and disabled profiles A→B→A to lock down config and module cache isolation.
_auto_create_thread's dedup pre-seed (is_duplicate(str(thread.id))) stops
the echo MESSAGE_CREATE Discord fires for the starter (id == thread.id)
from re-running the request. This PR turned the tracker persist into
`await self._threads.mark_async(thread_id)` and placed it BEFORE the
pre-seed; to_thread always suspends, so the echo's handler could reach
_discord_message_admission -> is_duplicate during the os.replace window,
claim the id first, and rerun the starter. On main both statements were
sync, so no window existed.
Move the pre-seed (with its comment) directly after
`thread_id = str(thread.id)`; the await now follows it.
Other mark_async call sites checked (discord :4717 slash create-thread,
discord :6140 pre-dispatch, matrix :1438 create_thread, matrix :2083
inbound): nothing after those awaits relies on state a concurrent event
could claim first, so no further reorder.
PROOF: tests/gateway/test_discord_double_dispatch.py::
test_thread_starter_duplicate_dropped now installs a recording mark_async
that asserts thread.id is already in _dedup._seen when it is awaited.
Red with the pre-swap order (1 failed), green after; ruff clean,
check-windows-footguns clean, real import of the adapter from the
worktree ok, scripts/run_tests.sh on the PR's test files green.
`record`/`record_media` do a read-modify-write of the JSON index ending in
`atomic_json_write`/`os.replace`, and every coroutine on the WhatsApp Cloud,
WhatsApp bridge and Telegram send/inbound paths called them inline on the
loop thread, stalling the whole gateway for a filesystem write nothing
awaits. Add `record_async`/`record_media_async` as thin
`asyncio.to_thread` wrappers (same precedent as
gateway/channel_directory.py) and await them from the seven coroutine call
sites; the sync functions stay for sync callers. The write completes before
the await returns, so lookup-after-send behaviour is unchanged.
Salvaged from #118883 by @Kyzcreig; re-shaped onto asyncio.to_thread (no
writer thread, no in-memory shadow, no queue), plus the missed
plugins/platforms/whatsapp/adapter.py record_media site.
`ThreadParticipationTracker._save` (gateway/platforms/helpers.py) ends in
`atomic_json_write` -> `os.replace`, whose duration is unbounded under
filesystem pressure. All five call sites are coroutines on the inbound-message
or slash-command path:
plugins/platforms/matrix/adapter.py _resolve_message_context
create_handoff_thread
plugins/platforms/discord/adapter.py _handle_message (x2)
_handle_thread_create_slash
So that rename was paid inline on the running loop, stalling every other
adapter's polling, every in-flight turn and every heartbeat for as long as it
took.
The fix is at the CHOKE POINT rather than at five call sites:
- `mark_async` does the in-memory insert synchronously and offloads only the
persist via `asyncio.to_thread`. The insert must stay synchronous because
both adapters gate on `thread_id in self._threads` immediately after
marking; deferring it would make mention-gating depend on executor
availability.
- all five coroutine call sites now await it.
- `mark` keeps its exact synchronous contract for the non-loop callers.
- an RLock is added in the SAME commit that introduces the concurrency: the
event loop used to serialize every caller by accident, and
`atomic_json_write` makes each write atomic without making
check/insert/trim/write atomic. Without it two concurrent marks lose one.
Enforcement is an AST class sweep, not an inventory: no `async def` under
gateway/ or plugins/ may call `<x>._threads.mark(...)`, and every
`mark_async(...)` must be awaited -- an un-awaited one never runs at all, so
the thread is neither persisted nor recorded in memory and mention-gating
re-prompts forever in a thread the bot already joined. A new adapter fails the
gate without anyone remembering a list.
tests/gateway/test_discord_thread_slash_expired_defer.py stubbed the tracker
with `SimpleNamespace(mark=...)`; it now uses the real tracker against a
tmp_path, so the test cannot rot silently the next time this surface moves, and
it additionally asserts the thread really was recorded.
Verified on this exact head, PYTHONPATH pinned to the worktree:
tests/gateway/test_thread_tracker_mark_off_loop.py 7 passed
expired-defer + admission-exemption + off-loop 10 passed
-k "discord or matrix or thread" over tests/gateway 1136 passed, 13 failed
The 13 failures are INHERITED: a clean worktree at upstream/main 75e9567ca7
with none of these changes fails the identical 13.
Every guard is gate-proven -- reverting the RLock, the to_thread, one call
site, one `await`, the sync insert, or the dedupe short-circuit each fails its
own test and only its own.
(cherry picked from commit 6fc8037d292ec051fe76ad213903b2ee5534080a)
`gateway.sticker_cache._save_cache` ends in `atomic_json_write` -> `os.replace`,
whose duration is unbounded under filesystem pressure. Its only production
caller is Telegram's `_handle_sticker` -- an inbound-message coroutine -- so
every sticker that missed the cache paid that rename inline on the event loop,
stalling every other adapter and every in-flight turn in the process for its
duration.
Adds `cache_sticker_description_async`, which dispatches the existing sync form
via `asyncio.to_thread`; the Telegram call site awaits it. The sync form keeps
its exact contract and is what the wrapper dispatches to, so it is unchanged for
non-loop callers.
`cache_sticker_description` is a read-modify-write (`_load_cache` -> merge ->
`_save_cache`). `atomic_json_write` makes each WRITE atomic, not the TRIPLE.
While the call was inline the event loop serialized every caller and the race
could not be observed; moving the write to a worker thread introduces real
concurrency, so `_CACHE_LOCK` is added in this same commit rather than deferred.
Tests use no wall-clock thresholds. The liveness witness is ORDERING: the rename
is held open on a barrier released only by a background timer, and the sibling
task must have ticked BEFORE that release. Verified RED first:
- revert `asyncio.to_thread` to an inline call -> 2 failed
("the cache write ran on the event-loop thread")
- revert `_CACHE_LOCK` -> the concurrency test fails on the DURABLE FILE:
"lost an entry ... keys = ['uid_a']"
`test_gate_proof_the_sync_form_does_block_the_loop` drives the identical barrier
through the sync form and asserts the loop DOES starve, so the liveness
assertion cannot pass vacuously.
GREEN: sticker off-loop + existing sticker cache tests -> 10 passed;
`tests/gateway -k "sticker or telegram"` -> 937 passed, 4 skipped; Ruff clean.
(cherry picked from commit 9752a9b1aba7536181d081cad9c1c32917a89680)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Declare `_TERMINAL_HEALTH_REASONS = frozenset({"socket_closed",
"client_closed"})` beside `_read_websocket_health` (the producer of those
literals at adapter.py:1705/:1707/:1715) and use it at the escalation
site in `_liveness_loop` (was `reason in ("socket_closed",
"client_closed")` at :1791).
WHY: the first-strike classification was a magic-string match against
literals produced 80 lines away with no shared definition. A rename in
the producer, or a new hard-closed reason, would silently demote the
escalation back to the threshold path with nothing failing. One named
set next to the producer makes the coupling visible. The
`tuple[bool, str]` return shape is unchanged (existing tests monkeypatch
the sampler with 2-tuple lambdas).
Proof: mutating the set to {"socket_closed"} fails
test_closed_transport_first_strike_forces_reconnect[client_closed];
restored, the liveness file is green.
Collapse the INFO/WARNING level switch on the probe-exit log
(plugins/platforms/discord/adapter.py:1755-1766) to one unconditional
logger.info and delete `expected_exit`. Drop
test_unexpected_probe_exit_logs_warning, which was the only caller that
could reach the WARNING half.
WHY: the WARNING branch is unreachable in production. It needs
`_running=True and not _disconnecting and _client is None` at the guard,
and `_client` has exactly two production setters to None:
- adapter.py:1269 inside connect(): synchronously followed by
`self._client = commands.Bot(...)` at :1271 with no await between, so
the probe coroutine can never observe the None.
- adapter.py:1941 inside disconnect(): runs after `_disconnecting = True`
(:1915) and after `await self._cancel_liveness_task()` (:1917), so the
probe is already cancelled (exits via CancelledError at :1752) and the
flag would read as an expected exit anyway.
No other `_client = None` in plugins/platforms/discord/, gateway/platforms/
base.py or gateway/run.py. `_running=False` and `_disconnecting=True` come
only from teardown paths, all "expected". The deleted test pinned this by
poking `adapter._client = None` directly — a state the code never produces.
This drops the level switch borrowed from #118504 but keeps its intent:
the probe exit is always logged with its full state (#118487, the probe
must never disappear silently).
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.
Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.
(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.
Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.
- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
Symptoms fixed (all live-reproduced on origin/main in a fake HERMES_HOME):
- `hermes plugins disable photon-platform` wrote `platforms/photon` while the loader keyed the
bundled adapter `photon-platform`, so the disable never applied (#27548). The loader now keys
bundled platforms `platforms/<dir>` like every other category (one call site in
plugins_discovery.py); the manifest name stays an accepted alias through gate_manifest.
- `plugins.manage toggle` / dashboard toggle wrote the raw identifier: enabling by bare leaf or
manifest name returned ok while a stale canonical key in plugins.disabled kept the plugin off.
dashboard_set_agent_plugin_enabled resolves the canonical key and purges every alias from the
opposing list (shared _activate_key, also used by cmd_enable/cmd_disable); the RPC reports the key.
- remove/uninstall (CLI, dashboard, RPC) left the plugin in plugins.enabled/disabled and
plugins.entries (#54336), left memory.provider dangling so the next agent init re-cloned the
provider from the catalog (uninstall silently reverted), and, for a symlink inside the plugins
dir, deleted the TARGET plugin and its metadata while the link stayed dangling. _remove_user_plugin
is the shared tail: unlink a link only, then forget config under every alias and reset
memory.provider (reported as cleared_memory_provider).
- `hermes plugins list` / dashboard hub / TUI hub called bundled backends, bundled platforms and the
live memory.provider "not enabled" (#73131, #82898): _plugin_status mirrors gate_manifest.
- A user-installed memory provider parked in plugins.disabled kept loading: load_memory_provider
honours the deny-list (name, dir name or manifest name) and says so once.
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995