Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.
This adds the dispatch-side signal that was requested instead:
- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
every parsed DISPATCH frame with no debug gate (verified live against
the real received_message path: 4/4 frames fired with
enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
incident report's field-proven operator bound); 0 opts out of this
dimension ONLY — the #109782 review failure put the knob in
_start_liveness_probe's all-or-nothing guard, killing the whole
watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out
Fixes#109521
(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
Add HTTP-Referer, X-Title, User-Agent (HermesAgent/<version>) and
X-Pplx-Integration: hermes-agent to the Perplexity web provider's
/search and /sdk/content/snippets requests. Same static headers Hermes
already sends Kimi and OpenCode; nothing per-user, no extra requests.
The openai-codex image provider rode a Responses call with a hosted
image_generation tool on a pinned chat model (gpt-5.5). Two failure
classes came with that shape: when OpenAI withdrew gpt-5.5 from an
account cohort every image call 404'd while chat kept working
(#105398, #107076), and the host model was free to answer in text
instead of calling the tool, so we streamed SSE, kept partial frames
and retried on empty streams.
Post to chatgpt.com/backend-api/codex/images/generations and
images/edits instead - the route the official Codex client uses
(codex-rs/ext/image-generation). No host model, no SSE, no
partial-frame handling; the response is a plain JSON body with
b64_json. Remote source URLs are fetched client-side and inlined as
data URLs because the backend's own downloader 400s on ordinary
public images.
The backend treats model/quality/size as advisory (#107233), so the
result now reports reported_quality/reported_size next to the
requested values plus the x-codex-imagegen-request-id for support.
GPT Image 2.5 is deliberately not added to this catalog: the backend
accepts any model id, including nonexistent ones, and generates with
its server-managed engine (C2PA reports gpt-image 2.0), so a 2.5 tier
here would be a label with no effect (#106708).
Gateway, dashboard, ACP server and the CLI already hold one registry-shared
SessionDB per state.db path, yet several in-process call paths still opened
a bare SessionDB() beside it. Each one is a full writer: schema init, write
lock, token-writer thread and a close-time WAL checkpoint. On a dashboard
serving overlapping requests that stacked up to the "5 live SessionDB
handles" precursor within seconds; per #110544 they are now harmless to each
other's WAL generation, but the leak itself remained.
Pure readers attach read_only=True (no writer connection, no write lock):
- plugins/hermes-achievements/dashboard/plugin_api.py::scan_sessions
(dashboard, per background scan and per /rescan; highest-frequency site)
- hermes_cli/console_engine.py::_session_db (dashboard console; list,
stats and export are reads; rename/optimize opt in to a writer)
- tools/process_registry_results.py::_owns_result (gateway, per retained
result load)
- hermes_cli/main.py::_session_db (last-session / title / cwd lookups)
- hermes_cli/terminal_breadcrumbs.py, hermes_cli/status.py,
hermes_cli/main_tui_launch.py (lookup one-shots)
Writers share the process's registry handle (hermes_state_registry.acquire;
close()/release_or_close release one refcount):
- acp_adapter/session.py::SessionManager._get_db — the AIAgent it builds
acquires the same path, so the ACP server held two writers per process
- hermes_cli/kanban_db_dispatch.py::_retag_legacy_worker_sessions
(gateway dispatcher tick)
- hermes_cli/main.py::_create_titled_session, hermes_cli/oneshot.py,
hermes_cli/foreign_sessions.py — the CLI acquires the same handle a
moment later
The "N live SessionDB handles" warning now counts only writable members:
read-only attaches are the sanctioned per-request shape for dashboard
routers and CLI lookups, and counting them turned a healthy topology into
an operator alarm (#100896 field reports of restart loops keyed on it).
Live repro (one registry writer + 4 overlapping dashboard/gateway paths in
one process): before 5 writable opens, 5 live handles, warning fired;
after 1 writable open (the registry handle), 0 from the request paths,
no warning.
Refs #100896#103339
FAL shipped google/gemini-omni-flash/v1.1/* (Aug 2026): the family gains a
text-to-video endpoint, 360p/720p/1080p/4k resolution enum, and keeps
3-10s integer durations with always-on native audio.
- plugins/video_gen/fal: bump gemini-omni-flash to the v1.1 endpoints,
declare the resolution enum, refresh display/strengths copy
- tests: family test asserts the versioned dual-modality endpoints; the
i2v-only clean-error guard survives via a synthetic family; catalog
invariant now requires both endpoints on every family
Schema verified against FAL OpenAPI (t2v: prompt required, 16:9/9:16,
360p-4k, duration 3-10 int; i2v adds image_url + optional end_image_url).
Pricing: $0.03/s 360p, $0.10/s 720p, $0.15/s 1080p, $0.30/s 4K.
Alibaba's Wan 3.0 generation went live on FAL (was "Coming Soon" since
the Wan 2.7 rollout). Adds two families to FAL_FAMILIES:
- wan-3.0: alibaba/wan-3.0/{text,image}-to-video — 2-30s integer
durations, 480p/720p/1080p, native audio, adaptive aspect handling
($0.05-$0.20/s by resolution)
- wan-3.0-prime: alibaba/wan-3.0-prime/{text,image}-to-video — premium
tier, same surface ($0.068-$0.28/s)
Schema quirks mapped from the FAL OpenAPI/llms.txt:
- i2v takes `start_image_url` (image_param_key, Kling-4K pattern)
- duration is a JSON integer (duration_int)
- audio toggle key is `audio`, not `generate_audio` — new generic
`audio_param_key` family flag in _build_payload (veo3.1 regression
asserted unchanged)
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.
- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
and it now honors the base_url argument instead of hardcoding
api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
incl. single-page and stuck-cursor termination; updated the two URL-pinning
pool-discovery tests for the ?limit=1000 contract
Live anonymous probes (2026-09-13, HermesAgent UA, 4 rounds): nemotron-3.5-lightning-free
returned zero bytes for >90s on every attempt and nemotron-3-ultra-free took ~40s, while
mimo-v2.5-free answered 200 in 2-4s each time. An aux model (titles, compression, memory)
that hangs is worse than a delisted one, so the default moves to the model that works.
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).
- hermes_cli/models_catalog_static.py: remove both slugs from the
opencode-free offline floor and the opencode-zen discovery floor;
document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
invariant to cover both.
The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.
- _render_slack_table_block(): projects rows as 'cell | cell' lines,
collecting text leaves from raw_text/rich_text cell subtrees; capped
at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
(thread history + attachment-nested blocks), the live inbound
attachment loop, and _extract_additional_text_from_slack_blocks
(top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
allowlist drops 'rows', so it only emitted an empty husk.
send_voice built its own discord.File from the audio bytes and never ran
the size preflight, so an oversized audio attachment still burned the
doomed 413 round-trip that #50846 is about. Route it through the same
_reject_oversized_upload helper as _send_file_attachment.
The limit constant and _discord_upload_limit_bytes lived on the adapter
facade while every consumer is in adapter_media.py; the facade+sibling
layout puts topic code in the sibling, so they move there.
Discord raised the default file upload limit from 10 MiB to 20 MiB for
users, bots, webhooks and interaction responses (developer changelog,
Sep 3 2026). The 25 MiB constant here predates the preflight salvage and
never matched the platform; more importantly discord.py 2.7.1 still
reports 10 MiB via guild.filesize_limit for unboosted guilds, so the
guild-aware path under-reported the cap and rejected 10-20 MiB files
Discord now accepts. Floor the guild value at the platform default so a
stale library constant can only widen, never shrink, the preflight.
The adapter facade is ~6.6k lines; new behaviour belongs in a topical sibling per the
facade+siblings layout. expand_link_entities() now lives in telegram_entities.py and
reuses the encode/decode UTF-16 slicing the adapter already uses for entity spans.
Also: skip inlining when the anchor text already is the URL (no 'url (url)' duplication),
trim the test file to the invariants and point it at the sibling.
Image.open() was never closed; convert()/resize() return new images so the
source file object lingered until GC (a real leak on Windows where the open
handle blocks later deletion of the original). Use the context manager.
- Keep main's media_write_timeout=60s (PR's HERMES_* env var dropped per
.env-is-secrets-only policy; main already fixed the timeout half).
- Replace stdlib imghdr (removed in Python 3.13) with a magic-byte sniff.
- Exclude GIFs: JPEG conversion flattens animations to one frame.
- Fix transparent-PNG handling: RGBA hit the len(getbands())==4 branch
before the white-background composite, rendering transparency black.
- Clean up temp JPEGs after send (both single and media-group paths);
the docstring promised caller cleanup that neither call site did.
- Write temp files via tempfile default dir instead of an undefined
DEFAULT_OUTPUT_DIR (NameError at runtime in the original PR).
- Add real-Pillow regression tests incl. a sabotage-verified
white-background test.
Behind an HTTP proxy (e.g. tgapi.indevs.in) the PTB
media_write_timeout (~20s) is exceeded by raw PNGs > 1-2MB, causing
TimedOut errors on both send_photo and the send_document fallback.
Add TelegramAdapter._compress_image_to_jpeg() which converts large
PNG/raster images (>1MB) to progressive JPEG at 85% quality, with
optional resize above 1600px. Applied in send_image_file() and in the
media-group path of send_multiple_images(). Compression is a no-op for
JPEGs, small files, and non-raster formats, and falls back gracefully
if Pillow is unavailable.
Co-authored-by: user
Swallowing FeatureUnavailable returned a bare False, so a hosted operator
saw the same generic "requirements not met / run hermes setup" line the
original report started with. The registry's _probe already logs a raised
exception with its message, so propagating the error is what puts
"target not writable" / "quarantine 404" in gateway.log.
Hosted/Docker images lock /opt/hermes/.venv, so --install-deps writing
site-packages fails with Permission denied and the adapter never starts.
Route Google Chat through lazy_deps (HERMES_LAZY_INSTALL_TARGET) and bake
the extra into the published image so a configured gateway can connect.
Rebase reconciliation with main's JSON-allowlist decoding (#109423): the two
remaining readers that consulted config.extra before the env var now follow the
per-profile precedence rule (explicit scoped env → the profile's YAML → default),
and ignored_threads still decodes a JSON-string list after the read.
The Matrix blank-YAML test asserted YAML-over-env, the old precedence #108440's
review flagged; it now pins the contract: explicit env beats YAML, a blank env
value is unset (YAML applies), YAML beats the default, and an explicit empty
list is a real "no rooms" value.
_bridge_env copied os.environ (the default profile's WHATSAPP_* values under
multiplex) and only overlaid scoped hits, so a secondary with YAML
`dm_policy: pairing` launched its Node bridge under the default profile's
`allowlist` policy and the bridge rejected valid pairing DMs before Python saw
them. The child env now carries the values the adapter resolved (scoped env →
own YAML → default); a scoped miss removes the key rather than inheriting it.
Refs #108440 (ehz0ah inline, plugins/platforms/whatsapp/adapter.py)
545e74d0ea correctly stopped writing telegram.proxy_url into TELEGRAM_PROXY
for a multiplexed secondary, but _build_ptb_requests still resolved the proxy
only from that env var, so the secondary silently connected direct (or via the
default's proxy). #100448 had deliberately left this bridge unscoped for that
reason; this finishes the consumer migration instead.
_apply_yaml_config seeds proxy_url into extra and resolve_proxy_url gains a
`configured` rung: scoped TELEGRAM_PROXY → the profile's YAML → HTTPS_PROXY/
HTTP_PROXY/ALL_PROXY (trust_env) → macOS system proxy, with NO_PROXY semantics
unchanged.
Refs #108440 (finding 6)
One reader (gateway.platforms._shared.extra_or_secret) now implements the
precedence every per-profile setting follows for the OWNING profile:
explicit scoped env/.env → that profile's config.yaml (PlatformConfig.extra)
→ the adapter's default. A scoped miss returns the default, never the launch
process's os.environ; single-profile / default-profile installs keep the
documented env-over-YAML contract.
Why: 545e74d0ea (#108705) stopped bridging a secondary's YAML into the
process env and moved readers to config.extra, but the shared reader and the
hand-rolled helpers in Discord/Slack/Matrix/Telegram consulted YAML FIRST and
then fell back to a scoped env read. Two bug classes followed (#108440
post-merge review by andrexibiza, #109032):
- an explicit env value could no longer beat YAML for the owning profile
(DISCORD_ALLOW_MENTION_EVERYONE=false lost to allow_mentions.everyone: true;
TELEGRAM_REACTIONS=true lost to the stock reactions: false);
- a secondary that OMITTED a key inherited the launch profile's bridged env
through the fallback (Matrix process_notices/session_scope, Discord
auto_thread/reactions/mentions, Slack reactions/ignored_channels).
Consumers migrated to the shared reader: Discord _build_allowed_mentions and
_extra_or_env_flag; Slack _slack_allow_bots, _reactions_enabled (the
_extra_or_env_* getters already used it); Matrix _extra_truthy, _extra_csv_set,
session_scope, reactions, require_mention parsers, and — new — the
allowed_users / ignore_user_patterns consumers that never read the seeded YAML
lists; Telegram _extra_bool, _extra_str_set, _reactions_enabled; Feishu
allow_bots; WhatsApp dm_policy/group_policy.
Refs #108440, #109032
A2A_PORT and A2A_ADVERTISED_TOOLSETS are already captured at
construction time (inside _profile_runtime_scope) via
_get_scoped_secret(), but A2A_PUBLIC_URL was still read with a bare
os.getenv() inside A2ARequestHandler._request_public_url() - which
runs on ThreadingHTTPServer's per-connection OS thread, not the
constructing thread.
Raw threading.Thread never inherits contextvars, so even swapping the
reader to _get_scoped_secret() at that call site would not help: the
request thread has no scope, secret_scope falls back to os.environ
either way. The value must be captured once at construction time
(which does run in profile scope) and threaded through as instance
state instead - same fix shape as A2A_PORT above.
A secondary multiplex profile without its own A2A_PUBLIC_URL now
falls back to the X-Forwarded-Host/Host-derived URL (or the bind
host) instead of silently advertising the default profile's public
URL in its Agent Card / discovery response.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 0c36aca5de53d88bbbc0b4cfceaed8307c744f7a)
#69090 scoped MATRIX_RECOVERY_KEY itself (via _scoped_recovery_key())
so a secondary profile resolves its own recovery key under multiplex,
but left its sibling, MATRIX_RECOVERY_KEY_OUTPUT_FILE, on a bare
os.getenv(). _recovery_key_output_path() is called from inside
_verify_or_bootstrap_cross_signing(), which runs fully inside
_profile_runtime_scope for a secondary profile: when that profile
bootstraps a new recovery key, it either doesn't get written to a
file at all, or gets written to the default profile's configured
path, depending on which one has the env var set.
Route it through the same _get_scoped_secret() helper _scoped_recovery_key()
already uses.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit fb765ee49b2f1a1853e52fb901d25f767180a2fc)
545e74d0ea made _reactions_enabled consult extra.reactions before the env
var, and _apply_yaml_config seeds extra["reactions"] whenever the YAML key
is present — including the stock reactions: false every install
materializes. The documented TELEGRAM_REACTIONS=true switch therefore
became a silent no-op after the 0.21.2 update (#109032), contradicting
yaml_env_setter's "explicit env wins over YAML" contract.
Read the scoped env first and fall back to the profile's own YAML: under
multiplex a scoped miss returns the default instead of another profile's
process-env value (#72348), so only a scoped/env hit counts as explicit
and per-profile isolation is unchanged.
Fixes#109032
(cherry picked from commit 2bd5a0a5c0a5f9630fd82f133def65f225f653a3)
`_orphan_timeout()` was `max(300, A2A_REPLY_TIMEOUT)` with no ceiling, so
an absurd value (1e18) meant the watchdog sweep could never fail an
orphan — the reply window is a floor for the grace, not a licence to
disable the sweep. Cap it at 86400s.
`disconnect()` failed and cleared `_pending`/`_pending_order` but left
`_active_tasks` populated, so a reconnected adapter would keep excluding
dead task ids from the orphan sweep forever. Clear it in the same locked
block.
The troubleshooting entry told users to raise A2A_REPLY_TIMEOUT for long tasks,
which did nothing against the hardcoded 300s orphan sweep (#106972). Now that the
sweep derives its grace from the reply window and skips tasks with a live waiter,
state that contract next to the variable.
545e74d0 (post-0.21.2) made the Matrix YAML bridge seed its values into
PlatformConfig.extra so secondary multiplex profiles read their own config. The
"csv" bridge kind seeds any non-None value, so `free_response_rooms: ''` now
reaches extra as '' — and the readers' `if raw is None` fallback no longer fires,
so MATRIX_FREE_RESPONSE_ROOMS is ignored and require_mention drops every
un-mentioned message. Before that commit the bridge only wrote env and the key
was absent from extra, so the env value applied.
Route the three identity-check readers (_extra_csv_set, _extra_truthy,
_resolve_max_message_length — the last a three-tier chain where '' also
short-circuited the plugin-registry default) through the shared
gateway.platforms._shared.extra_or_secret, whose default already treats a blank
string as unset (the idiom mattermost/dingtalk/slack readers use). Explicit
scalars, bools and lists (including []) stay authoritative.
Two invariant tests replace the salvaged suite (moved to
tests/plugins/platforms/matrix/ to mirror the source path): blank falls through
for all three readers; explicit values still beat env.
Fixes#109358
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Under gateway.multiplex_profiles os.environ holds the default profile's .env, so `_run_brv`
building the child env from raw os.environ curated a secondary profile's turns into the DEFAULT
profile's ByteRover cloud account (and prefetched the default's memories into the secondary's
context). The local half was already profile-scoped (`_get_brv_cwd`).
The child env now comes from `build_subprocess_env` and, under multiplex, strips the launch
profile's residue and sets BRV_API_KEY only from the served profile's secret scope — a miss means
no cloud key. Single-profile installs pass the process env through unchanged.
Closes#108993 (report and fix direction by @jonpol01).
`langfuse._secret` and `azure_identity_adapter._scoped_env` were changed to
raise rather than fall back, because swallowing `UnscopedSecretError` hides the
spawn-site bug the exception exists to surface. `_scoped_setting` looked like it
contradicted that, so make the split explicit and pin it.
Hindsight already follows the contract for everything that decides WHERE data
goes: `mode`, `apiKey` and the `bankId` partition read through bare
`get_secret`, so a scopeless multiplexed read raises. In `_load_config` that
raise happens on `HINDSIGHT_MODE` before any shaping value is reached, so the
swallow below cannot mask an isolation failure.
Presentation shaping is deliberately not in that class. `MemoryManager._each_provider`
logs an `initialize` failure at WARNING and drops the provider for the session,
so raising there would cost the whole memory provider because a speaker prefix
could not be resolved. It degrades to the provider's own default instead —
never to `os.environ`, which under multiplex is the default profile's.
The test names the offending key rather than asserting that something raised:
routing `mode` through the shaping helper shifts the failure to
`HINDSIGHT_API_KEY`, which a bare `pytest.raises` would still accept.
_load_config() reads the Hindsight bank, mode and retain tags through the profile
secret scope, but _apply_retain_settings() then discarded that answer and re-read
os.environ whenever the config value was falsy:
return cfg.get(key) or os.environ.get(env_var, default)
Under gateway.multiplex_profiles os.environ holds the DEFAULT profile's .env, so a
secondary profile's scoped miss came back as the default profile's retain tags,
observation scopes, source and speaker prefixes — the fallback-after-miss shape
gateway/AGENTS.md forbids. Tags are Hindsight's retrieval partition and
metadata.source is opt-in by design, so the secondary's memories were both
mislabelled and selectable by the default profile's tag filters.
Both halves now go through _scoped_setting(), which resolves the value with
get_secret() and falls back to the provider's OWN default — a miss is a miss, the
same rule embedded.py already applies to the daemon's key and base URL. The three
raw reads left inside _load_config() (retain_source, retain_user_prefix,
retain_assistant_prefix), directly under the comment declaring them per-profile,
go through it too.
Single-profile deployments are unchanged: with no scope installed get_secret()
still reads the process env, where the value IS this profile's own.
Fixes#108865
The adapter fix decoded `'["-100","-200"]'` before comma-splitting, but
the runner's central gate in gateway/authz_mixin.py::_coerce_allow_set
reads the same YAML-bridged env chain (TELEGRAM_GROUP_ALLOWED_CHATS,
TELEGRAM_ALLOWED_USERS via _auth_env) and still produced
{'["1"', '"2"]'}, so a group message admitted by the adapter could
still be rejected upstream.
Move the decoder to gateway/platforms/_shared.py, which both the adapter
and authz_mixin already import from (no plugin -> gateway cycle), and
route _coerce_allow_set through it. One invariant test on the runner
side, red before this change.
Same class as minimax-portal: aliyun, deep-seek, nim/build-nvidia/nemotron
and vertexai resolve in hermes_cli's alias tables but not in
providers.get_provider_profile(), so a profile lookup keyed on the alias
returned None and lost the profile's wire mode/extra_body/headers. qwen is
left out: the catalog maps it to alibaba while the qwen-oauth plugin already
claims it, a pre-existing disagreement outside this fix.
`_finite_positive_config_float` and `_config_int` were the same shape with the warn call
pasted six times and asymmetric sign checks. Collapse both onto `_liveness_knob(key, default,
cast)`: usable iff finite, >= 0 and exact for the cast; else warn and return 0; explicit 0
stays silent.
Two int-path holes closed on the way: `websocket_liveness_failure_threshold: .inf` raised
OverflowError inside `DiscordAdapter.__init__`, and `0.5` truncated to 0 and disabled the
probe silently — the bug class this PR exists to remove. `2.0` / `"2"` still resolve to 2.
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:
- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
`enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
`log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
`on_ready` and every healthy connection reads `event_silence` 300s later — a forced
reconnect every ~5 min. Live-verified against a real `commands.Bot` +
`DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
`ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
detect the "ESTAB + ACKing + zero events" incident by construction.
Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.
Two defects fixed:
1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
`_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
and ACKs included, so a legitimately quiet server is not flagged. The
health check gains an `event_silence` reason with its own bound,
`websocket_event_max_silence_seconds` (default 300s, 0 disables the
dimension alone). The stamp resets on `on_ready` so a reconnect never
inherits pre-restart silence. Trip path is unchanged: consecutive
failures -> retryable `discord_websocket_health_stale` -> the existing
reconnect watcher builds a fresh adapter.
2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
log line, permanently disabling the watchdog invisibly. Unparsable and
non-positive values now log one WARNING naming the knob and raw value.
The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
The delegation imported ``plugins.model_providers.deepseek`` — the only cross-plugin
module import in the tree, and one that resolves solely through the loader's
sys.modules shim (popped again if the deepseek plugin fails to load). Look the
profile up with get_provider_profile("deepseek") instead: already imported from
``providers``, honours a user override of the profile, degrades to the base no-op
without a try/except.
Drop the ``len(m) <= len("deepseek/")`` guard — the native profile returns
({}, {}) for an empty id anyway. Bind the expected value in the parity test and
assert it is non-empty so the equality cannot pass as ({}, {}) == ({}, {}).
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
`/model <x>` onto copilot-acp validates through `models_validate._static_catalog`, which
reads `provider_model_ids` with no disk cache. After the session probe landed, every such
switch spawned `copilot --acp`, ran the handshake, and killed it (1-3 s; up to the 15 s
probe timeout when the CLI is installed but the session stalls). The GitHub-API tier that
path used before sat behind a 5-minute in-memory memo; the ACP tier now has the same memo,
and it remembers failures too so a broken CLI is not re-spawned per switch.
The probe itself moves to `CopilotACPProfile.fetch_models` — the slot that already said
"model listing is handled by the ACP subprocess" and returned None — so hermes_cli/models.py
no longer hand-builds `CopilotACPClient` kwargs that `profile.create_client` owns.
Discovery failures are logged at debug instead of swallowed.
Tests: the two picker wiring tests collapse into one parametrized contract; a new test
proves three consecutive switch validations pay one probe and a failed probe is not retried
(fails when the memo read is removed).