Commit Graph

2346 Commits

Author SHA1 Message Date
teknium1
33f3d96b99 fix(disk-cleanup): never rmtree protected top-level dirs; never track kanban/
The tracked-item delete path in quick() called shutil.rmtree() on any tracked
directory without consulting the protection list, which only the empty-dir
sweep used. guess_category() files every path under cache/ as "temp", so a
terminal command that merely mentioned $HERMES_HOME/cache tracked the
directory itself, and 7 days later quick() removed cache/ wholesale — taking
cache/terminal (terminal snapshots) with it and breaking every later command
with a mktemp "No such file or directory".

Fix: _is_protected_dir() — a tracked DIRECTORY that is HERMES_HOME itself or
sits under an _EMPTY_DIR_PROTECTED_TOP_LEVEL tree is never tracked by
guess_category(), never listed by dry_run(), and skipped (logged SKIPPED,
entry dropped) by quick(). Files under cache/ still age out as before.

kanban/ (task attachments and workspaces have their own lifecycle) is added to
both _NEVER_TRACK_TOP_LEVEL and _EMPTY_DIR_PROTECTED_TOP_LEVEL; stale pre-fix
"test" entries under it are dropped by the existing re-validation instead of
deleted. The kanban row was first proposed in #80842 (@nicha16).

Fixes #114552
2026-09-18 09:22:13 -07:00
beardthelion
9ac4e76e0f fix(tools,tui_gateway,cli,plugins): parseable non-dict JSON no longer crashes the remaining file scans
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:

- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
  envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
  scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
  the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
  scalar snapshot reads as empty / returns the existing 5000 error instead of
  violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
  dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
  _combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
  unchanged.

Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.

(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
2026-09-18 09:19:04 -07:00
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
Victor Kyriazakos
5648f81431 fix(slack): reopen a native task card when Slack seals the stream mid-turn
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.

On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.

Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
2026-09-17 18:52:15 -07:00
Victor Kyriazakos
c57458a97c fix(slack): retain bearer on validated Enterprise Grid file redirects 2026-09-17 18:35:54 -07:00
kshitijk4poor
70addd3522 fix(gateway): keep default-off byte parity; scope only the policy reads
Four places changed behaviour for users who never touched the setting:

- `_interim_send` was stamped on every `warn` status and media-failure notice, and the
  Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry
  no interim mark at all, so the gap is class-wide (every status kind), and fixing it for
  warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class
  fix belongs in its own PR against gateway/AGENTS.md rule 3.
- The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger
  writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so
  delivery obligations landed in the routed profile's state.db while boot-time
  `_claim_pending_obligations` still reads the launch home. Only the policy reads
  (`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where
  main ran it.
- The turn-crash notice is rebuilt the same way: scope around the policy read, send outside.
- The "delivery failed after multiple attempts" notice is unconditional again: the requested
  result itself was lost and this line is its only signal, so it is not a diagnostic.

Tests that asserted the reverted behaviours are removed; the reviewer-round test file is
renamed for what it covers.
2026-09-18 01:43:35 +05:30
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
b675e6dee7 fix(discord): re-arm the bot handoff window per chunk; trim salvage to invariants
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.

Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
2026-09-17 12:07:09 -07:00
simpolism
2c5504bf26 docs(discord): clarify deployed bot handoff behavior and verify ingress 2026-09-17 12:07:09 -07:00
simpolism
5d220ed693 fix(discord): batch bot tag continuations 2026-09-17 12:07:09 -07:00
simpolism
f708640561 fix(discord): require hard mentions from bots 2026-09-17 12:07:09 -07:00
teknium1
df1074b4e5 fix(aux): structured-output rejection no longer kills fallback candidates or costs a doomed first request
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):

* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
  aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
  rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
  #89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
  paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
  provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
  https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
  that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
  and fallback paths — omits the field before the first request. Dropping rather than
  downgrading to json_object matches the end state the retry already produced; json_object
  needs a JSON-mentioning prompt and some relays return empty content under it.

New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.

Fixes #83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
2026-09-17 09:11:38 -07:00
kshitijk4poor
bb42672178 refactor(telegram): a confirmed stall logs one hand-off line, not a retry promise
_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.
2026-09-17 21:17:59 +05:30
kshitijk4poor
5ca3e5546c fix(telegram): a confirmed stall hands off before the reconnect backoff, not after it
Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.

WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.

The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.
2026-09-17 21:17:59 +05:30
kshitijk4poor
4aa76eb82e refactor(telegram): watchdog stall reuses _schedule_polling_recovery
`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.

WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.
2026-09-17 21:17:59 +05:30
kshitijk4poor
cb4063f76a refactor(telegram): type the polling-stall signal instead of matching log text
Replace `_looks_like_polling_stall` (a three-substring classifier over
`str(error)`) with `_PollingStallError(RuntimeError)`. The watchdog and the
post-reconnect verifier raise it at their two stall sites; the reconnect
ladder checks `isinstance` through `_iter_exception_graph` before handing
the adapter to the supervisor.

WHY: a substring match couples recovery routing to log wording (rename a
message, silently lose the fatal handoff) and can misfire on unrelated
errors that quote the same words. A typed exception keeps #113657's
behaviour with no behavioural change beyond the classifier's source.
2026-09-17 21:17:59 +05:30
joaomarcos
4a7ed0029d fix(telegram): rebuild adapter on confirmed polling stall instead of reusing unquiesced updater (#113618) 2026-09-17 21:17:59 +05:30
teknium1
7f03867523 fix(aux): parameter rungs hand a 429 on, and Fireworks gets top-level reasoning_effort
The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
2026-09-17 08:45:46 -07:00
Changhyun Min
98f758ae7e fix(upstage): default new Solar models to the 512K context window
Upstage /v1/models carries no context field, so Solar ids resolve through
DEFAULT_CONTEXT_LENGTHS, whose keys are substring matches. solar-mini4 hit
the legacy solar-mini key (32K, below the 64K minimum), so the agent refused
to start; solar-pro4 matched nothing and got the 256K fallback. The same
substring deny-list in the Upstage profile also treated solar-mini4 as
non-reasoning, silently dropping reasoning_effort for a model that reasons.

Solar keys now match only on an id boundary (key followed by end or one of
"-:.@"), after folding aggregator slugs such as solar-pro-3 into the native
solar-pro3. Legacy keys keep covering their dated, quant and variant ids
(solar-mini-250422, solar-pro3@q4) but not later generations. A "solar-"
family key gives Solar lineup ids (bare name "solar-<letter>...", org/
prefix allowed) 524288 tokens, per Upstage /v1/solar/models max_model_len;
open-weight ids like solar-10.7b-instruct keep the generic fallback, and an
explicit key still wins by longest-key-first order. The boundary-matched
prefixes are a module constant, so non-Solar keys are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 17:10:58 +05:30
teknium1
816cb37963 fix(web): ladder-derived firecrawl no longer forces the direct-only route
Closes the residual left open by the previous commit for gateway-ready
installs (#113017 review). _get_backend() and check_firecrawl_api_key()
already use read_selection semantics, but _get_firecrawl_client() still
keyed its "stored vendor selection: direct only" branch on
selection_exists("web"), which is true for a lone per-capability key. So
with only web.extract_backend set and the Nous Tool Gateway ready (no
FIRECRAWL keys), the search ladder resolved to firecrawl via the managed
slot and the client then refused it with a misleading "web is configured
to use firecrawl (set via hermes tools)" error — the other capability
still lost its normal cascade.

The direct-only branch is now keyed on an explicit firecrawl selection
(shared web.backend via read_selection, or a per-capability key that names
firecrawl via _is_explicit_firecrawl_selection). A per-capability key
naming another vendor is not a firecrawl selection, so a ladder-derived
firecrawl resolves exactly like a never-configured install: direct when
credentials exist, else managed. Explicit firecrawl selections keep their
direct-only (keyless-unlocking) behaviour; "nous" is untouched.

Test: test_ladder_derived_firecrawl_takes_managed_path_on_gateway_ready_install
— red on the previous commit (ValueError), green with this change.
2026-09-16 17:43:40 -07:00
teknium1
f41ad3bc0a fix(gateway,cron): key Matrix handoff/cron thread seeds the way Matrix replies are keyed
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).

The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).

- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
  `get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
  use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
  `group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
  asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
  contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.

Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before  create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
        cron seed matrix🧵… ≠ inbound
after   create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)

Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
2026-09-16 17:18:44 -07:00
ea-s21
3063d0f2bb feat(matrix): implement create_handoff_thread
The Matrix adapter inherited the base create_handoff_thread (returns None), so
cron attach_to_session and the gateway handoff watcher silently no-op on Matrix
while Telegram/Discord/Slack support them: the alert/handoff is delivered flat
instead of as a replyable thread, and the scheduler logs thread_id=None ->
"Mirror: no session found". A human reply then can't resume the session with
its seeded context.

Implement it Slack-style. Matrix has no channel-level create-thread API -- a
thread is just events whose m.relates_to / rel_type: m.thread reference a root
event's event_id -- so seed a root message and return its event_id as the thread
handle. The implementation reuses the adapter's own send() (chunking + E2EE
retry + formatting) and registers the root with self._threads.mark(), mirroring
inbound thread handling; _apply_relation_metadata already threads later sends off
a supplied thread_id, so the returned id is immediately usable. Returns None on
missing client / failed send (callers fall back to parent_chat_id).

Adds tests/gateway/test_matrix_handoff_thread.py (seed id returned; default seed
on blank name; None on no client, failed send, and missing event id).
2026-09-16 17:18:44 -07:00
teknium1
44657ff628 fix(achievements): apply the sticky-unlock floor to in-flight snapshots too
_compute_from_scan(is_partial=True) evaluated against an empty unlock
set, and _run_background_scan publishes those partials to
_SNAPSHOT_CACHE every progress_every sessions. So for the duration of
every background rescan an already-earned badge rendered as locked and
unlocked_count collapsed then climbed back — the exact 39<->40 flicker
from #112273 that the finished-scan floor alone did not remove.

Read state.json for partials as well (the floor applies); partials
still never record new unlocks or save_state, since an unlock time from
half a scan could be invalidated by a later session.

Also:
- tests/plugins fake SessionDB accepts include_compacted, which the
  scanner now passes (this was the red 'Python tests' check).
- engine tests patch get_hermes_home so _data_file's legacy migration
  cannot copy the developer's real state.json into the temp dir.

Part of #112273
2026-09-16 17:03:47 -07:00
teknium1
cf86c14128 fix(achievements): read the deduped display history; drop the TypeError shim
Trim of the #112301 salvage:

- `get_messages(sid, include_compacted=True)` instead of inactive+compacted:
  that is SessionDB's deduplicated display read (compaction-archived rows,
  one representative per display identity) and it deliberately excludes
  Undo/Rewind rows (active=0, compacted=0) — work the user undid is not
  credited. The reporter's own counter-check (1,447 raw vs 607 deduped
  file-tool calls) is the reason to prefer this read. Whether rewound turns
  should ever count is a product call left open; the sticky-unlock floor
  covers the badge side of it either way.
- No `except TypeError` fallback: the flag exists on every SessionDB build
  this plugin runs against.
- Sticky-unlock block collapsed to one `result.update(...)`.
- Second invariant test on a real SessionDB: stats survive
  `archive_and_compact`, and a schema-1 checkpoint entry is rescanned.

Part of #112273 (with the two preceding commits: Fixes #112273).
2026-09-16 17:03:47 -07:00
KoNit-K
bdd290d148 test(achievements): persisted unlock survives a rescan whose live metric shrinks
Cherry-picked from #112290 (@KoNit-K): the regression test for the sticky-unlock
half of #112273. The fix itself lands via #112301's commit.
2026-09-16 17:03:47 -07:00
kyssta-exe
b41b2d3e53 fix(achievements): keep unlocked badges sticky and scan full history so rewind/compaction never shrinks totals (#112273)
The dashboard rescanned only active rows (db.get_messages defaults:
include_inactive=False, include_compacted=False) so a rewind or
context compaction retroactively removed tool calls from lifetime sums
and per-session maxima (39->40 flicker, ghost badges).  Unlock state
was recomputed from live metrics each scan; state.json only stored
unlocked_at/evidence, so already-earned badges disappeared on rescan.

- Scan the full (inactive + compacted, deduped) history so totals are
  monotonic and rewind never drops file-tool/web counts.  Dedup via
  display-generation collapse avoids double-counting compaction tails.
  Bump checkpoint schema to 2 and force a one-time rescan of v1 caches.
- Make unlocks sticky: once recorded in state.json, an achievement
  stays unlocked/discovered even if a later aggregate dips below its
  threshold, preserving first_tier for display.

Part of #112273
2026-09-16 17:03:47 -07:00
teknium1
0631710ae4 fix(disk-cleanup): protect every per-profile user tree and trim tests to two invariants
Extends the salvaged #112860 entry (`workspace`) to the whole class: `plans` and
`home` are bootstrapped alongside `workspace` by `hermes_cli/profiles.py::_PROFILE_DIRS`
and listed as user data in `profile_distribution.py::USER_OWNED_EXCLUDE`, so a
`test_*`/`tmp_*` file inside any of them is a user's file, never scratch.

Trims the contributor's five tests to two invariants: the end-to-end
post_tool_call -> on_session_end path (workspace file survives, a root-level
`tmp_scratch.py` control is still removed) and the empty-dir sweep leaving a
directory under `workspace/` alone. Documents the protected trees in the plugin
README and the built-in-plugins page.

Dropped: test_workspace_project_tree_never_tracked, test_quick_keeps_workspace_test_file,
test_root_level_test_file_still_auto_deleted (folded into the E2E test as its control).

Part of #112859.
2026-09-16 16:54:19 -07:00
kevin
3b6b0e24e4 fix(disk-cleanup): keep test_*/tmp_* files under $HERMES_HOME/workspace
``workspace/`` is a user-owned project tree: every profile bootstraps it
(``profiles.py::_PROFILE_DIRS``), ``profile_distribution.py`` excludes it as
user data, and bundled plugins keep durable state there (``google_meet``
writes auth/registry JSON under ``workspace/meetings/``).

It was missing from ``_NEVER_TRACK_TOP_LEVEL``, so ``guess_category()``
returned "test" for ``workspace/<project>/tests/test_*.py``;
``_is_auto_delete("test", age)`` accepts any age, so the ``on_session_end``
sweep unlinked those files seconds after pytest ran them green - 20 file
deletions and 476 empty-dir removals (chrome-profile data dirs among them)
over three days on one install. It was also missing from
``_EMPTY_DIR_PROTECTED_TOP_LEVEL``, so meaningful empty dirs inside a project
tree were swept too.

Add "workspace" to both sets, beside the sibling user project trees
(``projects``, ``patches``, ``skins``, ``themes``, ``contributors``) that
#75403 / #32164 / #37721 already protect. No ``tracked.json`` migration is
needed: ``quick()`` re-validates stored categories through
``guess_category()`` and drops mismatches.

Regression tests: 4 of the 5 new tests fail against the unpatched plugin,
including the end-to-end ``post_tool_call`` -> ``on_session_end`` path; the
fifth pins that genuine ephemeral test files (root-level ``test_*.py``) are
still cleaned.
2026-09-16 16:54:19 -07:00
teknium1
45ef724575 fix(tools): one "Nous Subscription" image row instead of two rows that both read active
The Portal image plugin registered `provider_name="nous"` and rendered its own
"Nous Portal (image)" picker row. Selecting it wrote the same `image_gen.provider: nous`
as the managed FAL row, so both rows showed [active] for any managed selection and the
Portal pick was actually served by FAL.

The managed row is now the only Nous row (`imagegen_backend: "nous"`); its model picker
is the FAL + Krea + Portal union, filtered to the gateways the account is funded for
(free-pool accounts see FAL only). The Portal provider stays registered for pets but
no longer renders a row. CLI picker and dashboard model endpoints share the catalog.
2026-09-16 16:42:47 -07:00
Mariano Nicolini
c13276d9df refactor(image_gen): derive the managed-Krea model set from the Krea plugin so the two lists cannot drift 2026-09-16 16:42:47 -07:00
Sora-bluesky
9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
outpoints
94efd22177 fix(honcho): preserve title provenance through upstream peer routing 2026-09-15 22:30:11 -07:00
outpoints
44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
outpoints
231e1c1204 fix(honcho): resolve sessions against agent cwd, not process cwd
The Honcho provider resolved per-repo/per-directory session names from
os.getcwd(), which is the backend process launch directory on
Desktop/gateway hosts (typically $HOME), not the user's workspace. With
a manual sessions map entry for the home directory, every Desktop
conversation landed in that fallback bucket instead of the project's
per-repo session.

Use agent.runtime_cwd.resolve_agent_cwd() — the same single source of
truth already used for system-prompt and context-file discovery — so
Honcho session routing agrees with everything else about where the
agent logically lives. It honors the pinned session cwd, then
TERMINAL_CWD, then the launch directory; CLI sessions launched inside a
project resolve identically either way.

Adds a regression covering the Desktop-style case: backend launched in
$HOME, workspace elsewhere, home-directory manual map present.

Refs #24740

(cherry picked from commit cfb32757de2509dd404acf98311e357d73fa741e)
2026-09-15 22:30:11 -07:00
outpoints
3cbdc32565 fix(honcho): don't let auto-generated session titles override sessionStrategy
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.

Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.

Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.

Fixes #24740

(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
2026-09-15 22:30:11 -07:00
teknium1
7c5296ce1c fix(buzz): a clean relay close backs off and publishes retrying like any other disconnect
A relay that accepted, authenticated and subscribed and then closed cleanly
made the read loop return without raising, so _websocket_loop reconnected
immediately with no backoff and never flipped health to "retrying". The read
loop now raises ConnectionError on StopAsyncIteration so the clean close takes
the same backoff + degraded path as an idle or send-side disconnect.
2026-09-15 18:59:25 -07:00
teknium1
f529986abf fix(buzz): a dead socket ends the WebSocket connection from either side and publishes retrying
Follow-up to the cherry-picked watchdog from #112052 (@KoNit-K), finishing the
class the reporter of #112049 laid out:

- `_websocket_loop` runs the read loop and the discovery sweep as sibling
  tasks and ends the connection when EITHER finishes. The discovery sweep
  re-raises `ConnectionClosed` instead of logging it and retrying next tick:
  a send that sees the socket closed is proof the read the loop is parked on
  will never return. That is exactly the traceback the reporter watched for
  22-86 h while inbound stayed silent.
- Health is invalidated while reconnecting: the first disconnect publishes
  `retrying` (`_mark_degraded`) and a successful re-subscribe publishes
  `connected` again. Before, `connect()` wrote "connected" once and nothing
  ever changed it, so `/health/detailed` claimed delivery during the silence.
- The teardown awaits both tasks with `gather(return_exceptions=True)` instead
  of a bare `except (CancelledError, Exception): pass`, which could swallow a
  `disconnect()` cancellation landing mid-teardown.
- Slims the salvaged read-loop hunk: the extra "receive task remained parked"
  warning and the in-loop `_mark_degraded()` are dropped; the reconnect log
  line and the loop-level health flip cover both.

Docs: the Buzz page still described inbound as poll-only and the WebSocket
transport as a future optimization; it now describes the watchdog and the
`retrying` health state.
2026-09-15 18:59:25 -07:00
KoNit-K
8f1e0ec990 fix(buzz): recover from uninterruptible websocket reads 2026-09-15 18:59:25 -07:00
teknium1
8ebc3d420f fix(feishu): warn once when the empty-allowlist default drops group messages
Under multiplex a secondary profile reads FEISHU_GROUP_POLICY from its own
secret scope only (deliberate 0.21.3 isolation), so a profile whose .env
carries no FEISHU_* policy keys falls back to `allowlist` with an empty
FEISHU_ALLOWED_USERS and every human group message is rejected while DMs
keep working. That deny was logged only at DEBUG, making it look like the
events never arrived (#111420).

Keep the scoped read as is — no environ fallthrough. Instead, the first
group drop caused by the untouched allowlist default logs once at WARNING
naming the chat and the keys to set (FEISHU_GROUP_POLICY /
FEISHU_ALLOWED_USERS in the profile's own .env, or group_rules in its
config.yaml). Operator-configured denies (populated allowlist, per-chat
rule, non-allowlist policy) and later drops stay at DEBUG. The predicate
lives in the topical sibling feishu_admission_diagnostics.py.

Docs: the Group Message Policy section now states the per-profile read and
where to put the keys under a multiplexed gateway.

Co-authored-by: bear0328 <bear0328@users.noreply.github.com>
Co-authored-by: NanPan <111261006+poijygfdyy@users.noreply.github.com>
2026-09-15 18:58:59 -07:00
teknium1
79bf0be53b fix(slack): resolve a cold channel's workspace from the sole authenticated team
scope_id_for_chat only consulted the channel→team map, which is empty right after boot (and
after a reconnect) until an inbound event from that channel arrives. A /handoff into a Slack home
without a stored scope_id (SLACK_HOME_CHANNEL env homes, or config homes never re-set via
/sethome) therefore built a key without the team while every thread reply carries it — the
handed-off thread was still orphaned across a restart (#111896).

When the map has no entry and the channel is not known to be shared across workspaces, fall back
to the single authenticated workspace (filled by auth.test at connect); multi-workspace installs
keep returning None.
2026-09-15 18:56:45 -07:00
teknium1
fb975fb098 fix: slim the Slack edit-failure log to one line via the existing payload helper
Follow-up to the salvaged #111938 commit: `_slack_response_payload` already normalizes a
SlackResponse/dict body, so the new `_slack_api_error_code` helper and the two-branch
logger.error were redundant. One log line now always carries `api_error=<code|none>` so an
HTTP 200 + ok=false failure (e.g. message_not_found) is readable without exc_info.

Tests trimmed to one invariant per fix (session key on the response-ready line; API error
code on the edit failure); the `session=unknown` fallback test was a change-detector.
2026-09-15 18:56:20 -07:00
KoNit-K
6214769189 fix(gateway): improve response and Slack error logs 2026-09-15 18:56:20 -07:00
teknium1
6062ad5aee fix(platforms): QR fallback tip names Hermes' own uv when it is not on PATH
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
2026-09-15 18:55:47 -07:00
teknium1
c2b93ae5ac fix(platforms): QR fallback tip uses uv against the running interpreter
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.

Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:55:47 -07:00
Kevin Rajan
505b36b7af fix(platforms): make QR fallback install tip target the active interpreter
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.

Fixes #111695
2026-09-15 18:55:47 -07:00
teknium1
3250020b34 fix(telegram): slim the scheduled typing re-arm and trim its tests to two invariants
Fold the three re-arm helpers (_typing_retrigger_state, _clear_typing_retrigger,
_send_typing_quietly) into _retrigger_typing itself; the semantic change from #111886
is unchanged: schedule sendChatAction as a tracked background task instead of awaiting
it on the send path, one in-flight re-arm per chat, at most one per
typing_retrigger_min_interval_seconds (2s default, matching _keep_typing), and honour
typing_indicator: false, which previously only gated the refresh loop.

Tests: keep the two invariants that are red on origin/main — an intermediate send()
returns while sendChatAction is stalled, and a 20-chunk stream costs one sendChatAction
per chat — and drop the eight change-detector variants.

Co-authored-by: aurel282 <aurelien.gek@gmail.com>
2026-09-15 18:55:19 -07:00
Aurélien Gekiere
7c10c249ce fix(telegram): schedule the post-send typing re-arm instead of awaiting it
`_retrigger_typing` awaited `sendChatAction` inline on the send path, and
streaming re-arms after *every* intermediate send. `sendChatAction` is a
fire-and-forget UI hint whose result nobody reads, but awaiting it ran its
TLS round-trip on the same event loop as the `getUpdates` long-polls.

With several agents streaming concurrently the loop stayed pinned, the
long-polls were never serviced, and they decayed into CLOSE-WAIT while the
adapter still reported `connected` — a gateway that is deaf but healthy, which
`Restart=always` cannot recover because the process never exits.

py-spy put 20 of 20 MainThread samples in `send_typing` -> `send_chat_action`
-> `start_tls`. The handshakes are what cost: with `max_keepalive_connections=4`,
a re-arm per chunk churns the pool so most calls pay a fresh TLS handshake on
the loop thread.

Three changes, all in the re-arm path:

- Schedule the re-arm as a tracked task rather than awaiting it, so a
  round-trip never delays a send or a poll. It joins `_background_tasks`, so
  shutdown cancels it and it cannot outlive the adapter.
- One in-flight re-arm per chat, and at most one per
  `typing_retrigger_min_interval_seconds` (default 2s, `extra` knob; 0 restores
  a call per send). Telegram's bubble lasts ~5s and `_keep_typing` already
  refreshes every 2s, so the re-arm only has to cover the gap left by a landed
  message.
- Honour `typing_indicator: false`. Only `_keep_typing` consulted it, so the
  documented workaround still paid for a `sendChatAction` on every
  intermediate send.

Simulating 200 streamed chunks with a 10ms loop-blocking handshake:
200 `sendChatAction` calls and 2037ms of send-path time before, 1 call and
13ms after.

Fixes #111727

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 18:55:19 -07:00
teknium1
e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1
4b00842536 fix(web): keyless MCP honours SSE CR terminators and declared charsets
Review follow-up on the keyless Exa UTF-8 fix:

- `_parse_mcp_body` split SSE frames on `\n` only, so a body using bare CR
  line terminators (permitted by the SSE spec, and parsed fine before via
  `splitlines()`) failed with "Unrecognized MCP response shape". Split on
  the SSE terminators CRLF / CR / LF instead.
- `mcp_call` decoded the body as UTF-8 unconditionally, ignoring a charset
  the server did declare. Decode with the declared charset when the
  Content-Type carries one and fall back to UTF-8 only when it is absent.
- The HTTP >= 400 branch and `_keenable_request` still surfaced
  `response.text`, so a non-ASCII error body from a charset-less text/*
  response reached the user as ISO-8859-1 mojibake. Both now go through the
  same decode helper as the success path.
2026-09-15 18:43:28 -07:00
teknium1
28cdd72815 fix(web): keyless Exa decodes its SSE body as UTF-8 so CJK searches work
Exa's MCP endpoint answers `text/event-stream` without a charset, so
`requests` decoded `.text` as ISO-8859-1: every non-ASCII character came back
as mojibake, and for CJK results the UTF-8 continuation byte 0x85 became
U+0085, which `str.splitlines()` treats as a line break — the `data:` JSON
line was cut in two and the perfectly valid result surfaced as "Unrecognized
MCP response shape", after which the ring fell through to keyless Firecrawl
(403). Decode the bytes as UTF-8 (JSON-RPC and SSE are UTF-8 by spec) and
split SSE frames on newlines only. A parsed envelope that really carries no
text is now reported as "no text content" instead of an unrecognized shape.
2026-09-15 18:43:28 -07:00