Commit Graph

35004 Commits

Author SHA1 Message Date
teknium1
5bccc4e238 fix(model_switch): clear key_env only when the route changes; drop it on custom activation
`model.key_env` is not custom-only: the Desktop settings UI stores REGISTRY
provider keys there (e.g. HERMES_CUSTOM_LMSTUDIO_API_KEY with provider
lmstudio, #106336) and auth._model_level_key_env honours it. The previous
predicate (`not custom or route_changed`) therefore wiped that pointer on a
same-provider same-base_url model re-pick and silently broke the user's
credential. The pointer now clears ONLY when provider or base_url changed;
the inline api_key/api rule is unchanged.

Custom-endpoint activation (model_setup_flows_custom) popped base_url /
api_key but never key_env, so a stale pointer from a previous endpoint
outranked the credential it had just written — pop it alongside.

Tests: the same-route re-pick case now covers a registry provider (red on
the old predicate), and the two clear_model_endpoint_credentials tests are
folded into one invariant.
2026-09-15 03:39:45 -07:00
teknium1
1987b5e461 fix(model_switch): provider switch clears model.key_env/api_key_env in the canonical persist shape
Main no longer clears credentials in web_server's _apply_main_model_assignment;
every /model surface (CLI, gateway, TUI, dashboard) persists through
model_selection_config_updates(), which only dropped api_key/api on a route
change. Custom-endpoint activation writes model.key_env with NO inline key, so
a pointer-only model block survived every provider switch and routed the new
provider's requests to the old endpoint's env var (the PR's Bug 5) — live
repro on main: activate custom_myep -> /api/model/set openrouter left
key_env: CUSTOM_MYEP_API_KEY on disk.

Port the PR's web_server hunk to the one shape function: key_env/api_key_env
clear under the same route-changed rule as api_key (same-route re-pick keeps
them). The dashboard's _resolve_assignment_credentials re-adds the TARGET
provider's own pointer after this, so custom->custom still ends with the new
endpoint's key_env. One invariant test in the one-shape suite.
2026-09-15 03:39:45 -07:00
Teknium
68ebb2c996 fix(config): provider switch clears the stale model.key_env pointer
Custom-endpoint activation writes model.key_env, but
clear_model_endpoint_credentials never popped it and the switch-clears
trigger only fired on inline keys — a pointer-only model block survived
every provider switch, routing the new provider's requests to the old
endpoint's env var. key_env/api_key_env now clear under clear_api_key;
the trigger fires on pointer-only blocks too. All key_env writers set it
after the clear, so custom-to-custom switches keep the new pointer.
2026-09-15 03:39:45 -07:00
teknium1
88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
Hermes Agent
e27f16365d fix(agent): preserve streamed refusals as text (port of anomalyco/opencode#43343)
A model that declines mid-stream delivers the explanation on the
structured refusal channel (chat_completions delta.refusal; Responses
response.refusal.delta / refusal content parts) and leaves content
empty. The streaming accumulators dropped that channel entirely, so a
streamed refusal assembled into an empty message and fell into the
empty/invalid-response retry loops - burning paid retries reproducing a
deterministic refusal - while the non-streaming path had already fixed
this class in #46013.

- chat_completions streaming: accumulate delta.refusal (incl.
  model_extra), expose message.refusal on the assembled mock response so
  ChatCompletionsTransport.normalize_response applies the existing
  sole-payload -> content_filter promotion; count refusal deltas in the
  zero-chunk guard; carry refusal in the Relay final-response dict.
- Codex Responses stream consumer: collect response.refusal.delta as
  answer text so a refusal-only stream no longer raises 'did not emit a
  terminal response' with zero usable content.
- Responses normalizer: read type=refusal content parts in
  _extract_responses_message_text (attr and dict shapes).

Sabotage-verified: each new test fails with its wiring line disabled.
E2E: refusal-only stream -> terminal content_filter with explanation;
refusal-alongside-content stays a normal usable turn; plain-text
streams unchanged.
2026-09-15 03:39:07 -07:00
teknium1
021ab58a23 fix: hindsight update-time dep resolution survives a BOM in config.json
`_provider_pip_dependencies` still read ~/.hermes/hindsight/config.json with
strict utf-8 inside a bare `except Exception`, so a Windows-editor BOM made the
`mode` lookup silently fail and `hermes update` reinstalled only
`hindsight-client`, leaving the embedded daemon broken — the exact #70636
symptom this helper exists to prevent. Route it through the shared
`read_json_or_empty` (utf-8-sig, {} on missing/corrupt) like the other memory
readers in this PR, and add the reader to the parametrized BOM invariant.
2026-09-15 03:38:29 -07:00
teknium1
c97e36db3f test: one parametrized BOM invariant over the shared reader + 2 direct readers
mem0, hindsight and the honcho CLI all read through utils.read_json_or_empty,
so the seven per-loader tests collapse to one parametrized invariant plus the
Qwen creds case. The per-loader fix landed in the shared reader (one site, not
six), which is the salvage bar: <=2 invariant tests, no change-detectors.
2026-09-15 03:38:29 -07:00
Teknium
5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00
teknium1
e860b8e4e4 fix(context): compute-host /context and session.context_breakdown carry the per-file manifest; report blocked files
Why: the tui_gateway live formatter (`_format_live_context_output`, used when
the session runs on a compute host) renders its own summary and never got the
"Context files" block, and `session.context_breakdown` had no structured rows,
so Desktop's popover could not show them. The formatter now appends
render_context_file_lines() with the session cwd bound (the RPC thread has no
session context, so the discovery walk would key on the backend's cwd), and
the RPC payload gains a `context_files` list (contract + generated TS/OpenRPC
+ Desktop type). The docs sentence is scoped to the surfaces that render it.

A file whose content _scan_context_content replaces with a BLOCKED marker was
reported "loaded"; the manifest now runs the same scan and reports `blocked`.
The module docstring names the frontmatter-strip / chain-cap approximations
and drops the product-name attribution (credit stays in the PR body).
2026-09-15 03:37:49 -07:00
teknium1
f271ba09b0 fix(context): derive the /context file listing from the builder's own discovery walk
Review follow-up (Enough1122) on the salvaged #91272: the original
list_context_file_sources() hand-mirrored the priority ladder inside
build_context_files_prompt, so the two would drift the moment the builder
gained a context type or changed precedence — misreporting what the prompt
holds is worse than not showing it.

Now prompt_builder exposes one candidate finder per context type
(_CONTEXT_FILE_CANDIDATES → discover_context_files) and BOTH the loaders and
the manifest walk it. The manifest lives in the new sibling
agent/context_file_sources.py (not appended to the facade) and:
- reports empty / unreadable files truthfully instead of "✓ 0 tokens",
- mirrors the install-tree guard ("suppressed") so a Desktop session that
  fell back into the Hermes tree sees why nothing loaded,
- lists every .cursor/rules/*.mdc as loaded, matching the builder which
  concatenates all of them,
- measures truncation on the rendered "## label" section like the builder.

The block now renders on every surface that shows the /context category
table: CLI/TUI (hermes_cli/cli_info_mixin.py) and the messaging gateway
(gateway/slash_commands_status.py). The Desktop popover consumes the raw
session.context_breakdown payload (no text table) and is left as-is.

Tests trimmed to the two invariants: manifest/prompt parity across every
context type at once, and truncated/suppressed follow the builder.
2026-09-15 03:37:49 -07:00
Teknium
5349aa609d Inspired by Copilot CLI: /context now lists each context file with load status and token cost
Copilot CLI 1.0.81-6 shows each user instruction file separately in
/instructions. Hermes loaded AGENTS.md/.hermes.md/CLAUDE.md/.cursorrules/
SOUL.md through a priority ladder but gave the user no visibility into
WHICH files were discovered, which one won, which were shadowed, or how
much context each costs — the /context 'rules' category was one opaque
number.

- agent/prompt_builder.py: list_context_file_sources() — read-only
  manifest mirroring build_context_files_prompt discovery (priority
  ladder, AGENTS.md directory chain with AGENTS.override.md precedence,
  cwd-only CLAUDE.md/.cursorrules, SOUL.md from profile home) with
  per-file chars, est_tokens, and loaded/truncated/shadowed status
- cli.py /context: 'Context files' section rendering the manifest with
  status glyphs and shadowing/truncation notes; zero prompt/cache impact
- docs: reference/slash-commands.md /context row
- tests/agent/test_context_file_sources.py: 11 tests incl. E2E parity
  with build_context_files_prompt shadowing
2026-09-15 03:37:49 -07:00
kshitijk4poor
e112da7578 docs(desktop): preflight comments describe the OAuth-only branch; scope assertion via objectContaining 2026-09-15 16:07:43 +05:30
kshitijk4poor
2821cb2d8c refactor(desktop): one profile-order sort for the active strip and the at-rest groups
sortByProfileOrder moves to a pure lib module with a key selector so
buildRestGroups sorts its named squares directly, replacing the collator
sort that the component then re-sorted with a different comparator. The
fleet rail test no longer needs importOriginal (and three store mocks) to
reach the helper.
2026-09-15 16:07:43 +05:30
kshitijk4poor
8751b3edd4 test(desktop): keep the unscoped getProfiles invariant separate from the scoped one 2026-09-15 16:07:43 +05:30
kshitijk4poor
9bb985c5d4 fix(desktop): OAuth REST preflight gets its own dial budget
Sharing the remainder of SWITCH_DIAL_TIMEOUT_MS meant a slow-but-successful
socket dial left the preflight 0 ms and the switch failed as "Timed out
connecting" although the socket had just opened.
2026-09-15 16:07:43 +05:30
kshitijk4poor
b591df7c42 fix(desktop): at-rest rails read the roster only; active gateway keeps its single render path
The renderer's per-connection list no longer feeds buildRestGroups: with no
roster reconciler it could outlive a profile deleted elsewhere. The active
gateway is never rendered through FleetRestGroup — that path routed its own
squares through selectConnection (full dial + wipe) instead of selectProfile's
live swap. What survives from the original change is the order parity: at-rest
named squares follow $profileOrder like the active strip.
2026-09-15 16:07:43 +05:30
kshitijk4poor
256edfc1f5 fix(desktop): profile-list ownership follows the published source, without a roster reconciler
The per-connection list cache is kept only to repaint $profiles on re-home.
The $fleetRoster listener is dropped: a roster landing while the active
source's own /api/profiles read was in flight invalidated that read, so
$profiles stayed empty/stale after every switch or focus refresh that the
roster IPC won. Both are reads of the same backend; neither is "older".

A null descriptor is a reconnect blip (setConnection's contract) and keeps
the current owner instead of blanking the rail; the first published
descriptor adopts whatever list is already loaded. Legacy sources are keyed
by endpoint rather than a JSON tuple. Tests cover the roster race and the
null blip; the two use-session-actions tests now publish the descriptor
before seeding $profiles, matching the runtime order.
2026-09-15 16:07:43 +05:30
Zeus-Deus
5b49051276 fix(desktop): keep profile switches on the selected gateway 2026-09-15 16:07:43 +05:30
kshitijk4poor
6cd9fe1d5b chore: map Zeus-Deus contributor email (salvage #105139) 2026-09-15 16:07:43 +05:30
teknium1
d8053f4806 fix(video_gen): cap LTX 2.5 at 10s for 1440p/2160p; omit unset enum durations
fal's LTX 2.5 fast endpoints accept 6-20s only up to 1080p — "At 1440p and
2160p, all frame rates support up to 10 seconds" — so a 4K request with the
family's 20s ceiling was rejected by the vendor. Families can now declare
`duration_cap_by_resolution`, applied after the enum snap / range clamp on the
resolved resolution enum.

An unset duration on a duration_enum family also snapped to enum[0] (6s),
silently overriding the endpoint's own "auto" default; None now omits the key
for enum families exactly as it already did for range families.

test_managed_media_gateways asserts the alibaba/happy-horse/ namespace by
prefix rather than the exact v1.1 literal so the next version bump doesn't
flip an unrelated gateway test.
2026-09-15 03:37:11 -07:00
teknium1
1c1980dc1f fix(video_gen): make the duration enum explicit instead of sniffing tuple shape
`durations` carried two meanings told apart only by len==2 and gap>1: a
(min, max) range to clamp, or an enum to snap. A family with exactly two
legal values would have been misread as a range (review finding on #91311).
`durations` is now always the (min, max) window (what capabilities()/
list_models() read) and families with discrete values add `duration_enum`;
_clamp_duration takes the family and branches on the key, not the shape.

Also: restore the exact v1.1 endpoint assertion in the gateway namespace
test (a startswith/endswith check would not catch a silent version drift),
add the ltx-2.5 i2v snap case, and keep happy-horse on audio_native (the
schema test forbids audio+audio_native together, and v1.1 audio is always on).
2026-09-15 03:37:11 -07:00
Teknium
b30df1303c test: pin happy-horse namespace invariant, not exact 1.0 endpoint strings
The managed-gateway test asserted the literal v1.0 endpoint ids; its
stated purpose is verifying the alibaba/ (not fal-ai/) namespace. Assert
prefix+modality-suffix instead so version bumps don't break it.
2026-09-15 03:37:11 -07:00
Teknium
37286d3064 feat(video_gen): LTX 2.5 + Kling O3 families; Happy Horse upgraded to v1.1
Adds two new FAL video families and upgrades one:

- ltx-2.5 (cheap tier): lightricks/ltx-2.5/{text,image}-to-video/fast.
  Lightricks' open-source audio-video model. Native audio, 6-20s integer
  duration enum, 720p-2160p (i2v), $0.09/s at 720p. duration_int + 2k/4k
  resolution aliases; no seed key in the schema.
- kling-o3 (premium tier): fal-ai/kling-video/o3/standard/{text,image}-to-video.
  Kuaishou's frontier multi-shot model, 3-15s, optional native audio
  ($0.084/s off, $0.112/s on). String durations, i2v drops aspect_ratio,
  no seed/resolution keys.
- happy-horse upgraded from the sparse-docs 1.0 endpoints to
  alibaba/happy-horse/v1.1/{text,image}-to-video with the full published
  schema: nine aspect ratios, 720p/1080p, 3-15s integer durations, seed
  supported, audio native (no generate_audio key), i2v drops aspect_ratio.

All flags derived from each endpoint's llms.txt schema. Payload builder
asserted locally against the schemas; test for the old Happy Horse
"prompt-only" contract updated to pin the v1.1 schema, plus new payload
tests for ltx-2.5 and kling-o3.
2026-09-15 03:37:11 -07:00
teknium1
bdc7916196 fix(ux): plain-language, actionable user-facing messages (dashboard)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 03:36:22 -07:00
teknium1
754a9ef4e7 test(stt): give the silent-stall test the same idle budget
test_silent_stall_still_times_out kept the 0.1s idle window and flaked
once locally under heavy load (load avg ~250): the child was killed
before its single stderr line was read, so the pre-stall-output
assertion saw an empty stderr. Same class as the progress test this PR
de-flakes; give it the same 0.25s budget. The 30s sleep still trips the
window, so the must-time-out direction is unchanged.
2026-09-15 03:36:11 -07:00
teknium1
7ee9fa52c0 test(stt): correct the de-flake provenance in the idle-timeout comment
andrexibiza's review is right: the base test already printed tick 0
before its first sleep, so this change never removed a
sleep-before-first-tick race. The actual de-flake is the larger idle
budget (0.1s -> 0.25s) plus a longer heartbeat sequence whose ~400ms
runtime still exceeds the idle window. Say exactly that so the causal
record is accurate.
2026-09-15 03:36:11 -07:00
andrexibiza
a70ea873cf test(stt): stabilize idle-timeout progress test against spawn latency
The stderr-progress idle-timeout test used a 0.1s idle window with 0.04s
ticks — shorter than Windows process spawn, so the first chunk could never
arrive in time (deterministic failure on Windows, flake under Linux CI
load). Verified failing identically on pristine main before the change.

Fix: emit the first tick immediately, tick every 50ms for ~400ms total,
250ms idle window (5x tick period). The pass still depends on the progress
extension while tolerating real spawn/scheduling latency.

Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
2026-09-15 03:36:11 -07:00
teknium1
b00ebf6210 fix(skills): live-dashboard credits the human author and drops the product-name intro
Authoring standard 4 requires the human first in `author`; the skill was
drafted with Hermes so the tool was credited instead. The intro cited a
third-party product, which is allowed only in LICENSE/credit lines. Also
adds metadata.hermes.category and lowercases tags to match sibling
skills; docs page regenerated for this skill only.

Tests: the two prose tests asserted sentence literals (change-detectors);
they now assert structure — three Procedure phases, a "Done when" per step,
standard headings, and the tool wiring (cronjob/desktop_preview/[SILENT]).
2026-09-15 03:35:33 -07:00
teknium1
788601358d refactor(skills): live-dashboard becomes an optional skill reconfigured for the Desktop app
Why: the web dashboard is being deprecated in favour of the Electron Desktop
app, and a skill that ships a bespoke cron blueprint should not be bundled by
default. Reconfigure instead of just rebasing:

- Move skills/productivity/live-dashboard -> optional-skills/productivity/
  live-dashboard (install with `hermes skills install
  official/productivity/live-dashboard`); register it like every other
  optional skill: per-skill docs page under user-guide/skills/optional/,
  optional-skills-catalog row, sidebars entry.
- Desktop reality: add a "Show the dashboard" step — when `desktop_preview`
  is in the toolset (Desktop/GUI sessions) render index.html in the in-app
  preview pane after every build/tick and on request; otherwise report the
  absolute file path. Prerequisites section added (Enough1122 review).
- Drop the hard-wired cron/blueprint_catalog.py entry: the curated catalog
  is for bundled skills and would preload a skill that may not be
  installed. Use the skills-pipeline blueprint instead —
  `metadata.hermes.blueprint` on the SKILL.md registers a daily
  all-dashboards sweep as a /suggestions entry at install time (opt-in,
  never auto-scheduled), which is exactly the mechanism main provides for
  optional skills.
- Never hardcode ~/.hermes in prose the agent executes: refer to the Hermes
  home directory's dashboards/<slug>/ and write absolute paths into cron
  prompts (also answers the review's "tick prompt must name the state-file
  path" point).
- Tests follow the skill to optional-skills/, the catalog-blueprint tests
  are replaced by one parse_blueprint/blueprint_to_job_spec invariant and
  one desktop_preview-with-path-fallback invariant.
2026-09-15 03:35:33 -07:00
Teknium
2154458685 Inspired by Energy: one-sentence live dashboards — bundled skill + automation blueprint
Energy (getenergy.com) ships natural-language persistent dashboards:
describe what you want to see in one sentence and the agent builds a
self-updating status page fed by email threads, signed-in websites, and
files. This ports the concept onto Hermes's existing cron + connector
architecture:

- skills/productivity/live-dashboard: setup/tick split skill — pin the
  dashboard contract, verify one live read per source before scheduling,
  keep dashboard.json as source of truth with a self-contained HTML
  projection, stale-read discipline, deliver only on material change.
- cron/blueprint_catalog.py: live-dashboard automation blueprint
  (purpose/sources/time/recurrence/deliver slots) rendering to the
  dashboard form, /blueprint command, and hermes:// deep-link.
- tests/skills/test_live_dashboard_skill.py: skill standards + blueprint
  registration + real fill_blueprint E2E.
- docs: per-skill page, skills catalog row, sidebar entry.
2026-09-15 03:35:33 -07:00
teknium1
59c20a51aa test: drop product name from inbox-triage test docstring
Competitor product names are allowed only in LICENSE attribution and
salvage credit lines; the test docstring had an inspiration tag left over
from the port.
2026-09-15 03:34:55 -07:00
teknium1
0a208b6891 docs(skills): itemize inbox-triage voice calibration and bound its context cost
Why: the calibration step was a single ~120-word paragraph that models follow
less reliably than enumerable rules, and 20-50 full sent messages could
crowd inbox coverage out of context (Enough1122 review). Split sampling /
extraction / record / fallback into bullets, state that truncated excerpts
carry the style facts, add the Sent-folder-naming pitfall (`Sent`, `Sent
Messages`, `[Gmail]/Sent Mail`, localized) so a missing folder name does not
silently trigger the fallback, and compress the provenance parenthetical to
the operative fact. Docs page mirrors the SKILL.md body verbatim.
2026-09-15 03:34:55 -07:00
Teknium
7f2bf68308 Inspired by Energy: evidence-based voice calibration for inbox-triage reply drafts
Energy's cross-app reply agent analyzes ~100 of the user's past replies
before drafting, so drafts land in the user's actual voice instead of
generic-professional AI register. Port the mechanism into the
email-inbox-triage skill's drafting step:

- Step 4 now calibrates on a bounded sample (20-50) of the user's own
  sent replies — greeting/sign-off habits, length, formality, rhythm,
  per-audience differences, how the user pushes back — before drafting,
  with an explicit fallback when Sent is empty or inaccessible.
- New pitfall + verification item pinning the calibration discipline.
- Test locks the evidence-based calibration and fallback language.
2026-09-15 03:34:55 -07:00
teknium1
286e723db8 fix(skills): auto-load resolves under the agent's own home and skips internal forks
Why: build_auto_load_prompt read config via ambient load_config_readonly()
and looked skills up under the ambient SKILLS_DIR. Gateway bot threads lose
the HERMES_HOME ContextVar, so a bot profile's pinned skills came from the
launch profile — the docs promise profile scoping. _auto_load_parts now
passes home_override=_agent_home(agent) and build_auto_load_prompt binds it
for config, disabled-list and <home>/skills lookup, the same seam
_skills_prompt uses via skills_dir_override.

_auto_load_parts was unconditional and injected pinned SKILL.md bytes into
delegate children, curator/background_review forks and gateway hygiene
agents; it now mirrors _skills_prompt's gate (nothing without the skills
toolset) and returns [] when skip_context_files is set.

cli.py's HERMES_IGNORE_RULES check used == "1" while system_prompt used
is_truthy_value; both use is_truthy_value now.

Tests stay at 4: the build test asserts the home-scoped resolution, the
ignore-rules test also covers the subagent / no-skills-toolset gates.
2026-09-15 03:34:20 -07:00
Carl Taylor
075a256597 fix(skills): auto-load resolves once per agent and dedupes against -s
The system prompt must stay byte-stable for the life of a conversation:
`_auto_load_skills_result` is seeded in `_SESSION_STATE` and filled on
the FIRST prompt build only (HERMES_IGNORE_RULES captured then too), so
model switches, compression and static-prefix restoration reuse the
exact rendered bytes rather than re-reading config or skill files.

CLI: auto_load renders in the existing background `--skills` preload
thread (real session id for ${HERMES_SESSION_ID}), `-s` names dedupe
against the auto-loaded canonical names via
`build_preloaded_skills_prompt(excluded_loaded_names=)`, the activated
skills line shows auto_load first, and the lazily built agent is seeded
with the pre-resolved bytes. `--ignore-rules` skips auto-load with the
rest of the auto-injected context.

Re-implementation of #74060 by @ctaylor86 against current main.
2026-09-15 03:34:20 -07:00
ArcherQAQ
1976869c01 feat(skills): skills.auto_load pins skills into every new session's prompt
`skills.auto_load: [name, ...]` in config.yaml renders the listed skills
as fully loaded skill blocks in the system prompt of every new agent —
CLI, TUI, gateway, cron and API — the persistent counterpart of `-s`.
Missing or operator-disabled names are warned about and skipped; a
config typo never blocks session start.

Re-implementation of #26840 by @ArcherQAQ (via #74060) against current
main — the original patch targeted the pre-decomposition cli.py /
system_prompt.py god files; design and diagnosis preserved. The loader
reuses `_load_skill_blocks` (same disabled gate and Curator usage bump
as `-s`) instead of a parallel loop.
2026-09-15 03:34:20 -07:00
teknium1
bbe13059bc fix: gate kanban.default_assignee through the dispatch_profiles allowlist
_resolve_default_assignee imported the raw profile_exists instead of the
allowlist-gated predicate returned by _profile_exists_fn(). On a shared
board, a home with dispatch_profiles: [sage] and default_assignee: default
therefore passed the gate for "default" (every home has that profile), wrote
assignee=default plus an `assigned` event onto an unassigned card it may not
claim, and only then bucketed it skipped_nonspawnable. That is a row write on
a card another home owns. Routing through _profile_exists_fn() makes an
out-of-allowlist default_assignee resolve to None, so the card is left
untouched; the unimportable-profiles fallback is unchanged.

Review finding: default_assignee bypassed the dispatch_profiles gate and persisted assignee + assigned event onto unclaimable shared-board cards.
2026-09-15 03:33:53 -07:00
teknium1
923960fa4b fix(kanban): dispatch_profiles is config-only and fail-closed; trim tests
Follow-up to the salvaged #111004 commit, aligning it with the shape agreed
on #110995:

- Drop the HERMES_KANBAN_DISPATCH_PROFILES env bridge: non-secret behaviour
  lives in config.yaml only, like every other kanban.* key.
- Read the key via load_config_readonly() with the same fail-open config
  read as the sibling kanban.* readers (configured_max_in_progress).
- Fail closed when the key is set: the "none" sentinel is gone (an empty
  list already claims nothing), and an assignee that is not a valid
  profile id is never claimable instead of being lower-cased into the
  allowlist.
- Trim the regression file to two invariants (allowlist without `default`
  buckets the card as nonspawnable AND turns has_spawnable_ready off;
  unset key keeps upstream behaviour). Both drive the real dispatch tick
  against a real config.yaml + kanban.db; the first is red on origin/main.
- Docs: move the "Shared boards across homes" section out of the
  gateway-dispatcher paragraph, state that `default` collides by
  construction, add the config-reference row.
2026-09-15 03:33:53 -07:00
Kevin Rajan
2d46af3fa2 fix(kanban): per-home dispatch claim allowlist for shared boards
On a shared kanban.db, every home's profile_exists('default') is
unconditionally True, so any home's dispatcher could claim cards assigned
to 'default'. Wrap the _profile_exists_fn() predicate with an optional
per-home allowlist: kanban.dispatch_profiles (config.yaml, list or
comma-separated string) with a HERMES_KANBAN_DISPATCH_PROFILES env bridge.
Unset preserves upstream behavior; 'none' claims nothing. Foreign
assignees land in the existing skipped_nonspawnable bucket, and the gate
applies to the ready spawn path, _has_spawnable, and review dispatch alike.
Also documents the multi-home default collision in the kanban user guide.

Fixes #110995
2026-09-15 03:33:53 -07:00
teknium1
931387dd42 test(cron): two invariant ESTOP tests for the fire webhook and misfire backstop; docs
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.

Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
2026-09-15 03:33:13 -07:00
JulianCruzet
05e13a1fcc fix(cron): address Xipong's review nits on #110563
Two inline test-hardening suggestions + two follow-up coverage cases
(Xipong called them non-blocking but they're cheap and prove the
boundaries).

- tests/gateway/test_api_server_jobs.py::TestCronFireEstop::
  - test_fire_webhook_returns_503_when_estop_engaged: also patch
    `cron.scheduler_provider.resolve_cron_scheduler` and assert that
    `claim_fire`, `fire_claimed`, and `fire_due` are never reached.
    A 503 by itself does not prove admission never happened.
  - test_fire_webhook_401_when_verifier_crashes (new): crashing
    verifier → 401, ESTOP is never consulted. Proves auth runs before
    the ESTOP check so the sentinel state cannot leak to unauth callers.

- tests/cron/test_misfire_catchup.py::TestFireOverdueJobs::
  - test_estop_engaged_skips_backstop: spy on `provider.claim_fire`
    and `cron.scheduler_provider.threading.Thread`. Deterministic
    proof that no claim was attempted and no worker thread was
    constructed (the prior `wait_fired(timeout=0.5)` was timing-based).
  - test_estop_release_restores_backstop: spy on `provider.claim_fire`
    on the recovery path with an explicit `assert_called_once` so a
    regression that fires without claiming still fails the test.

All 58 tests across the affected files pass.
2026-09-15 03:33:13 -07:00
JulianCruzet
95fd5d6816 fix(cron): honor ESTOP in NAS fire webhook and misfire backstop
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:

- Door 2 (NAS fire webhook `_handle_cron_fire`,
  `gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
  the JWT/drain guards and `provider.claim_fire`. Added a
  `check_paused("cron-webhook")` guard inside the reservation block,
  returning 503 + Retry-After so the NAS retries after `hermes resume`
  rather than silently dropping the run.

- Door 3 (misfire backstop `fire_overdue_jobs`,
  `cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
  the function. Added an early-return `check_paused("cron-misfire")`
  guard. Self-healing — the next sweep after `hermes resume` catches
  everything up via the existing claim_fire path.

Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.

Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.

Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
2026-09-15 03:33:13 -07:00
kshitijk4poor
fef0bc56b2 refactor(cli): one notification payload, one sink per branch in _ring_bell
Review fold: build "\a" + sequence once and dispatch it to exactly one sink (app loop
when the Application runs, write_tty otherwise) instead of two branches with two writes
and two try/excepts. terminal_notify.notify() had a single caller (that fallback), so
write_tty is the public no-app entry and notify() is gone. The loop-side write keeps the
never-raises contract on a dead tty (EIO/closed file) the same way _pet_flush_kitty_frame
does, and _ring_bell skips entirely once _terminal_io_broken is set.

pty A/B re-run on this stack: 400 rings vs 90 frames, 0 aborted, 0 painted, 400/400 delivered.
2026-09-15 14:52:26 +05:30
kshitijk4poor
6b69d0f6fd fix(cli): /bg completion rings through _ring_bell like every other bell site
The side-worker (/bg, /btw, /login) hand-rolled its bell with a bare sys.stdout
'\a' and so never emitted the OSC 9 / OSC 777 desktop notification the other
end-of-work sites get. Route it through _ring_bell, which also owns the
bell_on_complete gate and the app-loop serialization.
2026-09-15 14:52:26 +05:30
kshitijk4poor
ac2359cd35 fix(cli): turn-end notifications no longer paint the pet's kitty frame as base64
Symptom (Ghostty, display.pet on, display.bell_on_complete on): at the end of a turn the
input line fills with several rows of base64 and a stale copy of the status bar + pet
stays above the response panel.

Root cause: _ring_bell runs on the agent thread and terminal_notify wrote the OSC 9 /
OSC 777 sequence through its own open("/dev/tty") (or sys.stdout). The prompt_toolkit
loop thread may at that moment be mid-write of a 12 KB kitty APC pet frame, which the tty
drains ~1 KB at a time. The second writer splices into the frame; the foreign ESC aborts
the APC and the terminal paints the remainder of the payload as text at the input cursor.
The wrapped garbage scrolls the screen, so the panel that follows is printed against a
stale cursor position and the old chrome survives above it.

Change: when the CLI's Application is running, _ring_bell hands "\a" + the notification
sequence to the app loop (_run_on_app_loop -> _write_terminal_sequence), serializing it
behind the renderer and the after_render frame writer. terminal_notify.notify() keeps the
/dev/tty path for callers without a running app; the sequence builder is split out as
notification_sequence().

Verification: pty A/B with the real Application + after_render frame writer, 400 rings
vs 91 frames — base: 3 leaks (4,324 base64 chars painted); fixed: 0 leaks, 400/400
notifications delivered, 0 aborted frames.
2026-09-15 14:52:26 +05:30
kshitijk4poor
2932195c22 fix(skills): give the provider cut one owner inside the parallel walker
The dashboard endpoint GET /api/skills/hub/search passes its user-supplied
`source` straight into parallel_search_sources and never applied the merged
provider cut, so ?source=nvidia returned a mixed set. That was the fourth
caller of the walker; the cut was copy-pasted at three of them and missing
at the fourth.

parallel_search_sources already computes the normalized provider filter, so
the cut now lives there — applied per source before results are counted and
merged. Every caller (CLI search via unified_search, CLI browse, TUI-gateway
browse, dashboard router) sees the same rule with no provider logic of its
own, source_counts stop reporting rows that are then dropped, and the three
duplicated call-site cuts are deleted. do_browse keeps its provider-specific
"No skills found for provider" message.

Also:
- HermesIndexSource.search now treats a whitespace-only provider_filter as
  "no filter", matching GitHubSource.search (the two adapters previously
  disagreed on the same keyword argument; unreachable through the walker,
  which pre-normalizes).
- The regression-test fixture seeds tap caches by github_provider_for label
  instead of case-sensitive repo literals, and serializes metas through
  _skill_meta_to_dict, so a DEFAULT_TAPS casing change can no longer silently
  unseed the fixture.

Validation: 121 targeted tests green; disabling the walker cut fails the
pre-existing test_unified_search_provider_filter_keeps_index_source with the
expected clawhub leak; 4/4 regression cases still red on unpatched main.
2026-09-15 14:15:59 +05:30
kshitijk4poor
69d181011b refactor(skills): skip wrong-provider taps and give the provider-filter idiom one owner
Follow-up polish on the provider-filter-before-limit fix:

- GitHubSource.search now skips taps whose repo maps to a different provider
  instead of enumerating every tap and filtering afterwards. A tap's repo fixes
  the provider of every result it yields (github_provider_for is the only source
  of extra.provider in this adapter), so the skip is lossless and avoids up to 23
  useless tap enumerations per provider-filtered search — real GitHub API calls
  against the 60/hr unauthenticated budget and the 30s overall timeout when the
  index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
  it replaces the four inline copies of the strip/lower/membership idiom in
  _select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
  the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
  pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
  provider cut as do_browse; it accepted a provider value but returned
  unfiltered results.

Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
2026-09-15 14:15:59 +05:30
Danylo Borodchuk
e75b5a2d38 fix(skills): filter providers before limiting search results 2026-09-15 14:15:59 +05:30
kshitijk4poor
a55c972e09 refactor(gemini): prefix rule only for no-id slot matching; share the provider-id guard
Slot arguments are always complete json.dumps output (Gemini re-sends full args), so the mid-stream JSON check could never fire; the id key needs no tool name; _new_call_id and the slot lookup now share _provider_call_id.
2026-09-15 13:27:24 +05:30
jmiguellucas
91665b5e79 fix(gemini): give each native tool call its own streaming slot
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.

Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.

Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
2026-09-15 13:27:24 +05:30