- appChromeBlockedTimers: 'status-chrome timers track the current overlay
model' (#12463). Each overlay state is mounted on the real StatusRule: the
in-flow prompts (approval, billing, clarify, confirm, secret, subscription,
sudo), agents/journey and the ambient dock keep the 1s clocks running, and
the occluding set pauses them. Gating clarify/ambient in
$isStatusRuleOccluded turns it red.
- appChromeBlockedTimers: the existing 're-syncs the elapsed read-outs'
test flaked under load (a fixed 20ms flush); it now polls for the reveal frame.
- appChromeStatusRule: session title renders without a raw accent background
fill (#82465 contrast bug).
- appChromeStatusRule: idle-since read-out is shown when idle, hidden while busy
and before the first turn completes.
#120071 deleted hud-modifier-native.test.ts (darwin-only, so it ran nowhere:
js-tests is ubuntu-only) and its native fixtures. hud-modifier-gesture.h is
the clean-tap state machine both the macOS and Linux XI2 helpers include and
is portable C, so the restored test compiles hud-modifier-gesture.test.c with
cc (xcrun clang on macOS) and runs it inside the electron vitest project,
which js-tests runs via check:test:desktop:platforms. Loosening the 500 ms
tap bound in the header turns it red.
The macOS event-adapter (.m), Windows C# (.cs) and node:test (.mjs) files are
not restored: they need a macOS/Windows toolchain and no JS lane runs there.
composer-text-guard.test.tsx is rebuilt on the REAL useComposerDraft hook (the
deleted file tested an in-test copy of the helper): the mount-time draft
restore must not throw while assistant-ui's composer core is unbound, and must
write through once bound (#49903 / #49488). Removing the hook's try/catch
turns it red.
The two inline-code tests render the real component with the real
styles.css injected and read the computed style of the rendered <code>; they
are cascade outcomes, not source greps. Re-scoping the inline-code rule to the
assistant slot (the original bug) turns both red.
- system-message.test.tsx: background report inline code uses the chat
inline-code tokens (#107486)
- group-chat-view.inline-code.test.tsx: room message inline code uses the chat
inline-code tokens (#114086)
Each entry point forwards spawnPriority through its own branch to the
Electron bridge (getConnection / getConnectionFor); nothing else asserts the
IPC payload for these doors. Sabotage check: dropping spawnPriority from
retainGatewayForAgent's plain-profile branch turns the null-connection test red.
Restored (apps/desktop/src/store/gateway-spawn-priority.test.ts):
- openGatewayForProfile without a priority never tags a dial as foreground
(pre-warms stay background slots, #102281)
- openGatewayForAgent forwards spawnPriority to every registry dial (#102281)
- retainGatewayForAgent tags the lease dial when the caller says foreground (#105104)
- retainGatewayForAgent without options never tags a registry dial (#105104)
- retainGatewayForAgent on the plain-profile route (null connection) still
dials foreground (#110354 branch, gatewayForProfile path)
These were skipUnless(lark-oapi), and lark-oapi is in no CI lane, so they ran
nowhere. They are mis-gated, not dead: the adapter logic around the SDK is
ours. Restored ungated with a fake of the lazily imported lark-oapi request
builders (the external boundary) and of _tenant_get_request, so they run on
every lane. Sabotage-checked against the adapter (red when hydration skips
on env ids, and when a failed Typing removal still stacks CrossMark).
- test_hydrated_bot_identity_wins_over_stale_env_values
guards: /bot/v3/info runs even with FEISHU_BOT_* set and the hydrated identity wins (#16993)
- test_bot_sender_name_is_fetched_via_basic_batch_and_cached
guards: bot names resolve via bots/basic_batch with repeated bot_ids and are cached
- test_human_sender_name_is_fetched_from_contact_api_and_cached
guards: human names resolve via the contact API (open_id id-type) and are cached
- test_processing_success_removes_typing_and_adds_nothing
- test_processing_failure_swaps_typing_for_cross_mark
- test_processing_failure_skips_cross_mark_when_typing_removal_fails
guard: the Typing -> (nothing | CrossMark) processing-reaction lifecycle, no contradictory badges
Each is the only remaining guard for its issue; all drive production code.
- test_allowlist_warning_platform_gate.py::test_allowlist_warning_requires_an_enabled_messaging_platform
guards: the "No env user allowlists" startup reminder fires only when a messaging platform is
enabled, never for an API-server-only gateway (#115439). Now builds the runner with
object.__new__ (the check reads only self.config): 4.1s -> ~0.2s.
- test_interim_only_consumer_warning.py::{test_stream_capable_consumer_still_warns,
test_interim_only_consumer_skips_duplicate_warning}
guards: no false "possible duplicate send" warning per turn for interim-only stream consumers,
control case keeps it for stream-capable ones (#105341)
- test_kanban_notifier.py::{test_block_loop_technical_kind_uses_neutral_orchestration_wording,
test_block_loop_owner_input_keeps_decision_wording}
guards: a technical block loop ping must not claim a human decision; needs_input keeps it (#111125)
- test_restart_resume_pending.py::TestResumePendingSystemNote::test_empty_message_noninteractive_note_continues_task
guards: webhook/API-server resumes tell the model to continue the interrupted task instead of
asking "what next" (#57056)
The purge dropped both /model offload tests as mechanism pins (they spied
asyncio.to_thread). Restored as liveness invariants instead: the blocking
call records the thread it ran on and must not be the event-loop thread.
Both go red when asyncio.to_thread is made to run inline.
- test_model_command_async_offload.py::test_picker_path_runs_provider_listing_off_the_event_loop
guards: list_picker_providers (can block on HTTP) never runs on the gateway loop (#41289/#41304)
- test_model_command_custom_providers.py::test_direct_model_switch_runs_off_the_event_loop
guards: `/model <name>` switch_model() never runs on the gateway loop (#20525)
Restored (repaired to drive production code, mocks only at the provider
boundary; each verified red when its fix is reverted):
- test_auxiliary_client.py::TestCodexAuxiliaryAdapterCompletedResponse::
test_completed_response_with_null_output_does_not_crash
A Codex-compatible host returning a completed Responses object with
output=None yields an empty 'stop' turn instead of TypeError (#33368).
Repaired: the old test patched the internal stream consumer; it now goes
through the real completed-response path via the fake client.
- test_aux_progress_streaming.py::TestContentBearingProgress::
test_content_free_frames_still_record_ttfp_timing
Keepalive frames still fire the provider-response (TTFP) hook after the
#96707 progress gating (#96945/#96963).
- test_context_compressor.py::TestStreamingClosedFailure::
test_premature_stream_close_is_transient_network_failure[x2]
httpcore 'incomplete chunked read' / 'response ended prematurely' are
classified by the real classifier as transient: network-failure flag set
(session preserved) and a cooldown shorter than a generic failure's
(#18458). Repaired: no longer patches _is_connection_error, and compares
cooldowns instead of pinning the 30s constant.
Review on #120071 (yoniebans) found tests deleted as 'dead' that were only
mis-gated, and regression tests with no remaining equivalent:
- Discord voice receive (real NaCl/RTP: wrong-key drop, DAVE decrypt
failure, malformed padding, non-allowlisted speaker) moves out of the
never-selected tests/integration/ to tests/plugins/platforms/, where no
conftest stubs discord.py; 34 pass with the messaging extra.
- Five *_windows_live.py files had a bare skipif(win32), so
list_os_marked_tests.py never selected them; now windows_only.
- iron-proxy token-swap E2E (opt-in gate) is back.
- Gateway regressions: cached agent iteration cap (#48127), clarify JSON
never renders as progress (#52374), handoff watcher DB work off the event
loop, delivery ledger single connection, checkpoint prune and memory trim
housekeeping ticks.
- Desktop: command-screenshot IPC rejection, inactive WSL bridge never
spawns wsl.exe, bootstrap runner git binary.
Review follow-up. The delete shield on adoptBotSectionsFromMeta was a
session-long Set: once a section was deleted here, a member filed back into
that id later (by another desktop that still has the section) could never
make it reappear on this one, and the module-level Set leaked between tests.
- The shield now lasts only while the delete's member clear is running and
expires when moveBotsToSection(members, null) settles. After that no member
cleared here carries the id, so one that does was filed there again and
adopting it is right.
- A per-delete token keeps an earlier delete's clear (delete -> Undo ->
delete) from lifting the shield while the later delete is still clearing.
- resetBotSectionDeletes() test seam, called from the user-sections beforeEach.
One invariant test: red on the PR head (shield never expires) and red with
the token guard removed (first clear lifts it early); green with this fix.
A delete clears its members one profile write at a time. While a write is
slow (a remote gateway, a pooled local backend waiting for a slot), the
not-yet-cleared members still carry the section's id and name in ui_meta,
and adoptBotSectionsFromMeta (#114983) rebuilt the section the user had just
deleted - under whatever name the member still carried, so a renamed section
came back with its pre-rename name.
- adoptBotSectionsFromMeta skips sections deleted on this desktop this
session; Undo takes the id back out.
- moveBotsToSection stops when its target section is deleted mid-loop, so a
rename re-stamp stuck behind a slow write cannot refile members into the
deleted section afterwards.
Under rollback-journal (DELETE) mode a reader needs a SHARED lock, and
every commit from another process blocks it across its journal+db fsyncs.
Read-only SessionDB handles were opened with a 1 s busy timeout, and
writer handles serve DELETE-mode reads on the writer connection, whose
1 s timeout exists for writes (they retry at application level). Under a
busy gateway, readers gave up after 1 s:
- dashboard GET /api/sessions -> 503 "Session store is busy", or 500 when
the lock surfaced through the FTS5 vtable constructor during the
read-only open probe;
- `hermes sessions list` -> "Could not open your session history
database. Run: hermes sessions repair" (a healthy store) or a raw
`database is locked` traceback;
- in-process reads on a gateway/CLI writer handle -> `database is locked`.
Reads now get the same 5 s SQLite busy budget the WAL read pool already
used (_READ_BUSY_TIMEOUT_S): read-only handles are opened with it, and a
DELETE-mode read on the writer connection raises busy_timeout for the
read and restores it after, so writes keep their short timeout and the
jittered application-level retry.
Since 10a78f9916 every member that knows its connection stamps its reply
with `from.source` (local members too, e.g. "This device"). The round's
own-entry watermark walk (e00aa13d7d, written against the pre-stamp base)
still only expected a source for remoteSource members, so a local member's
sourced reply did not read as its own: the watermark stayed before it, the
member got its own reply back as "new messages in the room" and answered
itself again, turn after turn, until the round cap. Live, a room that should
settle after one reply kept the Stop button up and posted extra replies.
authoredByMember now uses the same authorship rule as the `(you)` suffix
(isGroupChatSelf: gateway install_id first, then connection label/id, an
unsourced line only for a local viewer), so the two can no longer drift.
Hermes's 14-day `[tool.uv] exclude-newer` quarantine applies to Hermes's own
dependencies only (uv lock/sync, `hermes update`, LAZY_DEPS extras via
`ensure()`). A plugin's declared `python_dependencies` install under the
PLUGIN's policy: `install_specs(policy="plugin")` runs uv with `--no-config`
from any cwd, still inside the core constraints file.
Reverses item 3 of #118841, which ran the uv tier with cwd=<checkout> for
every install so the quarantine reached plugin deps from any cwd. That made
catalog re-pins floored on a <14-day release uninstallable (#120076:
"only hindsight-client<=0.9.2 is available"; #114530 held on the same gate).
Maintainer ruling (Teknium): "plugins dont have to abide by our 14 day rule
btw. They can have their own security policy on that. Only hermes'
dependencies themselves have to. We should recommend that they do this for
their plugins and we should give guidance to plugin devs that they should
though."
- tools/lazy_deps.py: INSTALL_POLICIES ("core" | "plugin"); `_uv_policy_args`
replaces `_uv_policy_cwd`; `_venv_pip_install(policy=)` defaults to core
(ensure/LAZY_DEPS), `install_specs(policy=)` defaults to plugin.
- hermes_cli/plugin_python_deps.py: `resolve()` passes policy="plugin".
- Docs: developer guide "Dependency security policy" section, catalog README
admission rule 9, AGENTS.md pinning policy — plugin authors are responsible
for their deps and strongly recommended to pin upper bounds, floor on the
oldest API-compatible version and run their own release quarantine
(`uv --exclude-newer` in their CI); operators can set UV_EXCLUDE_NEWER.
- Tests: the #118841 cwd test is replaced by two invariants — a plugin install
carries `--no-config` and no checkout cwd (red on base), a core lazy install
keeps the checkout cwd and no `--no-config`.
The Ink TUI already lists queued follow-ups above the composer; a standing
/goal only surfaced as transient status text after each judged turn. A one-
line row above the live-work dock now shows it: "⊙ goal · 3/20 turns ·
<title>", "⏳ goal parked · until 14:32 · <reason>" or "⏸ goal paused ·
<reason>", and leaves once the goal is done or cleared.
State comes from the existing session.control contract: one
session.control.read when the session changes, then the
session.control.update events the backend already publishes after /goal
commands and judged turns. No new polling of state.db.
The classic CLI dock above the status bar painted subagents and background
processes only. A standing /goal (and whether it is active, parked or
paused) was visible only as a "⊙ goal N/M" status-bar segment, and prompts
waiting in /queue could not be seen at all while the agent worked.
The dock now paints the goal's status line on its top row and a Queue block
(count, first three previews, +N more) on its bottom rows, closest to the
input they will be sent from; F7's summary line carries "goal active|parked|
paused" and "N queued". Goal state comes only from the GoalManager the CLI
already holds for the current session, so the 1 Hz dock refresh never opens
state.db and never shows a previous session's goal after /new.
/queue now dispatches inline while the agent runs, honouring its
busy_policy="dispatch": typed mid-run it was queued as the raw command, so
"/queue X" only re-enqueued X after the turn and /queue list|rm|edit could
not inspect the queue until it had drained.
gateway.run runs _ensure_ssl_certs() at import time and writes SSL_CERT_FILE
into os.environ. The macOS lane runs all macos_only files in one process, and
tests/gateway/* is collected first, so by the time test_actual_provider runs the
variable has leaked. ActualProfile deliberately yields to an explicit CA env
var, so both certifi-default tests saw {} / ca_bundle=None and failed.
Clear the four CA env vars the profile reads, via monkeypatch, in both tests.
The same fix covers dev Macs where SSL_CERT_FILE is exported in the shell.
The "display index not backfilled" probe was spelled twice in
hermes_state_messages.py, and the copy in the delete fence tested only
display_order while _ensure_display_order tests display_order OR
display_identity. Hoist one _DISPLAY_INDEX_MISSING_SQL beside
_DISPLAY_ACTIVE_CLAUSE and use it at both sites. The fence now refuses
whenever the read path would have backfilled instead of projecting; on
every reachable delete the two probes agree (the read path backfills both
halves before the export snapshot exists), so this only tightens fail-closed.
hermes_state_timeline keeps its own probe: it carries a role slot.
_cmd_export re-derived "which formats are human-read transcripts" as
`format == "html" or only`; use SAVE_TRANSCRIPT_FORMATS instead. Equivalent
on every reachable path: md/qmd without --only never reach _collect_sessions
(they route to _export_markdown), and --only forces the transcript view.
Bind include_compacted once in a local `_one` instead of threading it.
The IOERR test now injects on "WITH page AS", the display CTE's own opener,
so reshaping the SQL cannot let _ensure_display_order's SELECT probe absorb
the failure and make the case pass vacuously. Comments that narrated
removed guards now state the current reason.
The get_messages(include_compacted=True) extraction into
_display_rows_from_conn swapped `self._read_all(sql, params)` for a bare
`with self._read_ctx() as conn:`. _read_all routes through
_read_retrying_ioerr, which replays the SELECT on the same pooled mode=ro
reader across the WAL-transition `disk I/O error` window (#100871);
_read_ctx has no retry, so the TUI transcript load and every transcript
export would have surfaced a hard OperationalError where base recovered.
Route the display projection through _read_retrying_ioerr again. The
existing #100871 test file gains a get_messages case parametrised over
both read paths; the display CTE starts with WITH, so the flaky reader
gets a configurable statement prefix to target it.
Keep the behaviour contracts: a compacted session's markdown holds every
shown turn and the session is deleted (single + logical lineage), and a
same-count content rewrite after export is refused by the in-transaction
fence. Drop the boundary-append test (same fence, same refusal message) and
the html/--only/jsonl parametrization (sibling call sites of the same flag).
The /save test stays because it is the only coverage of the CLI and gateway
/save surfaces, trimmed to the md case per surface.
--delete-after-verified had three guards after the merge: #119935's pre-delete
message re-count, #120065's caller-side `previous != snapshot` compare between
exported items, and #120065's expected_display_messages compare inside
delete_session's BEGIN IMMEDIATE. Only the last one is race-free: it compares
the exact display rows against the store at DELETE time, so a same-count
rewrite, an append after the re-count, or inter-item drift is refused there.
The two caller-side checks run outside the transaction and can only ever
duplicate a verdict the fence already gives, so drop them and merge the
snapshots straight into expected_messages.
Verified with the S1 probe (--race): compacted_fires false, same_count_fires
false with the caller-side guards removed.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
In-place compaction is the default. It soft-archives every earlier row of
a session under the same id (active = 0, compacted = 1), and Desktop and
the dashboard still show those turns. The md/qmd export read the session
through export_session -> get_messages with the default live-only clause,
so it wrote only the compaction summary and the carried tail.
verify_export_file then compared the file with that same dict, and
delete_session removed every row of the session, including the archived
turns that never reached the file.
The md/qmd export now reads the display history (include_compacted), for
a single session and for --lineage logical. export_session and
export_session_lineage take include_compacted, off by default:
import_sessions inserts every message as live context, so the JSON export
and stranded-session adoption keep reading live rows only.
The other transcripts people read had the same hole without the delete:
`/save md` and `/save html` (CLI and gateway), `sessions export --format
html` (one session or all of them) and `--only user-prompts` each held 1
of 6 answers on a six-turn session after one compaction. They now read
the display history too (SAVE_TRANSCRIPT_FORMATS; export_all gains
include_compacted and reads per session then, since the display read
dedupes per session). `/save json`, JSONL and the dashboard's JSON export
stay live-only for the import reason above.
Before deleting, the verify step also re-counts the store's display rows
for every session the file covers and refuses on a mismatch. A message
that lands while the files are written, or a later export change that
reads a narrower view, now refuses the delete instead of being removed
unseen. Like the adoption retire loop, the re-count runs just before
delete_session, not inside its transaction. Rewind rows (undone turns,
the superseded originals of a carried tail) are still deleted without
being exported, as `hermes sessions delete` does: they are not part of
the history the session shows.
Measured through the real CLI on a session with 6 turns and one default
in-place compaction (15 rows, 13 shown): before, 3 messages were exported
and all 15 rows deleted, with answers 1-5 missing from the file; after,
13 messages are exported in display order, then deleted.
(cherry picked from commit adeaff1e33f1ae2b8a066ef2374bb29ade50b3b8)
Use a per-test byte sequence instead of repeating one random byte across the token. Keep the sequence shared across simulated tabs so duplicate-tab isolation does not depend on a random collision, while preserving the reload invariant.
The stamp cache made a miss (openai, groq, mistral, ... have no provider
profile) call _home_layer twice, re-running get_hermes_home() and
hermes_home_key() on the forced re-check. It now resolves the layer once
and re-checks stamps on it, and skips the re-check when the periodic check
already ran in the same call. A miss costs the same as on main (one stat
pair) while hits stay on the 1 s cadence; the install-then-lookup contract
test still passes.
get_provider_profile('openai'), 20k calls: 201 us before this commit,
152 us after (main 144 us); 'nous' hit 46 us (main 128 us).
_disk_serve_tier owns the ollama clamp now, and get_model_capabilities /
get_model_info already treat config=None as load-it-yourself.
Six config=config pass-through ternaries in agent/models_dev.py whose callees
have no test double (_configured_catalog_provider, _models_dev_id,
_get_provider_models) collapse to a single call; the _cfg_get and
_load_model_overrides ternaries stay because tests replace those functions.
custom:<route> names always miss the exact registry lookup before resolving
to the generic custom profile, so every one forced a plugin-dir stamp re-check
and the 1 s stamp cache never applied to them. The picker resolves them once
per model. They now wait for the periodic check (a plugin registering a named
route is still picked up within the stamp TTL); other misses keep the
immediate re-check. 50 lookups of custom:lab: 50 -> 0 stamp reads,
52.3 -> 9.8 us per call.
Keep the stale-serve gate proof (TTL-expired but servable rows are not
prefetched; rows past _PROVIDER_MODELS_STALE_SERVE_MAX are) and drop the
five ollama/credential permutations that re-test _cache_entry_valid,
which has its own coverage.
The parallel prefetch classified any cache entry past
_PROVIDER_MODELS_CACHE_TTL as needing a fetch and blocked the picker on
the thread pool until every one returned. But cached_provider_model_ids()
has two non-blocking tiers, not one: past the TTL and inside
_PROVIDER_MODELS_STALE_SERVE_MAX it still returns the cached list
immediately and revalidates off-thread. Prefetching those slugs traded a
non-blocking serial call for a blocking parallel one.
Since _PROVIDER_MODELS_STALE_SERVE_MAX is far longer than
_PROVIDER_MODELS_CACHE_TTL, every picker open more than a TTL after the
previous one paid for it: locally, 26 providers with a cache 60s past TTL
took 3.1-4.3s, bounded only by the slowest provider's round-trip. Gating
on usability instead of freshness brings that to 0.45s.
The gate now keys on the row the serial call actually reads via
_normalized_cache_slug: a bare "ollama" stays its own cache key rather
than folding into "custom", because the local native catalog has its own
TTL and its own empty-is-authoritative rule. And an empty catalog counts
as servable only for ollama inside _OLLAMA_LOCAL_MODELS_CACHE_TTL —
cached_provider_model_ids returns it with no round-trip, so prefetching
it is redundant work the picker waits on. Past that TTL an empty row gets
no stale-serve window and really does block, so it stays in the prefetch.
Entries the serial path genuinely cannot serve — missing, fingerprint
mismatched, or past the stale-serve window — still prefetch in parallel,
so a cold cache is unaffected. refresh=True already skips the prefetch
entirely, so explicit refresh still forces every provider.
python -m pytest tests/hermes_cli/test_model_cache_parallel_prefetch.py -q
19 passed
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit fb8954a37816002bc5db9a592ebc2d8f1af87673)