The Sep-2026 facade decomposition moved _detect_image_mime_type_from_bytes and
_normalize_to_supported_image into tools/vision_tools_image_prep.py; import from
the defining module per repo convention (compat pointers are off limits in-tree).
Self-review finding (hostile redteam pass), and it defeats the exact guarantee
the previous commit advertised.
The brand scan is bounded by the declared ftyp box size so a token in a FOLLOWING
box cannot upgrade a HEIC to AVIF. But the bound fell back to the whole 64-byte
sniff window whenever the declared size was outside `16 <= size <= len(header)`.
That failed OPEN in precisely the attacker-controlled case:
declared size 0 -> image/avif (genuine HEIC, 'avif' in a later box)
declared size 4 -> image/avif
declared size 8 -> image/avif
declared size 12 -> image/avif
declared size 15 -> image/avif
declared size 24 -> image/heic (truthful size, correct)
A size too small to hold any compatible brand now scans none, rather than
scanning everything. A size >= 16 that overruns the sniffed window is a truncated
read rather than an attack, so that case clamps to the bytes actually available
and a genuine oversized-ftyp AVIF is still detected.
The prior test only exercised the truthful size, so it proved the happy path
rather than the invariant it was named for. Added parameterized malformed-size
cases (0/1/4/8/12/15), a misaligned-size case (17-23, non-multiples of 4), and an
honest-oversized case. All six malformed-size assertions fail if the fail-open
bound is restored; verified by mutation, not assumed.
Two defects in the HEIF/AVIF sniffing added by this PR.
1. Major-brand-only detection mislabeled AVIF as HEIC.
AVIF encoders routinely stamp the generic still-image brand 'mif1' as the
MAJOR brand and declare the AV1 codec only in the compatible-brand list, so
every mif1-major file was reported as HEVC-coded HEIC. Parse the
compatible-brand list (bounded by the declared ftyp box size, so an 'avif'
token in a following box can't upgrade a real HEIC) and let AV1 brands win
when both families appear.
2. AVIF was rejected outright when pillow-heif was missing.
Pillow >= 11.3 bundles a native AvifImagePlugin, while pillow-heif wheels
are commonly built with no AV1 codec at all (libheif_info() reports
AVIF: ''). Gating AVIF on a pillow-heif import therefore refused files
Pillow could already decode, and pointed the user at a library that cannot
decode them. Registration is now best-effort: the decode attempt is the
arbiter, and failures emit codec-specific guidance naming the backend that
actually serves that format.
Tests: mif1+avif and mif1+av01 regressions, a box-size bound guard, a
truncated/bogus-size guard, AVIF-converts-without-pillow-heif, and AVIF error
text. Verified against real pillow-heif-encoded HEIC and Pillow-encoded AVIF
files, not just synthetic headers; both new brand tests fail if the
compatible-brand scan is reverted.
AGENTS.md:559-576 requires every PyPI dependency to carry an upper bound
rather than an exact runtime pin. Replace "pillow-heif==1.4.0" with
"pillow-heif>=1.4.0,<2" and regenerate uv.lock (resolves 1.5.0, with
hashes), confirming the range admits later 1.x releases.
vision_analyze rejected iPhone photos with 'source is not a recognized
image'. iPhones capture HEIC (ISO-BMFF/HEVC) and upload pipelines often
mislabel it .jpg; the magic-byte sniff had no HEIF branch, so the bytes
were rejected before reaching normalization.
- _detect_image_mime_type_from_bytes: sniff the ISO-BMFF 'ftyp' box and
recognize HEIF/HEIC (heic/heix/heim/heis/hevc/hevx/mif1/msf1) and AVIF
(avif/avis) brands, returning image/heic or image/avif. Non-image ftyp
brands (mp4, ...) stay unrecognized.
- _normalize_to_supported_image: register the pillow-heif Pillow opener,
then re-encode HEIF/AVIF to PNG before embed (same soft-dependency
pattern as SVG rasterization). Missing decoder yields an actionable
'pip install pillow-heif' error, not a generic failure.
- Add pillow-heif==1.4.0 as a core dep alongside Pillow (prebuilt wheels
bundle libheif; no system libs needed).
- Tests: brand detection (incl. AVIF + mp4-not-misdetected), end-to-end
HEIC->PNG normalization, and the decoder-missing error path.
The vision resolver's magic-byte sniff stays authoritative (no extension
trust); HEIF now reaches the existing normalize step instead of 400ing.
Twenty-three tests imported helpers from sibling test modules by dotted
path (tests.run_agent.test_run_agent, tests.cli.test_cli_init,
tests.gateway.test_42039_duplicate_user_message); those follow the moves
and renames. The run_agent autouse fixture that zeroes jittered_backoff
now covers all of tests/agent, where one pre-existing test asserts the
real backoff is positive — it opts out via a real_retry_backoff marker,
the same pattern real_concurrent_gate and real_agent_prewarm use.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Follow-ups from the #109962 review that landed after the merge:
- A hand-edited manifest (`"secondaries": null`, a record without profile/home,
`"default": []`) crashed with a traceback AFTER the flag had already been
flipped. It is now refused up front, with the same message from the dry-run
plan and the real run.
- The dry-run plan said "restore its recorded standalone gateway" for a record
with neither service nor pid, where the run did nothing; both now say so.
- `--standalone --dry-run` with no manifest exits 1 like the real path and
prints the same hint.
- One "Rollback incomplete" message; a failed default restart now tells the
user the multiplexer is still running and how to stop it.
- Drop the dead `prefixes and` guard; an empty profile name can no longer
produce a bare ":" platform prefix.
`_finite_positive_config_float` and `_config_int` were the same shape with the warn call
pasted six times and asymmetric sign checks. Collapse both onto `_liveness_knob(key, default,
cast)`: usable iff finite, >= 0 and exact for the cast; else warn and return 0; explicit 0
stays silent.
Two int-path holes closed on the way: `websocket_liveness_failure_threshold: .inf` raised
OverflowError inside `DiscordAdapter.__init__`, and `0.5` truncated to 0 and disabled the
probe silently — the bug class this PR exists to remove. `2.0` / `"2"` still resolve to 2.
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:
- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
`enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
`log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
`on_ready` and every healthy connection reads `event_silence` 300s later — a forced
reconnect every ~5 min. Live-verified against a real `commands.Bot` +
`DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
`ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
detect the "ESTAB + ACKing + zero events" incident by construction.
Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.
Two defects fixed:
1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
`_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
and ACKs included, so a legitimately quiet server is not flagged. The
health check gains an `event_silence` reason with its own bound,
`websocket_event_max_silence_seconds` (default 300s, 0 disables the
dimension alone). The stamp resets on `on_ready` so a reconnect never
inherits pre-restart silence. Trip path is unchanged: consecutive
failures -> retryable `discord_websocket_health_stale` -> the existing
reconnect watcher builds a fresh adapter.
2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
log line, permanently disabling the watchdog invisibly. Unparsable and
non-positive values now log one WARNING naming the knob and raw value.
The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
The manifest is kept after a partial rollback so it can be re-run, but the
detached-gateway branch spawned unconditionally, so every re-run doubled the
secondary (service start/restart is idempotent, Popen is not).
Tests: dedupe the fixture-local profile-name helper to module scope.
Rollback started each per-profile gateway while the still-live multiplexer's
gateway_state.json listed it as served, so `hermes -p X gateway run` refused
with exit 78 and RestartPreventExitStatus parked the unit for good — the exact
symptom #109473 reports, one step later. Clear served_profiles (and the
profile's platform entries) first; recorded [] is authoritative for the guard.
A failed secondary previously skipped the default restart with the flag already
off, leaving config saying standalone while the process kept multiplexing. The
default now restarts regardless (it is the last operation either way, so a
cgroup-kill of this process can no longer strand later steps); the manifest is
kept for a re-run only when something failed.
Tests: the fake service layer now runs the real served-by-multiplexer guard at
each secondary start, and a failed-secondary contract test; both red on the
prior stack.
An interruption during the first secondary stop/uninstall left no recovery
metadata on disk. Write the complete manifest first; the dict never changes
afterwards, so the per-secondary and post-flag rewrites were byte-identical
and are dropped.
Based on #109490 by @JoaoMarcos44.
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
The delegation imported ``plugins.model_providers.deepseek`` — the only cross-plugin
module import in the tree, and one that resolves solely through the loader's
sys.modules shim (popped again if the deepseek plugin fails to load). Look the
profile up with get_provider_profile("deepseek") instead: already imported from
``providers``, honours a user override of the profile, degrades to the base no-op
without a try/except.
Drop the ``len(m) <= len("deepseek/")`` guard — the native profile returns
({}, {}) for an empty id anyway. Bind the expected value in the parity test and
assert it is non-empty so the equality cannot pass as ({}, {}) == ({}, {}).
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.
Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.
Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
-> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After: -> _reasoning_config={"enabled": False}.
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
turn prologue now raises instead of silently falling back to per-request wall time, which
would re-create the mid-turn drift the fix removes. Only production caller
(assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
untrustworthy-stamp contract is its own test.
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):
- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
grep hit) was classified as a killed run — the read-only block vanished from the
model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
send/replay parity still held because both consumers made the same mistake. Require
the shape every executor actually produces: the marker is the LAST line of `output`,
inside a JSON envelope with a non-zero exit code (or a bare text result). Real
interruptions from tools/environments (130), managed_modal (130) and
code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
stamp (clock skew, corruption) had negative age and kept the confirmation plus its
api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.
Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:
- Prefix-only on the send path. build_api_messages now canonicalizes only
messages[:current_turn_user_idx]; rows the current turn appended (its own tool
calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
the previous iteration had already sent whenever a tool result matched the
interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
to remove, and it also made the dangling-tail transform order-dependent on when
the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
"exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
`grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
result to an orphan notice (or dropped a read-only block). That heuristic was
tolerable at resume time only; it now runs per request. Match the executors'
bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
live on the send path while replay expired it. _reset_per_turn_agent_state stamps
time.time() once at admission; the three other writes (bind identity, stage
message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
`"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
treat an unknowable age as expired; missing stamps (legacy rows) are still left
alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
_reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
the durable list untouched, live tail preserved; admission-clock freeze across
iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
(send path unpatched, whole-list canonicalization, per-request clock, fail-open,
loose heuristic).
The 1e9 → '1000M' test row read as a snapshot of a missing rung; the
formatter comment now states the cap (token/cost figures stay well under a
billion; a B suffix would collide with the bytes reading) and the row cites it.
scripts/dump_desktop_slash_registry.py --check exits 1 (no write) when the
committed apps/desktop/src/lib/desktop-slash-registry.json differs from
desktop_surface_registry(); the docstring now cites the test that actually
fails (test_desktop_slash_registry.py, not test_commands.py). The pytest
runs --check as a subprocess (green on the committed file) and calls
check() on a hand-edited copy (red).
scripts/ci/classify_changes.py: _PY_SKIP includes apps/, so an apps-only
edit of that JSON (or apps/shared/src/gateway-events.json) skipped the
Python lane and the only equality check never ran in CI. A table of
cross-language contract files now forces python:true; two classifier rows
pin it.
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.
Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
The desktop model picker moved from foldIncludes (searchFold: lower-case +
`[-_.]` → space on text AND query) to the shared fuzzyRank, which only
lower-cased. `gpt.4o`, `claude_3` and `qwen3-8` returned zero rows where
main listed gpt-4o / claude-3-opus / qwen3.8-flash, while HighlightMatches
still folded — filter and highlight disagreed. fuzzyScore now folds both
sides with a length-preserving fold, so positions still index the original
target and all three surfaces rank a separator variant identically.
Tests: three separator rows in fuzzy.test.ts and in the desktop picker
test. Sabotage (lower-case only): all six fail.
34 of the 46 `NO_DESKTOP_SURFACE` rows in desktop-slash-commands.ts were a
byte-for-byte copy of `desktop=` on the matching CommandDef in
hermes_cli/commands.py (7 more were aliases of those rows). The live
`commands.catalog` already carries that metadata; the static list was the
offline fallback and would silently drift on the next registry edit.
Now `hermes_cli/commands.py::desktop_surface_registry()` is the one author
of `/name -> desktop` (aliases included). `scripts/dump_desktop_slash_registry.py`
writes it to apps/desktop/src/lib/desktop-slash-registry.json, which the
desktop imports as its offline fallback (`registryUnavailableSpecs`). Five
names the Python registry has never heard of stay in an explicit
`TS_ONLY_NO_DESKTOP_SURFACE` with the reason WHY: `/density /details /logs
/mouse` are Ink-process-local display toggles (handled in
ui-tui/src/app/slash/commands/core.ts; advertised via `_TUI_EXTRA`), and
`/pets` is the plural typo of the desktop's own `/pet` action. `/switch`
was never a block-list row (it is a `/resume` alias) — not a finding.
Cross-language contract: tests/hermes_cli/test_desktop_slash_registry.py
asserts the committed JSON == desktop_surface_registry() and that every
alias carries its canonical value; desktop-slash-commands.test.ts asserts
every dumped row is unavailable/unsuggested offline with the dumped reason
and that the TS-only set is disjoint from the dump. Both sides fail on
drift (sabotage: flipping one `desktop=` in commands.py -> Python test
"stale"; adding `/clear` to the TS-only list -> vitest disjointness fails;
dropping `registryUnavailableSpecs()` -> 4 existing vitest cases fail).
Behavior change: none for users. `/model` (`desktop="hidden"`) keeps its
local picker spec; `hidden` is a popover flag read from the live catalog.
Sites: apps/desktop/src/lib/desktop-slash-commands.ts::NO_DESKTOP_SURFACE
(46 rows) -> hermes_cli/commands.py::desktop_surface_registry (41 rows via
the dump) + TS_ONLY_NO_DESKTOP_SURFACE (5 rows). apps/desktop/src/AGENTS.md
updated.
web/src/lib/slashExec.ts and web/src/components/SlashPopover.tsx had zero
importers since the React composer was replaced by the PTY-embedded TUI
(f49afd3122) — exactly what web/AGENTS.md forbids, now orphaned. Their
parseSlash still carried the `(.*)` newline bug and lacked the `prefill`
variant. Desktop and the TUI each hand-rolled the same slash split and the
same command.dispatch narrowing; the multi-line fix (#41323, #55510) had to
be applied to each copy separately.
Sites:
web/src/lib/slashExec.ts::executeSlash/parseSlash/parseCommandDispatch -> deleted
web/src/components/SlashPopover.tsx::SlashPopover -> deleted
apps/desktop/src/lib/chat-runtime.ts::parseSlashCommand -> apps/shared/src/slash.ts::parseSlashCommand
apps/desktop/src/lib/chat-runtime.ts::parseCommandDispatch -> apps/shared/src/slash.ts::parseCommandDispatch
apps/desktop/src/lib/chat-runtime.ts::SLASH_COMMAND_RE -> apps/shared/src/slash.ts::SLASH_COMMAND_RE
apps/desktop/src/app/types.ts::*CommandDispatchResponse (5 interfaces) -> apps/shared/src/slash.ts
ui-tui/src/domain/slash.ts::parseSlashCommand/looksLikeSlashCommand -> apps/shared/src/slash.ts
ui-tui/src/lib/rpc.ts::asCommandDispatch -> apps/shared/src/slash.ts::parseCommandDispatch
ui-tui/src/gatewayTypes.ts::CommandDispatchResponse -> apps/shared/src/slash.ts
9 desktop importers + 3 TUI importers repointed.
Behavior change: desktop `parseSlashCommand` now lower-cases the command
name like the TUI, backend `resolve_command` and `slash.exec` already do
(`/Help` resolved before via the case-insensitive backend; local desktop
action lookups were case-sensitive). TUI's parsed result no longer carries
the redundant `cmd` echo (no consumer read it).
Tests: apps/shared/src/slash.test.ts (parseSlashCommand multi-line /
newline-boundary / degenerate cases; parseCommandDispatch every variant +
malformed rejection). Sabotage: restoring `(.*)` in SLASH_PARTS_RE fails
2 tests; restored -> 7 pass. Desktop chat-runtime.test.ts and TUI
asCommandDispatch.test.ts cases moved here; slashParity.test.ts repointed.
ui-tui/src/lib/color.ts called itself "the twin of the desktop app's
src/themes/color.ts" and the two had already drifted: the desktop measured
readableOn but used a coarse 0.2x5 ensureContrast ladder and returned 0 for
unparseable luminance; the TUI had the fine 0.05x20 ladder and null-for-garbage
but a luminance>0.5 threshold readableOn. Both now import the primitives from
apps/shared/src/color.ts (`@hermes/shared/color`, also exported from the root
index); each surface keeps only what is specific to it. No palette / preset /
skin VALUE changes anywhere — only math.
Sites (path::symbol → canonical):
apps/desktop/src/themes/color.ts::hexToRgb → @hermes/shared/color::parseColor (deleted)
apps/desktop/src/themes/color.ts::rgbToHex → @hermes/shared/color::toHex (deleted)
apps/desktop/src/themes/color.ts::mix → @hermes/shared/color::mix
apps/desktop/src/themes/color.ts::relativeLuminance → @hermes/shared/color::relativeLuminance
apps/desktop/src/themes/color.ts::contrastRatio → @hermes/shared/color::contrastRatio
apps/desktop/src/themes/color.ts::readableOn → @hermes/shared/color::readableOn (desktop wrapper readableInk pins ['#161616','#ffffff'])
apps/desktop/src/themes/color.ts::ensureContrast → @hermes/shared/color::ensureContrast
ui-tui/src/lib/color.ts::{Rgb,parseColor,toHex,mix,relativeLuminance,contrastRatio,readableOn,ensureContrast,lighten,darken}
→ @hermes/shared/color (same names)
Stays desktop-only (apps/desktop/src/themes/color.ts): luminance, normalizeHex, readableInk, OKLCH set
(hexToOklch, oklchToHex, oklchToSrgb255, maxChroma, hueDelta, harmonize, mixOklab, withHue, ensureContrastOklch).
Stays TUI-only (ui-tui/src/lib/color.ts): liftForContrast, grayOf, desaturate, toHsl, fromHsl, retone,
boostSaturation, color()/ColorChain.
Importers repointed (17): apps/desktop/src/{sdk/index.ts, themes/context.tsx, themes/retint.ts,
themes/retint.test.ts, themes/skin.ts, themes/vscode.ts, themes/vscode.test.ts};
ui-tui/src/{theme.ts, sdk/index.ts, sdk/apps/weather.tsx, app/createGatewayEventHandler.ts,
components/agentsPanel.tsx, components/branding.tsx, components/loaders.tsx,
components/overlayPrimitives.tsx, lib/color.ts, lib/color.test.ts}.
Wiring: apps/shared/package.json exports './color'; apps/shared/src/index.ts re-exports;
apps/desktop/tsconfig.json paths + vite.config.ts alias for '@hermes/shared/color'
(ui-tui resolves the subpath via the workspace package exports, like './billing').
Behavior change (1): relativeLuminance / contrastRatio return null for unparseable
input on the desktop too (previously 0, which made garbage measure like pure
black). Desktop SDK export `contrastRatio` therefore widens to `number | null`.
Only ensureContrastOklch relied on the number: it now treats null as "already
passing / can't measure" and returns the input unchanged. Every other desktop
caller passes 6-digit hex.
Behavior change (2): readableOn MEASURES both candidate inks and returns the one
with the higher contrast ratio (desktop semantics; the threshold version got
mid-lightness accents wrong: white on #4f9e5e is 3.29:1 vs near-black 5.50:1).
Signature is readableOn(bg, inks = ['#000000', '#ffffff']); the desktop passes
its own pair via `readableInk` so desktop output is byte-identical. The TUI
switches from the luminance>0.5 threshold to measurement: over the 185 distinct
hexes in ui-tui/src/theme.ts (DARK/LIGHT seeds + built palettes) and
hermes_cli/skin_engine.py, 56 flip from '#ffffff' to '#000000' — all
mid-lightness accents (L 0.18–0.49, e.g. #cd7f32, #4caf50, #ef5350, #ffa726,
#4dabf7) where black measures 4.6–10.8:1 against white's 1.9–4.5:1. Note the
TUI never called readableOn directly; it only reaches ensureContrast's pole
choice (below), and ensureContrast is only reachable via the color() chain and
the theme.ts re-export (no production caller today).
Behavior change (3): ensureContrast steps 0.05 x 20 from the ORIGINAL color toward
the measured readableOn pole (TUI semantics). The desktop previously stepped
0.2 x 5 toward a threshold-chosen pole, so desktop-derived accents that needed a
lift (skin/VS Code imports whose accent fails 4.5:1 on the sidebar, and
--dt-primary-solid) may now land up to 0.15 closer to their original hue —
they stop at the first passing rung. Palette VALUES are unchanged; only
synthesized colors move.
Also: parseColor accepts #rgb shorthand where desktop hexToRgb rejected it —
strictly more permissive; the only desktop path fed raw user hex is
normalizeHex, which already expands shorthand itself.
Tests: apps/shared/src/color.test.ts (moved TUI parse/mix/contrast cases +
two invariants):
- "readableOn(%s) returns the ink with the higher measured contrast" — computes
contrastRatio for each candidate in the test and asserts the returned ink is
the max (a contract, not a hardcoded hex) over #4f9e5e (both ink pairs),
#cba6f7, #ffffff, #101014.
Sabotage: reverted readableOn to the luminance threshold → 3 red
(#4f9e5e x2, #cba6f7); restored → green.
- "ensureContrast(%s on %s) clears %s" — 5 failing pairs end ≥ min; plus
"leaves passing and unparseable colors byte-identical".
Sabotage: truncated the ladder to 3 rungs → 5 red; restored → green.
ui-tui/src/lib/color.test.ts keeps only the color() chain case.
Validation:
apps/shared: npx tsc -p . --noEmit (0) && npx vitest run → 3 files, 30 tests passed; npm run lint clean
apps/desktop: npx tsc -p . --noEmit (0); npx vitest run --project ui → 798/800 files, 7563/7572 tests;
the 9 failures (src/app/messaging/index.test.tsx x8 12s-timeouts, src/lib/markdown-blocks.test.ts
property fuzz 36s) are load-induced flakes under the full parallel run: both files pass in
isolation on this branch (16/16) and on origin/main; neither imports color math. npm run lint 0 errors
ui-tui: npm run build:ink; npx tsc -p . --noEmit (0) && npx vitest run → 168 files, 1764 tests passed; npm run lint 0 errors
git diff --check clean; no new gitignored .d.ts.
Handoff: desktop vs web preset palettes diverge for the four shared ids
(web presets carry a 3-slot palette {background, midground, foreground(alpha 0)}
+ warmGlow, not the desktop's 24-slot set, so only the comparable slots are
listed; web `foreground` is #ffffff alpha 0 on all four — a glow/overlay
slot, not text ink). Design call for Teknium; nothing changed here.
preset slot desktop web
cyberpunk background #000a00 #040608
cyberpunk accent #00ff41 (primary/ring/mid) #9bffcf (midground)
cyberpunk foreground #00ff41 #ffffff (alpha 0)
ember background #160800 #1a0a06
ember accent #d97316 (ring/midground) #ffd8b0 (midground = desktop fg/primary)
ember foreground #ffd8b0 #ffffff (alpha 0)
midnight background #08081c #0a0a1f
midnight accent #8b80e8 (ring/midground) #d4c8ff (midground)
midnight foreground #ddd6ff #ffffff (alpha 0)
mono background #0e0e0e #0e0e0e (match)
mono accent #9a9a9a (ring/midground) #eaeaea (midground = desktop fg/primary)
mono foreground #eaeaea #ffffff (alpha 0)
Desktop and web each re-implemented the same locale plumbing: the
TranslationOverride<T> partial-catalog type, isRecord (four copies across
the two apps), mergeTranslations, the RTL_LOCALES={'ar'} set with the
documentElement.lang/dir effect, and the endonym table for the language
picker (6 entries on desktop, 17 on web, overlapping and hand-synced).
The generic parts now live once in apps/shared/src/i18n.ts (exported from
the root index and the `@hermes/shared/i18n` subpath). It is generic over
the catalog type — no Translations, no `en` — so translation catalogs stay
per-app (content decision, deliberately not merged here).
Sites (path::symbol → canonical):
apps/desktop/src/i18n/define-locale.ts::TranslationOverride, isRecord,
mergeTranslations → @hermes/shared/i18n; defineLocale is a one-liner
web/src/i18n/define-locale.ts::TranslationOverride, isRecord,
mergeTranslations → @hermes/shared/i18n; defineLocale is a one-liner
apps/desktop/src/i18n/runtime.ts::isRecord → shared isRecord
apps/desktop/src/i18n/context.tsx::isRecord, RTL_LOCALES,
applyDocumentLocale → shared isRecord / applyDocumentLocale
web/src/i18n/context.tsx::RTL_LOCALES + inline lang/dir effect
→ shared applyDocumentLocale
web/src/i18n/context.tsx::LOCALE_META literal (17 names)
→ derived from shared LOCALE_ENDONYMS (same exported shape)
apps/desktop/src/i18n/languages.ts::LOCALE_OPTIONS.name (6 names)
→ LOCALE_ENDONYMS.<id>; englishName/configValue columns stay
The six desktop endonyms were byte-identical to web's before the move.
Tests: apps/shared/src/i18n.test.ts — mergeTranslations keeps untouched
sibling keys under a nested partial override and replaces functions/arrays
wholesale without mutating the base; RTL_LOCALES ⊆ keys(LOCALE_ENDONYMS);
applyDocumentLocale is a no-op without a document. The existing desktop
context.test.tsx RTL/lang assertions keep covering the effect.
Behavior change: none.
Three TS surfaces each carried their own ANSI stripper with different
coverage. The TUI's (OSC, DCS/SOS/PM/APC strings, complete and truncated
CSI, multi-byte non-CSI ESC sequences, stray ESC, C0 controls) is now the
single implementation at apps/shared/src/ansi.ts, exported from the root
index and the new `@hermes/shared/ansi` subpath (ui-tui has no DOM lib, so
it imports the subpath like it does for billing/skin).
Sites (path::symbol → canonical):
ui-tui/src/lib/text.ts::stripAnsi, sanitizeAnsiForRender, hasAnsi
→ moved to apps/shared/src/ansi.ts (text.ts now imports stripAnsi
from '@hermes/shared/ansi' for its own trail helpers)
ui-tui: 13 importers repointed from '../lib/text.js' to
'@hermes/shared/ansi' (createGatewayEventHandler.ts,
components/messageLine.tsx, 11 __tests__ files)
apps/desktop/src/lib/ansi.ts::stripAnsi (2 regexes) → deleted;
parseAnsi/ansiColorClass/hasAnsiCodes stay (styled-segment parser)
apps/desktop/src/app/session/hooks/use-prompt-actions/index.ts
→ imports stripAnsi from '@hermes/shared/ansi'
apps/desktop/src/components/assistant-ui/tool/fallback-model/index.ts
private SGR-only stripAnsi → deleted; imports the shared one
Tests: the TUI 'ANSI sanitizers' cases move from
ui-tui/src/__tests__/text.test.ts to apps/shared/src/ansi.test.ts, plus
one invariant: an OSC-8 hyperlink + DCS string + SGR + partial CSI tail
strips to exactly the visible text with no ESC/BEL left.
Behavior change: desktop chat system messages (use-prompt-actions) and
inline-diff chrome (stripInlineDiffChrome) now also lose OSC hyperlink
payloads, DCS strings, truncated CSI tails and C0 control bytes that the
weaker regexes let through. TUI behavior is unchanged.
Three hand-rolled compact-number formatters and two mirrored copies of the
reasoning-effort value set collapse into apps/shared/src/format.ts and
apps/shared/src/reasoning-effort.ts, exported from the package root and as
the subpaths `@hermes/shared/format` / `@hermes/shared/reasoning-effort`
(the TUI compiles with lib ES2023 and imports subpaths only). Surfaces keep
their own label maps and UI helpers. No re-export shims remain.
Convention for compactNumber (desktop's implementation, moved verbatim):
lowercase 'k', uppercase 'M', promotion-guarded thresholds (>= 999.5 -> k,
>= 999_950 -> M) so rounding can never print "1000k", trailing ".0"
stripped, non-finite / <= 0 -> "0".
Sites (path::symbol -> canonical):
apps/desktop/src/lib/format.ts::compactNumber -> apps/shared/src/format.ts::compactNumber (moved; file deleted)
web/src/lib/format.ts::formatTokenCount -> deleted
ui-tui/src/lib/text.ts::fmtK -> deleted (text.ts's own callers use compactNumber)
apps/desktop/src/app/agents/index.tsx -> @hermes/shared
apps/desktop/src/app/chat/sidebar/chrome.tsx -> @hermes/shared
apps/desktop/src/app/chat/sidebar/session-row.tsx -> @hermes/shared
apps/desktop/src/app/command-center/index.tsx -> @hermes/shared
apps/desktop/src/app/shell/context-usage-panel.tsx -> @hermes/shared
apps/desktop/src/app/shell/titlebar-controls.tsx -> @hermes/shared
apps/desktop/src/app/skills/index.tsx -> @hermes/shared
apps/desktop/src/app/skills/mcp-tab.tsx -> @hermes/shared
apps/desktop/src/components/ui/tab-dropdown.tsx -> @hermes/shared
apps/desktop/src/lib/statusbar.tsx -> @hermes/shared
apps/desktop/src/sdk/index.ts::compactNumber -> re-exported from @hermes/shared (plugin SDK surface unchanged)
apps/desktop/src/plugins/kanban/{board,drawer}.tsx -> unchanged (import via @hermes/plugin-sdk)
web/src/components/ModelInfoCard.tsx::formatTokenCount -> @hermes/shared::compactNumber
web/src/pages/ModelsPage.tsx::formatTokenCount -> @hermes/shared::compactNumber
ui-tui/src/components/appChrome.tsx::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/components/thinking.tsx::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/app/slash/commands/session.ts::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/__tests__/text.test.ts::fmtK suite -> apps/shared/src/format.test.ts (table incl. promotion guard)
apps/desktop/src/lib/reasoning-effort.ts::REASONING_EFFORTS/REASONING_EFFORT_VALUES/
DEFAULT_REASONING_EFFORT/ReasoningEffort/isReasoningEffort -> apps/shared/src/reasoning-effort.ts
(SHORT_LABELS, reasoningEffortLabel, isThinkingEnabled, resolveReasoningEffort stay local)
apps/desktop/src/app/settings/constants.ts -> @hermes/shared
apps/desktop/src/app/settings/model-settings.tsx -> @hermes/shared
apps/desktop/src/app/shell/model-catalog-menu.tsx -> @hermes/shared (+ local reasoningEffortLabel)
apps/desktop/src/app/shell/model-edit-submenu.tsx -> @hermes/shared (+ local UI helpers)
apps/desktop/src/app/shell/model-menu-panel.tsx -> @hermes/shared
apps/desktop/src/lib/model-status-label.ts -> @hermes/shared (+ local reasoningEffortLabel)
apps/desktop/src/sdk/index.ts -> value set re-exported from @hermes/shared; label helper stays from '@/lib/reasoning-effort'
apps/desktop/src/lib/reasoning-effort.test.ts -> value-set + isReasoningEffort cases moved to apps/shared/src/reasoning-effort.test.ts
web/src/lib/reasoning-effort.ts::EFFORT_OPTIONS -> labels mapped over shared REASONING_EFFORT_VALUES (same order: none, then 7 levels)
web/src/lib/reasoning-effort.ts::VALID_EFFORTS -> Set(REASONING_EFFORT_VALUES); normalizeEffort falls back to DEFAULT_REASONING_EFFORT
Semantics kept: web `none` is selectable; desktop `none` resolves to ''
(thinking off); desktop isReasoningEffort still trims + lowercases.
Behavior change:
- web: token counts on the Models page and ModelInfoCard now print a
lowercase 'k' and are promotion-guarded: 128_000 "128K" -> "128k",
999_999 "1000.0K" -> "1M", 1_500 "1.5K" -> "1.5k". 'M' is unchanged.
- TUI: fmtK used Intl compact notation; compactNumber differs only in
suffix case and the guard: 1_000_000 "1m" -> "1M", and billions no
longer get a 'b' suffix (1_000_000_000 "1b" -> "1000M"). Sub-million
values are identical ("999", "1k", "1.5k"). Non-positive values now
print "0" instead of "-1k".
- desktop: none (its formatter moved verbatim).
Tests: apps/shared/src/format.test.ts::"compactNumber" (table incl.
999_999 -> "1M", 999_949 -> "999.9k"; fails when the promotion guard is
removed) and apps/shared/src/reasoning-effort.test.ts::"reasoning-effort"
(no duplicate values, `none` is the only non-level, default is a member;
fails on a duplicated level or a `none`-accepting isReasoningEffort).
Three byte-identical (modulo prettier and a "keep in sync" header comment)
copies of model-search-text.ts and two of fuzzy.ts collapse into one copy
each under apps/shared/src, exported from the package root and as the
subpaths `@hermes/shared/fuzzy` / `@hermes/shared/model-search-text` (the
TUI compiles with lib ES2023 and imports subpaths, never the DOM-typed
root). The vitest suites move with the code; no re-export shims remain.
Sites (path::symbol -> canonical):
ui-tui/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank -> apps/shared/src/fuzzy.ts (moved)
web/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank -> deleted
ui-tui/src/lib/model-search-text.ts::modelSearchText -> apps/shared/src/model-search-text.ts (moved)
web/src/lib/model-search-text.ts::modelSearchText -> deleted
apps/desktop/src/lib/model-search-text.ts::modelSearchText -> deleted
ui-tui/src/lib/fuzzy.test.ts -> apps/shared/src/fuzzy.test.ts (moved)
ui-tui/src/lib/model-search-text.test.ts -> apps/shared/src/model-search-text.test.ts (moved)
ui-tui/src/components/modelPicker.tsx::fuzzyRank, modelSearchText -> @hermes/shared/fuzzy, @hermes/shared/model-search-text
web/src/components/ModelPickerDialog.tsx::fuzzyRank, modelSearchText -> @hermes/shared
web/src/lib/model-picker-filter.ts::fuzzyScoreMulti -> @hermes/shared
apps/desktop/src/components/model-picker.tsx::modelSearchText -> @hermes/shared (+ fuzzyRank, see below)
The header comment now names only the cross-language twin
(hermes_cli/model_search.py) as the thing to keep in sync.
Behavior change (desktop only): the desktop model picker used to filter
model rows with `foldIncludes` substring matching and keep the curated
order; it now ranks them with the same `fuzzyRank(models, query,
modelSearchText)` the web and TUI pickers use. What a user sees
differently while typing a query:
- subsequence queries match: "g4o" now finds "gpt-4o" (previously only
a literal substring such as "gpt-4" or "4o" matched);
- the best match floats to the top instead of rows staying in curated
order (exact > prefix > word-boundary > contiguous > scattered);
- a query that matches the provider name/slug still shows that
provider's full curated list in order, exactly as before;
- an empty query still shows the curated list verbatim.
The in-row highlight is unchanged (substring emphasis via HighlightMatches),
so a fuzzy-only hit renders without emphasis rather than mis-highlighting.
Tests: apps/desktop/src/components/model-picker.test.tsx::"orders model
rows exactly as the shared fuzzyRank does" asserts the rendered row order
equals the shared fuzzyRank order for the same inputs (fails on both the
old substring filter and a reversed ranking).
Also widen the two real-subprocess list_models tests from a 2 s to a 30 s response deadline:
spawning a Python interpreter under 16 parallel test workers occasionally exceeded 2 s and the
file flaked (TimeoutError in _request). The deadline only bounds a failure; the happy path
returns as soon as the fake server answers.
`/model <x>` onto copilot-acp validates through `models_validate._static_catalog`, which
reads `provider_model_ids` with no disk cache. After the session probe landed, every such
switch spawned `copilot --acp`, ran the handshake, and killed it (1-3 s; up to the 15 s
probe timeout when the CLI is installed but the session stalls). The GitHub-API tier that
path used before sat behind a 5-minute in-memory memo; the ACP tier now has the same memo,
and it remembers failures too so a broken CLI is not re-spawned per switch.
The probe itself moves to `CopilotACPProfile.fetch_models` — the slot that already said
"model listing is handled by the ACP subprocess" and returned None — so hermes_cli/models.py
no longer hand-builds `CopilotACPClient` kwargs that `profile.create_client` owns.
Discovery failures are logged at debug instead of swallowed.
Tests: the two picker wiring tests collapse into one parametrized contract; a new test
proves three consecutive switch validations pay one probe and a failed probe is not retried
(fails when the memo read is removed).
_render_peers_map_view called _sibling_resolutions per account, and each call
ran _all_profile_host_configs(), which parses every profile's config.yaml via
list_profiles() and re-reads its honcho.json; show() re-runs after every
workspace switch. cmd_peers_map now scans once and threads the rows through.
_seen_gateway_accounts turned a locked or corrupt state.db into the same [] that
means "no gateway traffic yet"; it now notes the error on stderr before returning
the empty list. The missing-table fallback for old schemas is unchanged.
Iterating a honcho SyncPage walks every following page, so _api_workspace_peers
pulled the whole workspace on its first call, never hit the `< 50` break, re-walked
pages 2..N and returned duplicates; the 200-peer cap that exists so a public bot's
workspace cannot stall the CLI was defeated and the `p<N>` picks pointed at the
wrong peers. It now reads `.items` per page and stops at the cap; `_api_workspaces`
reads `.items` too.
The wizard's `new_host` probe looked only at the host block, so an install that
keeps peerName/enabled/workspace at the root with no hosts.hermes block read as
fresh and Enter defaulted to pinning every account onto one peer. The probe
includes the root.
`_seen_gateway_accounts` dropped rows whose origin said `is_bot`, but
SessionSource.to_dict never serializes that field, so the filter was dead; removed
with its test row. `_sanitize_peer_id` was a copy of session_peers.sanitize_peer_id.
A confirmed workspace repoint is persisted, so the exit line no longer says
"Nothing changed" after one.
The identity step treated any host block without a mapping key as a new install and defaulted the choice to single peer. An install with enabled, workspace and peerName then had Enter write pinUserPeer: true and merge every gateway account onto the operator. cmd_setup now decides new-ness before its prompts populate the block, and only a block with none of the mapping, peerName, workspace or enabled keys defaults to single.
The docstring said grouping session rows by (source, user_id) enumerates every account the gateway handled. record_gateway_session_peer overwrites a row's user_id, so a shared thread keeps only its last author. The docstring now says that, and a test pins it.
clone_honcho_for_profile copied sessionPeerPrefix but not sessionAiPeerPrefix. A new profile cloned from a default block with the AI prefix on fell back to unprefixed session names and collided with the default profile's gateway sessions.
peers map showed sanitize(prefix + user_id) for prefixed accounts. The runtime appends a sha256 suffix when sanitizing changed the id or it collides with an explicit peer, so the preview named a peer the gateway never writes to. The preview now builds a HonchoSessionManager with no client and asks it.
peers map can repoint a host block's workspace, but the gateway agent cache signature did not include it. A live gateway kept writing to the old workspace until an unrelated eviction or restart. The signature now carries honcho.workspace.
Clearing the last alias in peers map popped userPeerAliases from the host block, and the host then inherited the root aliases again. The host block now keeps an empty map, which both config readers treat as an explicit override. The summary says when that empty map hides root aliases.
The seven _preview_peer_resolution tests become one parametrized test.
The seven _seen_gateway_accounts tests become two: one database whose
rows exercise grouping, ordering, bot and null-user filtering, legacy
rows without session_key, and both label sources; one for a missing
database or table.
The cmd_peers_map runner is a module function. It derives the profile
list from the config's host blocks, so the multi-profile tests no longer
build one by hand. Tests that share an assertion are parametrized: alias
written to the host block, dash clears an alias, nothing changed writes
nothing, and messages printed in the view. test_raw_runtime_id_entry
was a subset of the offline test and is folded into it. The two
_classify_workspace_peers tests merge into one assertion over the full
label map. The three setup-wizard tests lose their answer comments and
long docstrings; the two default-choice tests are parametrized.