Commit Graph

41113 Commits

Author SHA1 Message Date
teknium1
315424d027 chore: map hermes@dasg.ltd to dasgltd for attribution 2026-09-23 16:25:21 -07:00
teknium1
9612a1a249 test(gateway/stream): pin the stream-duplicate invariants to two tests
- test_stream_consumer_oversized_leftover.py: keep the user-visible invariant
  (no line of the answer reaches the channel twice as a new message) and drop
  its lane-size twin, which pins the same re-split.
- test_stream_consumer.py: a failed first send leaves no uneditable partial
  preview; the complete reply is the only message that reaches the chat.

Both are red on origin/main and green with the fixes.
2026-09-23 16:25:21 -07:00
teknium1
80aec2223e fix(gateway/stream): no uneditable preview after a failed first send
When the stream consumer's first send failed (e.g. a Telegram timeout that
never reached the platform), _first_send disabled edits but left the
message id unset, so the next tick sent another first send: a partial
preview that could never be edited. The final reply then went out as a new
message and the truncated preview stayed on screen next to it (on every
first-send timeout, not only under load). With edits disabled, skip
non-final first sends; the final is delivered once, complete.

Found by the C12 exactly-once suite (fault_matrix[stream_timeout_first_send-
fk_tg|fk_dc]).
2026-09-23 16:25:21 -07:00
teknium1
4f24811617 fix(gateway/stream): no stray "(n/n)" in the live overflow preview
Two follow-ups to the re-split after _seal_overflow_heads (previous commit),
found by the exactly-once E2E suite driving the real GatewayRunner on
platforms whose limit a streamed reply overflows (Discord 2000):

1. _split_first_send kept the adapter's " (n/n)" chunk indicator on the
   tail chunk it reuses as the live preview, so later deltas were appended
   after it: the user saw "...wo (3/3)rd0575..." embedded mid-reply and the
   visible text no longer matched the transcript. Strip the indicator from
   the kept tail.

2. The re-split after a seal now `continue`s like the gate above it when the
   turn is not finished: the tail is still unsent, and falling through into
   the segment-break reset would clear it.

tests/gateway/test_stream_consumer.py::...fence_aware_split asserted the old
behaviour (the tail starts with the full indicator-suffixed chunk); it now
asserts the indicator is dropped from the tail and never appears in an edit.
2026-09-23 16:25:21 -07:00
Craudinho
2f5f1a2c03 fix(stream): a seal must not leave an oversized leftover on the non-final lane
_seal_overflow_heads clears the edit target mid-iteration, so the overflow gate
evaluated at the top of run() is already stale when the leftover is pushed. The
seal loop exits after ONE successful seal (its condition includes
_message_id is not None), so a buffer far past the limit leaves most of itself
behind. That leftover then reaches _push_update -> _first_send with no message to
edit, on a NON-final tick, and adapter.send() chunks and caps it itself, numbering
the pieces (i/n). The turn-final lane later publishes the same text again with its
own denominator, because the two lanes use different budgets and different payload
shapes.

Observed in production on Discord: one inbound message, one API call, no tool
turns, a 36642-char answer, and two interleaved sequences on screen - (i/10) and
(i/9) - whose first chunks were byte-identical once the indicator was stripped and
whose concatenations shared a prefix. The gateway's normal final send was
correctly suppressed (streamed=True, content_delivered=True), which places the
duplicate inside the consumer rather than in the gateway's delivery ledger.

Fix: re-check the overflow gate AFTER the seal and hand a still-oversized leftover
to _split_first_send, which is the consumer's own cap-aware, ledger-aware splitter
and already owns sealing. This removes the same-iteration invalidation rather than
compensating for it downstream, and it does not add a fourth delivery flag: the
existing flags behaved correctly here.

Negative control on this branch: with the change reverted, both new tests fail,
including the one asserting that no line is published twice.
2026-09-23 16:25:21 -07:00
teknium1
21999f0bb3 test(e2e): cron soak fails on a broken fire-claim heartbeat; #119970 xfail scoped to its deviation
ev_long_run used to flag a missing claim refresh as stalled_heartbeat
and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop
left all five scenarios green. Every virtual 30 s of the 15-minute hold
now (a) waits for the run's heartbeat to restamp the claim and fails if
it never does, and (b) has a contender call the real claim_job_for_fire
and asserts it loses, also once the first stamp is older than the TTL.
With the heartbeat replaced by True all 5 scenarios are red on the
refresh assertion; on main they are green.

The #119970 cell no longer carries a whole-scenario strict xfail. While
its behavioural probe reproduces, only the oracle's fired-set and
next_run_at checks may deviate: each deviation is recorded, the model
follows the stored slot, and the soak runs every virtual day with all
other checks live, XFAILing at the end only if a deviation was seen
(and failing if the probe says open but none was). The #120314 entry
is gone (merged). A scenario that fails mid-hold now releases its held
run before teardown so it cannot bleed into the next scenario.
2026-09-23 16:21:19 -07:00
teknium1
1a6b0825aa test(e2e): drop probes for merged delivery fixes; run-time xfails for gaps with no fix PR
#120314, #120377, #120444 and #120450 are on main: their PROBES entries,
every expect_gap naming them and the gap_open(120444) audit branch go, so
those cells are plain tests again (two of those probes read source text,
which the suite must not do). Cell 5's README row no longer claims a
strict xfail.

The two static strict xfails with no probe and no fix PR
(STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time
xfails via _pending_fixes.known_failure: only the final assertions run
under it, after every wait (restart, catch-up, settle) has succeeded,
and only an AssertionError matching the gap's own signature XFAILs; any
other failure stays red and a fixed tree simply passes. The whole-run
audit now skips only tokens whose cell actually XFAILed this run.
2026-09-23 16:21:19 -07:00
teknium1
7498546291 test(e2e): two replicas contend for every cron fire claim
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
2026-09-23 16:21:19 -07:00
teknium1
6cc967a0fd test(e2e): merge-order-safe live-gap xfails and deterministic delivery cells
A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
2026-09-23 16:21:19 -07:00
teknium1
1fbf297459 ci(e2e): name the delivery suites in the 900 s per-file budget
The e2e job already discovers tests/e2e/core/delivery/ (C12 messaging
exactly-once, C13 cron virtual-clock soak) through
`run_tests.sh --include-integration tests/e2e`; both need the 900 s
per-file budget under load (C12 150-590 s on a loaded 20-core box).
Cell 5 lives under tests/conformance/ and runs in the unit job.
2026-09-23 16:21:19 -07:00
teknium1
b477dd96c3 test: messaging exactly-once through the real gateway (C12)
A child-process GatewayRunner (own HERMES_HOME, SIGKILL + restart) runs the
real agent against the scripted fake LLM provider and three instances of a
FakePlatformAdapter that implements the gateway/platforms/base.py contract
(4096/2000 limits, edits vs no edits, threads, streaming on/off). Only the
transport is fake; its fsynced op journal is ground truth. Faults per op:
timeout, ack lost, 429 retry_after, connection reset, too long, and park
points for SIGKILL before/after the platform applies an op or between
provider completion and the delivery ledger.

Oracle per inbound: exactly one complete unmarked visible reply (extra
copies only with the ledger's duplicate marker), visible text == persisted
assistant text, user row persisted once, model turn run once, no stray or
stale partial reply; plus a whole-run audit across chats. Scenarios: 13
faults x 3 platforms, /queue and interrupt follow-ups on a slow turn,
parallel threads, re-delivered inbound ids (during/after a turn, after a
runner reconnect), five crash points with restart catch-up, and an unclean
restart with nothing in flight.

Red-proven against the reverted stream fixes (#120315), a hand-revert
of c961e5bb69 (#91653), retrying timed-out sends, boot redelivery without
the marker, double-sent finals, a dropped queued follow-up and disabled
inbound dedup. Five strict xfails pin four live gaps (streaming
ack-lost duplicate, dedup lost on reconnect, streaming crash after accept,
and the unclean-restart recency fallback that re-answers every session
active in the last 120 s). Eight cells stay red until #120315 lands (every
streamed reply over the limit, and the first-send timeout) and carry xfail 'fixed by #120315':
strict where the bug fires every run, non-strict where it depends on the
stream/edit tick. The Director counts a resume note merged into the original
tagged user message as a resume, not a second run of that inbound.
2026-09-23 16:21:19 -07:00
teknium1
67d969d30f test: persistence conformance cell 5 - delivery-outbox effect exactly-once
Replaces the wave-2 skip stub. Drives the real delivery ledger, the real
adapter record/finalize methods and the real GatewayRunner boot claim +
redeliver halves in child processes on a temp HERMES_HOME, SIGKILLing at
recorded-unsent, attempting-unsent, sent-unacked and inside a boot's own
redelivery, plus 3 concurrent rebooters over 12 rows. Only the transport is
fake (fsynced journal = ground truth). Invariants per obligation: <= 1
unmarked copy, >= 1 copy, terminal ledger row after the first clean boot,
later reboots claim/send nothing, integrity_check ok.

FIRE (strict xfail, not fixed): sweep_recoverable claims a 'pending' row
without moving it to 'attempting' and the boot redelivery never calls
mark_attempting, so a boot killed after the platform accepted a PLAIN
redelivery leaves the row pending and the next boot resends it unmarked -
two unmarked copies. Red-proven against dropped recovery marker, dropped
claim CAS, skipped mark_delivered and re-claiming delivered rows.
2026-09-23 16:21:19 -07:00
teknium1
ac5577c09f test: virtual-clock soak of the real cron ticker loop (C13)
Runs the real InProcessCronScheduler loop (due scan, pending slots, fire
claims + heartbeat, executions ledger, run_one_job, delivery routing,
mark_job_run, manual run / trigger_job, a second process on the same
HERMES_HOME, SIGKILL + boot recovery) under a file-backed virtual clock for
~30 virtual days per scenario across DST transitions and TZ configs (unset,
foreign process TZ, Asia/Shanghai +08:00, America/New_York), outages,
restarts mid-run, manual runs between ticks, runs longer than the fire-claim
TTL and two contending processes. Only run_job and the platform send are
faked.

Oracle, independent of cron/jobs.py: croniter over wall time in the job's
zone. After every tick the fired-occurrence set equals the model (no early,
late, duplicate or skipped fire beyond the catch-up policy), every stored
next_run_at equals the model and is in the future; at the end every
execution is terminal and truthful (delivered <=> the fake platform got
exactly that execution's message once). Red-proven against re-injected
#105690 manual-run re-stamp, the fire-claim heartbeat deadlock shape, UTC
next_run, skipped boot recovery, missing cross-process exclusion and the
DST fold due-compare bug (#120314; newyork_on_shanghai_fall carries a strict
xfail 'fixed by #120314' until it lands). #119969 is pinned as a strict xfail.
2026-09-23 16:21:19 -07:00
hermes-seaeye[bot]
9822170b19 fmt(js): npm run fix on merge (#120737)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 23:07:57 +00:00
brooklyn!
a88121ac8c fix(desktop): keep the custom-model action out of the model id namespace
Radix Select items for models carry a `model:` prefix, so a slug like
`__custom__` selects the model instead of entering typing mode.
2026-09-23 18:02:19 -05:00
brooklyn!
e1295e33d9 fix(desktop): offer a custom model only when the query matches nothing
A trailing custom-model section under live catalog matches was noise, and
made Enter pick a slug the user was still narrowing toward. Offer it once
the query matches no row, or when asked via the new "Add custom model…"
button the Switch model dialog now shares with the composer menu.
2026-09-23 18:02:19 -05:00
brooklyn!
8a67351872 feat(desktop): custom model entry in the Settings model selects
Every model Select (main, auxiliary, MoA slots, fallbacks) gets a
"Custom model…" row that swaps the control for an inline text field;
Enter or blur remembers the id. withActive moves next to the new
ModelSelect.
2026-09-23 18:02:19 -05:00
brooklyn!
7fe3318ed5 feat(desktop): add a custom model from the composer menu and pickers
Typing an id no provider lists offers one row per configured provider
(current first); Enter switches to it and remembers it. The composer
menu also gets an explicit "Add custom model…" footer row that turns
the search box into slug entry, and Edit models can add or remove
custom rows. Only mount the catalog list when it has rows so sections
below it no longer sit under two separators.
2026-09-23 18:02:19 -05:00
brooklyn!
d24d5556ed feat(desktop): remember typed model slugs as custom models
The provider catalog is a hint list; a slug it lacks (a newer release, a
custom endpoint) is still a model to the backend. Store typed ids per
provider in localStorage and merge them into the catalog rows so every
picker offers them as normal entries. Strings for all locales.
2026-09-23 18:02:19 -05:00
teknium1
f6b17351a4 test(e2e): surface_switch no longer a known break; the tools[] pin fix is on this branch 2026-09-23 15:43:51 -07:00
teknium1
d956f0ae57 fix: key the tools[] pin by code version; never re-add config-excluded tools
Review follow-up on the byte-identical tools[] pin.

- The pin records the code identity that built it (checkout/build sha, else
  the release version). Written by the same code, every pinned tool that is
  still available keeps its pinned bytes, including tools whose parameters
  are derived per surface (delegate_task, text_to_speech, memory, patch).
  The per-tool "parameters differ -> take current" rule replaced those bytes
  on every surface hop and rewrote the ~44KB pin each time. A pin from other
  code (`hermes update`, legacy name lists) takes the current definitions
  once and is re-pinned.
- A pinned tool this process did not build is carried forward only while
  this agent's toolset selection allows it (enabled minus disabled toolsets
  and role reservations, before check_fn). It must also pass the session
  schema gates on the merged array, so browser_exec never comes back once
  terminal is gone. Client-surface toolsets (desktop_ui, project) still
  carry across hops: no config choice removed them there.
- The rotation compaction child inherits the parent's pin in the publish
  transaction.
- `hermes sessions recover` keeps pin rows in its system_prompts sweep and
  clears dangling pin hashes, as lost-and-found now does too. Profile moves
  carry the pin like the prompt. A continuing session whose pin is missing
  or unreadable (a row swept by an older build) pins the tools it sends on
  that turn, so later hops stay stable.
2026-09-23 15:43:51 -07:00
teknium1
da70c46124 fix: a pinned tool whose parameters changed takes the current definition
Surfaces differ only in tool descriptions; a parameters change means
the handler's contract moved (hermes update), and a resumed session
must not keep offering the model the old signature. One cache miss per
contract change, none per surface hop.
2026-09-23 15:43:51 -07:00
teknium1
7a31c365c3 fix: one session sends byte-identical tools[] across TUI, oneshot and gateway hops
The session tools pin (sessions.tool_names) stored names only, so every fresh
process re-materialized the bytes from its own surface and every surface hop
of one durable session was a full prompt-cache miss:

* tool_search's deferred catalog is built per process ("Search 6 additional
  tools" in the TUI gateway vs 5 in -q);
* a pinned tool missing from the fresh build (skill_manage under the -q
  footprint) came back from the static registry schema, without its
  dynamic_schema_overrides;
* a -q --resume that rebuilt the stored prompt (model switch, cwd drift)
  persisted its own pruned array over the pin.

The pin now stores the full definitions and restore replays a pinned tool that
is still available byte-for-byte (deregistered tools drop, new ones append at
the tail, legacy name-only pins still work). A continuing session whose prompt
is rebuilt applies the pin before building it, matching the freeze policy
(tools[] only changes on /new, /reload-mcp, compaction). The array is
content-addressed in the existing system_prompts store like the prompt itself,
so identical arrays across sessions are stored once; get_session resolves it.
2026-09-23 15:43:51 -07:00
teknium1
01217c5fc2 fix(tui_gateway): withdraw an approval that ends before its settle hook attaches
_emit_approval_request sends the frame, then attaches the settle hook that
withdraws it when the queue wait ends. A client answering by approval.respond
(or another surface resolving it) inside that window ends the wait first;
register_gateway_settle returns False and the request was left in
open_requests, so every later session.resume replayed a prompt nobody was
waiting on. Honour the False: withdraw at once.
2026-09-23 15:36:00 -07:00
teknium1
3dca25b6f6 test(tui): heap-trim stub accepts the hydrate_secrets kwarg the turn scope now passes 2026-09-23 15:21:15 -07:00
teknium1
d9e10e0eb1 fix(tui): a secondary profile's key save and session model stay in that profile
One Desktop backend serves several profiles. Two session-bound paths read or
wrote the LAUNCH profile instead of the session's:

- model.save_key was not @_profile_scoped: a key saved from a secondary
  session (session_id) or for an explicit profile landed in the launch
  profile's .env, and the handler then exported it into the shared
  os.environ. It now binds the profile scope like model.options, and the
  explicit os.environ publish is gone (save_env_value already publishes to
  the bound scope, and to os.environ only for the launch profile).
  reconcile_record() already skips a non-launch home, so the profile-param
  guard around it is dropped. model.disconnect had the same gap (it removed
  the launch profile's credentials) and gets the same decorator.
- session.create's info.model and the first state.db row resolved the
  default model from the launch profile's config. _session_default_model()
  resolves it under the session's own profile scope; the same launch-model
  fallback in the lazy resume info, the fallback session info, the live
  session identity and the branch row now use it too.

Found by the two-tenant Desktop backend canary
(tests/e2e/core/tenancy/test_two_tenant_desktop_backend.py): alpha's saved
key appeared in default's .env, and alpha's session.create reported
default's model.
2026-09-23 15:21:15 -07:00
Kevin Rajan
51ff85255a fix(desktop): reload language for the active profile and connection
Adapt #113996 to the current profile boundary without a shared-context store
import cycle. Reset explicit language intent per config owner, retain GET
origin for saves, and fence stale outcomes across owner changes.

(cherry picked from commit b6ccdcf8327c516668962d480a272a424d035d5d)

Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: forztf <forztf@users.noreply.github.com>
Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-09-23 17:17:54 -05:00
brooklyn!
4faa44267e fix(desktop): refresh Kanban labels when the locale changes
Preserve #87589's translated navigation and open-board labels. Add a
plugin-owned locale subscription so late config loading and runtime
language switches also refresh palette and keybind contributions without
re-registering the board route or changing saved shortcuts.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-23 17:16:09 -05:00
brooklyn!
4f9f3c3072 fix(desktop): preserve Unicode settings descriptions
Keep Unicode letters, combining marks and numbers when suppressing
redundant descriptions. Normalize equivalent Unicode spellings without
collapsing distinct Chinese, Arabic, Cyrillic or Indic copy to an empty
comparison. Cover Echo Transcripts and existing duplicate suppression.

Follow up the missing-copy diagnosis in #87588; its catalog keys already
exist, so this fixes the remaining shared ConfigField rendering issue.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-23 17:15:58 -05:00
brooklyn!
1766f65df8 fix(desktop): localize sidebar filters without changing selections
Adapt the filter-menu translations from #85952 and #96422, retaining the
current grouping order, profile scope and advanced-mode gates. Include the
Russian filter catalog from #112983 without its unrelated date/tab work.

Co-authored-by: hdy2001 <hdy2001@gmail.com>
Co-authored-by: YUAN <sociaphobia@163.com>
Co-authored-by: daio yue <dy3426983@foxmail.com>
2026-09-23 17:15:44 -05:00
brooklyn!
1e740989eb fix(desktop): retire settled live bubbles a durable fold already carries, with partial or missing receipts (#118670, #119540, #118671) 2026-09-23 17:13:53 -05:00
brooklyn!
4e950948e4 fix(desktop): match preserved error turns by durable rowId, not renderer id (#119326) 2026-09-23 17:13:53 -05:00
brooklyn!
8228bc0a1e fix(desktop): never lose a refused or failed steer's text (#68927) 2026-09-23 17:13:53 -05:00
brooklyn!
3b5c5cc7fb fix(desktop): keep a reasoning-only stream instead of hydrating over it (#118755)
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-23 17:13:53 -05:00
teknium1
9135bd35c7 fix(aux): title calls on reasoning-mandatory routes stop sending a doomed disable
The title lane asks for thinking off (reasoning_config={"enabled": False}). A route
whose model cannot disable reasoning answers 400 "Reasoning is mandatory for this
endpoint and cannot be disabled"; the ladder then retried at the floor effort and
memoised the route, but only in memory. OpenRouter and the Nous Portal already
publish reasoning.mandatory per model in /v1/models, and Hermes mirrors that
catalog to cache/reasoning_caps.json. The aux floor never read it, so an aux-only
OpenRouter route (nothing else warms that catalog) paid the failed round-trip in
every process.

known_reasoning_floor now also floors a thinking-off aux call when the route's
catalog (memory, then disk mirror; never HTTP) flags the model mandatory, and
kicks the background warm on a cold catalog so the mirror answers every later
call and process. Optional-reasoning models keep their disable.
2026-09-23 15:13:16 -07:00
teknium1
a5c389e16d test: collecting a Bedrock test no longer pip-installs boto3 mid-run
agent/bedrock_adapter.py calls lazy_deps.ensure("provider.bedrock") at
import time. The HERMES_DISABLE_LAZY_INSTALLS kill-switch was only set by
a per-test fixture, so collecting any test module that imports the
adapter ran a real `uv pip install boto3` into the shared CI venv (the
unit job never synced the bedrock extra). test_bedrock_adapter.py raced
it: when the install had not landed yet, its botocore tests skipped and
test_call_converse_replays_thinking_botocore_accepts failed with
"No module named 'botocore'" (FLAKY on this PR's second CI run).

- tests/conftest.py sets the kill-switch at import, before collection.
- tests.yml syncs --extra bedrock with the other lazy-install extras the
  suite exercises, so the botocore tests keep running, deterministically.
- The unguarded botocore test importorskips like its siblings.
- Invariant: test_lazy_deps.py asserts the switch is set at collection
  (red on origin/main's conftest, green here).
2026-09-23 15:11:36 -07:00
teknium1
8ab10cc541 ci: run the unit job on a SQLite that uses WAL
The tests.yml pinned uv 0.9.28, which resolves `uv python install 3.11`
to CPython 3.11.14 linking SQLite 3.50.4. That SQLite has the WAL-reset
bug, so Hermes deliberately falls back to DELETE journal mode and the
unit job never exercised WAL: ~2,600 tests that reach a WAL SessionDB on
a current SQLite ran DELETE, and the 38 requires_wal tests were skipped.

Bump the pin to uv 0.12.13 (resolves 3.11.16 / SQLite 3.53.1) in both
jobs, and add a step that fails the job when the venv's SQLite is
WAL-reset vulnerable, so a later pin change cannot silently revert it.
2026-09-23 15:11:36 -07:00
brooklyn!
65ad5296ea fix(cli): make function-key workflows reliable on macOS (#120666)
* fix(tui): preserve function-key events

* fix(cli): add portable dock shortcut

* test: update dock shortcut expectations
2026-09-23 22:05:49 +00:00
teknium1
b0369c49a1 fix(config): a save refused over bad YAML says to fix the file, not "try again"
The FailedConfigRead save refusal always said "could not be read ... Try
again", including when the fallback came from a YAML parse error that no
retry fixes. The refusal now matches the cause: bad YAML gets the same
"has a formatting error ... `hermes config edit`" guidance as
require_readable_config_before_write; a PermissionError gets the
fix-permissions hint; only other read errors (EMFILE/EIO) say try again.
2026-09-23 15:05:44 -07:00
teknium1
714f646eb9 fix(config): an unreadable config.yaml no longer rebuilds its fallback on every load
A read error (EMFILE/EIO/EACCES) was never cached so the next load would retry
the file, but that also re-parsed the last-known-good backup (or rebuilt the
defaults) on every load_config() while the file stayed unreadable (~170x a
cache hit in review; 12-32 ms vs 0.2 ms per load here). The fallback is now
cached like a parse-error fallback, and a hit on a read-error fallback first
re-reads the raw file bytes (no parse): the fallback is served only while that
read still fails, so a cleared EMFILE (same file signature) reads the real file
at once.
2026-09-23 15:05:44 -07:00
teknium1
327dc2a721 refactor(config): move read-failure reporting into hermes_cli/config_read_errors.py
The parse-failure banner, the active-failure record, FailedConfigRead and the
write refusals are one topic; config.py had grown past 4,100 lines with this
PR. Pure move (no behaviour change): config.py drops to 3,999 lines, below
origin/main. Callers outside config.py now import from the new module.
2026-09-23 15:05:44 -07:00
teknium1
061da070d2 fix(config): a transient read error is no longer remembered as a corrupt config.yaml
One EMFILE/EIO on an intact file recorded the file's signature in _CONFIG_PARSE_FAILURES,
and since nothing edits the file the record never expired: get_active_config_parse_failure()
kept returning the errno text for the rest of the process, so provider auto-resolution
refused with `corrupt_config` on every turn while the file read fine again. The good file
was also copied away as a `.corrupt` backup and the banner told the user to fix a
formatting error.

A successful read now clears the record (the flag means "the current file failed the last
attempt", whatever the cause, so the refusal still holds while the file is unreadable), a
read error keeps its intact file out of the `corrupt` backups, and the warning says the
file could not be read instead of pointing at YAML.

Live: PR head — one injected EMFILE, retry loads `skin: mono`, yet
get_active_config_parse_failure() == '[Errno 24] Too many open files' and
_refuse_env_adoption_if_config_corrupt() raises AuthError(corrupt_config); with this
commit the flag is None after the retry and no `.corrupt.*` copy is written.
2026-09-23 15:05:44 -07:00
teknium1
db2f07c5cf fix(config): a transient read error no longer wipes config.yaml
Every fail-open config reader turned a read error on an intact file into
a stand-in ({} from read_raw_config and the TUI's _load_cfg_raw, defaults
or a possibly stale last-known-good from load_config). Writers then
mutated that stand-in and saved it; the writer re-reads the file, finds it
fine, and merges by deletion, so one EMFILE/EIO during a TUI config.set,
a dashboard save, a migration step or any load_config()->save_config()
caller replaced the whole file with the stand-in plus one key.

- Fallbacks from a failed read are now FailedConfigRead (a dict subclass
  carrying the error). Readers are unchanged; save_config() and
  atomic_config_write() refuse to persist one, so the error reaches the
  caller and the file stays byte-identical. The subclass rides the
  load->mutate->save round trip, so every writer is covered without
  touching each of them.
- A read error's fallback is no longer cached (nor recorded as the next
  last-known-good), so the retry reads the file again.
- PUT /api/config merges into a new dict, so it reads strictly instead.
- Gateway /verbose and /footer wrote the effective view back (fail-open
  to {} on error, ${VAR} values expanded); they now round-trip the raw
  file strictly.
2026-09-23 15:05:44 -07:00
teknium1
0698be87bc fix(config): a sparse save no longer writes an empty agent: section
_normalize_max_turns_config set config["agent"] unconditionally, so any
save of a raw document without an agent section (Desktop partial PUT,
every migration persist) appended `agent: {}` to config.yaml. Only set it
when the section exists or legacy root max_turns is being moved into it.

The v20 support-floor golden pinned that phantom: its `agent: {}` only
survived because the next migration step re-added it after the strip pass
emptied the section.
2026-09-23 15:05:44 -07:00
teknium1
9c97fb6e34 test(config): trim the explicit-null salvage to one invariant test
Keep the cross-surface case (full save, partial save, migration) that pins the
property; the helper-level parametrizations restate it. Moved to a focused
module because test_config.py is past the file-size gate.
2026-09-23 15:05:44 -07:00
teknium1
24f074cecd test(config): import RetirementIssue in the xai migration fold regression
The rebase reconciled the import block against main's trimmed list and
dropped the one name the new test constructs, so the regression died with
a NameError before reaching apply_migration.
2026-09-23 15:05:44 -07:00
Hukla
2f0cab5141 fix(config): keep the xai migration emitter on the same no-fold width 2026-09-23 15:05:44 -07:00
Hukla
7f59505685 fix(config): preserve long quoted scalars 2026-09-23 15:05:44 -07:00
Shenrui Ma
93c5856bd5 fix(config): prevent saves from dropping explicit null settings
Use a distinct stripped-node sentinel so valid null values survive configuration saves and migrations. Keep existing default comparisons and explicit-path preservation rules.

Scope-risk: narrow
Tested: 14 new cases pass; 7 fail on unmodified main. Broader macOS run: 1266 passed, 7 skipped across 61 files.
Not-tested: Full repository suite and native Linux/Windows execution.
2026-09-23 15:05:44 -07:00
brooklyn!
92d13e83d7 fix(desktop): translate stock intro bodies through locale catalogs
Keep the stock selector and English JSONL authoritative. Add zh, zh-hant and ja display copy at each personality rotation position; preserve custom personality names. Reuse core skills-hub translations for advanced settings. Reported by NealZhouPanda in #88798; related broader work #90810 by Oliver Hees and #101305 by Euterer remains independent and is not superseded.
2026-09-23 17:01:11 -05:00