Commit Graph

42417 Commits

Author SHA1 Message Date
ethernet
7dee4cc068 fix(activate): keep PATH, HOME and the temp dirs in POSIX form under MSYS
The pm env is read by a native Windows Python, which the MSYS/Cygwin
runtime hands PATH, HOME, TMPDIR, TMP and TEMP in Windows form. activate
exported them verbatim, so bash split 'C:\a;C:\b' on ':' and found no
commands: scripts/run_tests.sh died with 'env: command not found' right
after activation. Convert them back with cygpath, resolved before the
export replaces PATH.
2026-09-23 19:56:46 -04:00
ethernet
cde48023ee test(pm): accept the --test-environment arg activate passes to setup
babbec1c4c made activate always pass --test-environment[=extras] to
setup-hermes.sh. The isolated checkout's stub still required
--runtime-only alone, so every native-Windows bash activation test
failed with 'setup failed'.
2026-09-23 19:56:46 -04:00
ethernet
c0bd183eef fix(release): build stable draft notes from commits, not generate-notes
GitHub's generate-notes lists merged pull requests since the last
published release. A fork merges none, so the stable draft body was only
a "Full Changelog" link, and it lost the HERMES_BUILDS_TABLE marker that
the stable workflow renders the download tables into.

The cut now builds the body with generate_changelog() over the commits
from the published stable commit (or the seed version's tag before the
first publication) to the cut commit.
2026-09-23 19:56:46 -04:00
ethernet
bcf1e1b93c fix(e2e): accept -Phase verify-stamp in windows-e2e.ps1
dd86f27636 added the verify-stamp phase dispatch and the workflow step that calls it, but not the ValidateSet entry, so every Windows leg failed parameter binding after a successful update (fork E2E run 35933334927).
2026-09-23 19:53:10 -04:00
ethernet
97de4b4ef8 Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:

- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
  double-quoted scalar is never folded after an escaped backslash. pm-clean builds
  every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
  (ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
  GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
  read-only probe (source_git_env) now refuses promisor lazy fetches, and the
  partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
  Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
  PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
  --include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
  restore joins the darwin branch of the existing after-pack.mjs, and its test
  loads the hook from electron-builder.config.cjs and imports PlatformPackager
  from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
  win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
  pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
2026-09-23 19:51:34 -04:00
lgy1027
093db2b190 fix(desktop): restore packaged macOS locale markers
Adapt #93931 to the current packaging lifecycle with real packager paths.
Retain the non-blocking restore from 0fbb1bc337539902d1132af4a02412b52e73d6ed,
exclude Chromium gender packs, and leave Windows stamping in afterExtract.

(cherry picked from commit c37a15c2fe5665081e7c9e6bfbe93a4c2f3aadfe)

Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-09-23 18:32:26 -05:00
brooklyn!
cb7b1b32dd fix(desktop): keep English stop defaults, Unicode-safe matching, configurable barge-in (#117801)
Builds on the salvaged voice.stop_phrases matcher:

- /api/config merges DEFAULT_CONFIG, so an untouched install reports
  stop_phrases: ["stop"] instead of omitting the key. Compare against
  /api/config/defaults so that case keeps the built-in English list
  ("goodbye", "never mind", ...) and only a list the user changed replaces
  it. Parse malformed values (mapping, null) as the backend does: default.
- Normalise transcripts and phrases with NFKC and Unicode punctuation
  (\p{P}), so NFD Cyrillic from STT, «guillemets» and CJK full stops match.
- Wire the existing voice.barge_in_threshold_multiplier into the desktop
  barge-in monitor: it scales the quiet-floor multiplier and the playback
  clamp (PLAYBACK_MIN_TRIGGER_LEVEL) by its ratio to the backend default
  (3.0), so stock behaviour is unchanged and a quiet BT HFP headset can
  interrupt with a lower value. No new setting or env var.
2026-09-23 18:30:27 -05:00
aydnOktay
802a0cf6ff fix(desktop): honour voice.stop_phrases in the spoken-stop matcher (#117801)
The desktop voice loop advertised configured stop phrases in the notice but matched only a hardcoded English list, so non-English STT phrases (e.g. отбой) were submitted as prompts. Wire the matcher to the same voice.stop_phrases config the backend uses.
2026-09-23 18:30:27 -05:00
brooklyn!
529d27e5be fix(desktop): read table headers aloud instead of dropping tables silently
Read Aloud removed Markdown tables with no spoken trace (#86602). Speak the
header row in their place ("Model, Price.") and keep skipping the body rows,
so the listener hears that a table is on screen and what it compares
without cell-by-cell data. The header is the reply's own text, so it is
always in the voice's language; a localized "table omitted" notice would
follow the UI locale instead and reintroduce the mismatch when the two
differ. A table whose header cells are all empty stays silent.

Also keep the sentence punctuation after an inline MEDIA: token on both
the desktop and backend paths ("see MEDIA:/x.py. Then" no longer loses
its full stop), per review on #89367.
2026-09-23 18:30:12 -05:00
Ricardo Mendes
25f7bac318 fix(tts): keep file links and code-block placeholders out of speech
- MEDIA:/path tokens (which render as 'Open <hyphen-slug>.xlsx') are stripped
  in both the backend normalizer and desktop sanitizer; a voice cannot say
  them and ElevenLabs-style models loop on the hyphenated slug ('eeeeee').
- Desktop no longer speaks the 'code block omitted' / 'link' placeholders;
  unspeakable tokens are silence, not English words (#86602).
- Line-final colons close to periods before newlines flatten, on both sides,
  so 'here is the list:' followed by a code block is spoken as two sane
  sentences instead of hanging the voice on an open colon-pause.

Regression tests on both sides pin each behavior, including the exact
colon-plus-code-fence reproduction.

Salvaged from #89367 (79db15c1a8). The symbol/dash/euro expansion and the
bare-path "the path" placeholder were not carried: on the desktop they put
English words into non-English voices, which is the #86602 bug class.
2026-09-23 18:30:12 -05:00
brooklyn!
58694737e2 fix(time): repair surrogate-bearing locale zone names at every strftime site
On Windows (fr-FR, es-AR, de-DE reports) the zone name arrives in the ANSI
code page but is decoded under a UTF-8 LC_CTYPE (UTF-8 mode, or Piper/espeak
flipping the process locale mid-run) with surrogateescape. datetime.strftime
splices tzname() in as UTF-8, so "%Z" raised UnicodeEncodeError while the
system prompt was being built and every new/compressed conversation died.

Rework safe_strftime into a small repair: output is untouched for valid text
(the system prompt stays byte-identical), surrogateescape'd bytes decode back
through the ANSI code page ("heure d'été"), anything else degrades to U+FFFD,
and a raising "%Z" is rendered from the repaired tzname(). hermes_time no
longer imports agent.* at module load.

Route the remaining locale-name sites through it: cron quota-hold notice,
auxiliary cooldown notice, cron session titles, session_search dates,
insights, learning graph and billing renew dates. Tests use real datetimes
with a surrogate zone name instead of stubbed strftime.

Fixes #102910

Co-authored-by: Aniruddha Adak <127435065+aniruddhaadak80@users.noreply.github.com>
2026-09-23 18:29:35 -05:00
fangliquan
85cb540e6d fix(time): tolerate Windows locale strftime errors 2026-09-23 18:29:35 -05:00
brooklyn!
b3e2e9fa25 test(desktop): unmounting I18nProvider cancels a pending locale retry
The bounded startup locale retry (#96177's second half) already landed on
main; this pins its cleanup path, which had no coverage: once the provider
unmounts, the scheduled retry is cleared instead of polling /api/config for
a tree nobody renders. Adapted from the unmount case in #96213.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-23 18:28:20 -05:00
brooklyn!
b4bad6fca1 fix(desktop): gate the boot WS probe extension on child liveness, not output
Follow-up to the progress-aware probe (#96177):

- Wait while the spawned child is alive instead of "alive and wrote output
  in the last 30s". The stall this fixes is silent: the backend answers
  HTTP, then holds the GIL importing gateway platform modules and prints
  nothing until the loop recovers (the web_server heartbeat only logs
  "event loop stalled" afterwards). Output recency counted down during
  exactly that window, so a long stall still failed. A live loopback child
  whose listener accepted the TCP connect is busy, not refusing; refusals
  and auth rejections are immediate error/close events, not timeouts.
  Drops the now-unused BackendOutputTail.lastActivityAt/backendMakingProgress.
- One policy home: spawnedBackendProbeOptions() in gateway-ws-probe.ts
  (base 10s, cap 90s = the port-announcement cold-start budget); both the
  primary and pool spawn paths use it. Remote gateways, "Test remote" and
  host-backend attach keep the fixed budget.
- Without keepWaitingWhile the probe is again a single one-shot timer (no
  1s polling during the base budget). Timeout reasons name what ended the
  wait: base budget, backend stopped after the budget, or the cap.
- Tests drive every deadline with fake timers (no wall-clock races) and
  cover the production policy: a 30s stall succeeds while alive, a child
  that exits mid-wait fails at the next check, a live backend that never
  answers fails at the cap.
2026-09-23 18:28:20 -05:00
Finn763
4fc04e5a4b fix(desktop): make boot WS probe wait on backend progress instead of a fixed 10s cap
Windows cold start can hold the backend's GIL for 12-28s while gateway
platform modules import (feishu/lark, weixin, telegram C extensions),
leaving the /api/ws upgrade unanswered well past the probe's fixed 10s
connect budget. The probe then fails, the desktop declares the backend
unhealthy, spawns a second backend, and boot lands in the
'Ignoring stale Hermes backend exit (1)' cascade (#96177).

The probe now supports backend-progress-aware waiting: once the quiet
connect budget passes, a keepWaitingWhile callback is polled and the
deadline only fires when the backend stops reporting progress (hard
capped by maxConnectWaitMs so a hung backend still fails the boot).

- gateway-ws-probe.ts: keepWaitingWhile / progressCheckIntervalMs /
  maxConnectWaitMs options; behavior is unchanged when the callback is
  omitted (remote-gateway probes keep the fixed timeout).
- main.ts: both local-backend boot paths (primary window + pool) pass
  keepWaitingWhile = backendMakingProgress(child, outputTail) with a
  90s cap aligned with the existing port-announcement cold-start budget.
- backend-claim.ts: BackendOutputTail tracks lastActivityAt; new
  backendMakingProgress() signal: alive child + (no output yet, i.e.
  inside the block-buffered import window, or recent output within 30s).
- Regression tests: probe waits past the base budget while progress is
  reported, fails once progress stops, respects the hard cap, fails
  closed on a throwing callback; tail activity + progress-signal units.

Benchmark (simulated 13s import stall, production budgets):
  before: FAIL at 10.0s -> boot cascade     after: OK at 13.8s
  warm start: OK 0.96s                      after: OK 0.96s (no delta)
  dead backend: FAIL 10.0s                  after: FAIL 10.1s (same)
  hung backend (alive, never ready):        after: FAIL 90.8s (hard cap)

Closes #96177

(cherry picked from commit e8dbe9e0a2bfb6853401d55a2bbda02393560de0)
2026-09-23 18:28:20 -05:00
teknium1
8aa1e450b8 fix(anthropic): keep OAuth rotation on accepted native proxies
The anthropic.com-only refresh guard stopped token rotation on hosts the
resolver itself accepts as native Anthropic (*.claude.com, /anthropic
proxies), which already receive the Anthropic token at startup. Long
sessions there hit 401 and the recovery refresh was refused too.

Official hosts (anthropic.com, claude.com) rotate as before; any other
host rotates only when the current key is already an Anthropic
credential (sk-ant- or OAuth), so the #17829 case (custom key swapped
for ANTHROPIC_API_KEY / the OAuth token) stays blocked. Azure keeps its
static-key exclusion. Also corrects the _resolve_openrouter_runtime
docstring, which still claimed OPENAI_BASE_URL is never consulted.
2026-09-23 16:27:37 -07:00
teknium1
fd1f47d8f4 chore: map limu.mzm@bytedance.com to leemove for attribution 2026-09-23 16:27:37 -07:00
teknium1
da031c31c1 fix(runtime_provider): don't send an OPENAI_BASE_URL-bound OPENAI_API_KEY to OpenRouter
The OpenRouter key ladder falls back to OPENAI_API_KEY (legacy home of an
OpenRouter key). When OPENAI_BASE_URL binds that key to another host, a
bare `provider: custom` with no endpoint fell through to openrouter.ai and
shipped the OpenAI-proxy key there instead of failing fast. Use it for
OpenRouter only when OPENAI_BASE_URL is unset or names the same host.

Hosts compare through utils.base_url_hostname, so a scheme-less
OPENAI_BASE_URL (proxy.corp:8080/v1) still counts as binding the key.

Found by the C11 routing truth table (tests/e2e/core/tenancy).
2026-09-23 16:27:37 -07:00
teknium1
a4d8f173a9 fix(anthropic): never refresh Anthropic credentials onto a foreign host
_try_refresh_anthropic_client_credentials runs before every request and on
401 for provider 'anthropic'. A URL-bearing model alias under the built-in
'anthropic' label (#28660) points that provider at a foreign host with no
key; the pre-request refresh then swapped in ANTHROPIC_API_KEY / the OAuth
token and sent it to that host.

The third-party guard above matched "anthropic.com" as a substring of the
URL, so a host like 127.0.0.1:8080/anthropic.com still passed as native and
got the refresh. Refresh only when the endpoint's hostname is Anthropic's own
(or no endpoint is set).

Found by the C11 routing truth table (tests/e2e/core/tenancy): a /model
switch to such an alias carried the vendor key to the alias host.
2026-09-23 16:27:37 -07:00
leemove
96ff42ff66 fix(anthropic): skip credential refresh for third-party endpoints
With provider 'anthropic' pointed at a third-party Anthropic-compatible
endpoint, _try_refresh_anthropic_client_credentials only skipped Azure and
otherwise re-resolved native Anthropic credentials (ANTHROPIC_API_KEY, the
stored OAuth token) and rebuilt the client with them, so the endpoint got
a credential it was never given. Skip the refresh for any endpoint
build_anthropic_client already classifies as third-party.

Ported from run_agent.py onto agent/client_lifecycle.py, where the method
lives now; the separate Azure test is dropped (Azure is one such endpoint).

Salvaged from #17829.
2026-09-23 16:27:37 -07:00
teknium1
94c0e49307 docs: scope the ACP toolset parity wording
ACP resolves toolsets like the gateway for the same platform config,
which includes plugin toolsets. An empty platform_toolsets.acp list still
adds enabled plugin toolsets, so the docs point at agent.disabled_toolsets
for removing them.
2026-09-23 16:27:05 -07:00
teknium1
316d6fb8f7 fix(acp): admit config MCP servers by platform_toolsets.acp like the gateway
A fresh ACP agent appended mcp-<server> for every enabled config MCP server
unconditionally, so a platform_toolsets.acp allowlist of server names and the
no_mcp sentinel were ignored on ACP while the gateway honoured both.

The MCP half now comes from the same _get_platform_tools(config, "acp") call
as the base toolsets: its server names (default every enabled server, a listed
allowlist, or none for no_mcp) are keyed as mcp-<server>. Editor-provided
session/new servers are unchanged.

Docs: `hermes tools` has no ACP platform entry, so drop the claim that it
configures platform_toolsets.acp; document the MCP rules with a config example.
2026-09-23 16:27:05 -07:00
teknium1
cccea211a9 chore: map tymrabchuk@gmail.com to 5uck1ess 2026-09-23 16:27:05 -07:00
teknium1
5fb5e743b3 fix(acp): resolve ACP toolsets through the shared platform resolver
A fresh ACP agent hardcoded enabled_toolsets=["hermes-acp"], so
platform_toolsets.acp never narrowed the editor tool surface, unlike the
gateway, cron and api_server which all resolve via
hermes_cli.tools_config._get_platform_tools. Resolve the ACP base the same
way (ACP keeps appending its own mcp-<server> entries), and treat only None,
not an explicit empty list, as "use the hermes-acp default" in the /tools
and MCP-refresh rebuilds so a deny-all list cannot re-widen mid-session.

With the unconfigured default the resolved tool definitions are
byte-identical to the hermes-acp composite, so existing sessions keep the
same tool list and prompt cache.

The fresh-session assertion in test_make_agent_prefers_passed_toolsets_over_config_servers
now checks membership of the config MCP entry: the resolver returns the
expanded toolset keys rather than the bare composite name.

Refs #74582, #79516. Credit: #64045 (@israellot), #80309 (@thatssoheil),
#106834 (@nicolasramos) proposed the resolver routing.
2026-09-23 16:27:05 -07:00
teknium1
9a07debfbf test(acp): stub load_config on the real hermes_cli.config module
test_acp_real_agent_gets_session_db_for_recall replaced the whole
hermes_cli.config module with a one-attribute stub, so any real import
from it inside _make_agent (now hermes_cli.tools_config, for the shared
toolset resolver) died with ImportError. Patch the one function instead;
the assertions are unchanged.
2026-09-23 16:27:05 -07:00
teknium1
b1b017d079 test(acp): trim the salvaged disabled_toolsets tests to one invariant
Keep the behavioural /tools check (real get_tool_definitions, red on main);
drop the kwargs-forwarding and memory-provider cases, which the fresh-agent
tool-surface test in the next commit covers end to end.
2026-09-23 16:27:05 -07:00
Tym Rabchuk
d2a073f93b fix(acp): filter disabled_toolsets from /tools listing; add behavioral coverage
Review follow-up: _cmd_tools rebuilt its listing without
state.agent.disabled_toolsets, so a config-disabled toolset was filtered
from execution but still advertised by /tools. Pass it through, matching
the session tool-surface rebuild.

New tests exercise the real get_tool_definitions (no patching) and assert
a disabled toolset is absent from both the /tools listing and the rebuilt
valid_tool_names — with a baseline assertion that the toolset is present
when nothing is disabled, so the check cannot pass vacuously.
2026-09-23 16:27:05 -07:00
Tym Rabchuk
2c12e7ae84 fix(acp): pass agent.disabled_toolsets to AIAgent — config-disabled tools stayed executable over ACP
The CLI (cli.py: CLI_CONFIG['agent'].get('disabled_toolsets')) and the
gateway (gateway/run.py) both read agent.disabled_toolsets from config
and pass it to AIAgent. The ACP adapter's _make_agent never did, so
state.agent.disabled_toolsets was always None and the ACP tool-registry
rebuild (acp_adapter/server.py -> get_tool_definitions) included every
tool in the enabled toolsets — a toolset the user disabled in config
(todo, browser, ...) remained fully executable in editor/ACP sessions.

Observed live: a 'todo' tool call executed from a profile whose config
lists todo in agent.disabled_toolsets.

Read agent.disabled_toolsets in _make_agent and pass it through,
mirroring the CLI and gateway paths.

Claude-Session: https://claude.ai/code/session_01YNvCUipheR7yx4VorUL2jW
2026-09-23 16:27:05 -07:00
teknium1
6d150c7e78 fix(gateway/stream): a failed tail send at a tool boundary no longer loses the tail
After an overflow split the heads are sealed and the message id is cleared,
so the tail goes out as a first send in the tick that carries the tool
boundary. When that send failed, _end_segment skipped the unseen-tail flush
(it required a truthy _message_id) and _reset_segment_state cleared
_accumulated: ~1.1k chars vanished. The same loss existed for a plain failed
first send at a segment break.

The flush now runs whenever the boundary update did not land, except for the
__no_edit__ sentinel (its continuation still goes out once via the fallback
final). With a real message id the condition is unchanged, so no new send
there. The flush also honours the egress guard like the other new-message
fallbacks, since it can now run on the path where a draft was declined.

Tests: the oversized-leftover file now pins one invariant, parametrized over
none / commentary / segment_break / tail_send_fails: every tail line reaches
the channel, never twice as a new message, and post-boundary text is not
glued onto the pre-boundary preview. tail_send_fails is red on f361b39
(24 lines lost), green here.
2026-09-23 16:25:21 -07:00
teknium1
a09b36e82c fix(gateway/stream): a boundary drained in the overflow-split tick still applies
Review finding on the seal re-split: `continue` after `_split_first_send`
skipped the rest of the tick, so a commentary drained with the oversized
burst was dropped and a tool boundary was lost (post-tool text then
extended the pre-tool preview). The initial-overflow gate had the same
`continue`.

Seal first (a no-op without a message id), split once, then push the
tail through `_push_update` in the same tick so commentary, segment end
and the flush barrier run normally after it. `continue` is kept only when
a head send failed and the full text must stay for the fallback final.
2026-09-23 16:25:21 -07:00
teknium1
315424d027 chore: map hermes@dasg.ltd to dasgltd for attribution 2026-09-23 16:25:21 -07:00
teknium1
9612a1a249 test(gateway/stream): pin the stream-duplicate invariants to two tests
- test_stream_consumer_oversized_leftover.py: keep the user-visible invariant
  (no line of the answer reaches the channel twice as a new message) and drop
  its lane-size twin, which pins the same re-split.
- test_stream_consumer.py: a failed first send leaves no uneditable partial
  preview; the complete reply is the only message that reaches the chat.

Both are red on origin/main and green with the fixes.
2026-09-23 16:25:21 -07:00
teknium1
80aec2223e fix(gateway/stream): no uneditable preview after a failed first send
When the stream consumer's first send failed (e.g. a Telegram timeout that
never reached the platform), _first_send disabled edits but left the
message id unset, so the next tick sent another first send: a partial
preview that could never be edited. The final reply then went out as a new
message and the truncated preview stayed on screen next to it (on every
first-send timeout, not only under load). With edits disabled, skip
non-final first sends; the final is delivered once, complete.

Found by the C12 exactly-once suite (fault_matrix[stream_timeout_first_send-
fk_tg|fk_dc]).
2026-09-23 16:25:21 -07:00
teknium1
4f24811617 fix(gateway/stream): no stray "(n/n)" in the live overflow preview
Two follow-ups to the re-split after _seal_overflow_heads (previous commit),
found by the exactly-once E2E suite driving the real GatewayRunner on
platforms whose limit a streamed reply overflows (Discord 2000):

1. _split_first_send kept the adapter's " (n/n)" chunk indicator on the
   tail chunk it reuses as the live preview, so later deltas were appended
   after it: the user saw "...wo (3/3)rd0575..." embedded mid-reply and the
   visible text no longer matched the transcript. Strip the indicator from
   the kept tail.

2. The re-split after a seal now `continue`s like the gate above it when the
   turn is not finished: the tail is still unsent, and falling through into
   the segment-break reset would clear it.

tests/gateway/test_stream_consumer.py::...fence_aware_split asserted the old
behaviour (the tail starts with the full indicator-suffixed chunk); it now
asserts the indicator is dropped from the tail and never appears in an edit.
2026-09-23 16:25:21 -07:00
Craudinho
2f5f1a2c03 fix(stream): a seal must not leave an oversized leftover on the non-final lane
_seal_overflow_heads clears the edit target mid-iteration, so the overflow gate
evaluated at the top of run() is already stale when the leftover is pushed. The
seal loop exits after ONE successful seal (its condition includes
_message_id is not None), so a buffer far past the limit leaves most of itself
behind. That leftover then reaches _push_update -> _first_send with no message to
edit, on a NON-final tick, and adapter.send() chunks and caps it itself, numbering
the pieces (i/n). The turn-final lane later publishes the same text again with its
own denominator, because the two lanes use different budgets and different payload
shapes.

Observed in production on Discord: one inbound message, one API call, no tool
turns, a 36642-char answer, and two interleaved sequences on screen - (i/10) and
(i/9) - whose first chunks were byte-identical once the indicator was stripped and
whose concatenations shared a prefix. The gateway's normal final send was
correctly suppressed (streamed=True, content_delivered=True), which places the
duplicate inside the consumer rather than in the gateway's delivery ledger.

Fix: re-check the overflow gate AFTER the seal and hand a still-oversized leftover
to _split_first_send, which is the consumer's own cap-aware, ledger-aware splitter
and already owns sealing. This removes the same-iteration invalidation rather than
compensating for it downstream, and it does not add a fourth delivery flag: the
existing flags behaved correctly here.

Negative control on this branch: with the change reverted, both new tests fail,
including the one asserting that no line is published twice.
2026-09-23 16:25:21 -07:00
teknium1
21999f0bb3 test(e2e): cron soak fails on a broken fire-claim heartbeat; #119970 xfail scoped to its deviation
ev_long_run used to flag a missing claim refresh as stalled_heartbeat
and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop
left all five scenarios green. Every virtual 30 s of the 15-minute hold
now (a) waits for the run's heartbeat to restamp the claim and fails if
it never does, and (b) has a contender call the real claim_job_for_fire
and asserts it loses, also once the first stamp is older than the TTL.
With the heartbeat replaced by True all 5 scenarios are red on the
refresh assertion; on main they are green.

The #119970 cell no longer carries a whole-scenario strict xfail. While
its behavioural probe reproduces, only the oracle's fired-set and
next_run_at checks may deviate: each deviation is recorded, the model
follows the stored slot, and the soak runs every virtual day with all
other checks live, XFAILing at the end only if a deviation was seen
(and failing if the probe says open but none was). The #120314 entry
is gone (merged). A scenario that fails mid-hold now releases its held
run before teardown so it cannot bleed into the next scenario.
2026-09-23 16:21:19 -07:00
teknium1
1a6b0825aa test(e2e): drop probes for merged delivery fixes; run-time xfails for gaps with no fix PR
#120314, #120377, #120444 and #120450 are on main: their PROBES entries,
every expect_gap naming them and the gap_open(120444) audit branch go, so
those cells are plain tests again (two of those probes read source text,
which the suite must not do). Cell 5's README row no longer claims a
strict xfail.

The two static strict xfails with no probe and no fix PR
(STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time
xfails via _pending_fixes.known_failure: only the final assertions run
under it, after every wait (restart, catch-up, settle) has succeeded,
and only an AssertionError matching the gap's own signature XFAILs; any
other failure stays red and a fixed tree simply passes. The whole-run
audit now skips only tokens whose cell actually XFAILed this run.
2026-09-23 16:21:19 -07:00
teknium1
7498546291 test(e2e): two replicas contend for every cron fire claim
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
2026-09-23 16:21:19 -07:00
teknium1
6cc967a0fd test(e2e): merge-order-safe live-gap xfails and deterministic delivery cells
A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
2026-09-23 16:21:19 -07:00
teknium1
1fbf297459 ci(e2e): name the delivery suites in the 900 s per-file budget
The e2e job already discovers tests/e2e/core/delivery/ (C12 messaging
exactly-once, C13 cron virtual-clock soak) through
`run_tests.sh --include-integration tests/e2e`; both need the 900 s
per-file budget under load (C12 150-590 s on a loaded 20-core box).
Cell 5 lives under tests/conformance/ and runs in the unit job.
2026-09-23 16:21:19 -07:00
teknium1
b477dd96c3 test: messaging exactly-once through the real gateway (C12)
A child-process GatewayRunner (own HERMES_HOME, SIGKILL + restart) runs the
real agent against the scripted fake LLM provider and three instances of a
FakePlatformAdapter that implements the gateway/platforms/base.py contract
(4096/2000 limits, edits vs no edits, threads, streaming on/off). Only the
transport is fake; its fsynced op journal is ground truth. Faults per op:
timeout, ack lost, 429 retry_after, connection reset, too long, and park
points for SIGKILL before/after the platform applies an op or between
provider completion and the delivery ledger.

Oracle per inbound: exactly one complete unmarked visible reply (extra
copies only with the ledger's duplicate marker), visible text == persisted
assistant text, user row persisted once, model turn run once, no stray or
stale partial reply; plus a whole-run audit across chats. Scenarios: 13
faults x 3 platforms, /queue and interrupt follow-ups on a slow turn,
parallel threads, re-delivered inbound ids (during/after a turn, after a
runner reconnect), five crash points with restart catch-up, and an unclean
restart with nothing in flight.

Red-proven against the reverted stream fixes (#120315), a hand-revert
of c961e5bb69 (#91653), retrying timed-out sends, boot redelivery without
the marker, double-sent finals, a dropped queued follow-up and disabled
inbound dedup. Five strict xfails pin four live gaps (streaming
ack-lost duplicate, dedup lost on reconnect, streaming crash after accept,
and the unclean-restart recency fallback that re-answers every session
active in the last 120 s). Eight cells stay red until #120315 lands (every
streamed reply over the limit, and the first-send timeout) and carry xfail 'fixed by #120315':
strict where the bug fires every run, non-strict where it depends on the
stream/edit tick. The Director counts a resume note merged into the original
tagged user message as a resume, not a second run of that inbound.
2026-09-23 16:21:19 -07:00
teknium1
67d969d30f test: persistence conformance cell 5 - delivery-outbox effect exactly-once
Replaces the wave-2 skip stub. Drives the real delivery ledger, the real
adapter record/finalize methods and the real GatewayRunner boot claim +
redeliver halves in child processes on a temp HERMES_HOME, SIGKILLing at
recorded-unsent, attempting-unsent, sent-unacked and inside a boot's own
redelivery, plus 3 concurrent rebooters over 12 rows. Only the transport is
fake (fsynced journal = ground truth). Invariants per obligation: <= 1
unmarked copy, >= 1 copy, terminal ledger row after the first clean boot,
later reboots claim/send nothing, integrity_check ok.

FIRE (strict xfail, not fixed): sweep_recoverable claims a 'pending' row
without moving it to 'attempting' and the boot redelivery never calls
mark_attempting, so a boot killed after the platform accepted a PLAIN
redelivery leaves the row pending and the next boot resends it unmarked -
two unmarked copies. Red-proven against dropped recovery marker, dropped
claim CAS, skipped mark_delivered and re-claiming delivered rows.
2026-09-23 16:21:19 -07:00
teknium1
ac5577c09f test: virtual-clock soak of the real cron ticker loop (C13)
Runs the real InProcessCronScheduler loop (due scan, pending slots, fire
claims + heartbeat, executions ledger, run_one_job, delivery routing,
mark_job_run, manual run / trigger_job, a second process on the same
HERMES_HOME, SIGKILL + boot recovery) under a file-backed virtual clock for
~30 virtual days per scenario across DST transitions and TZ configs (unset,
foreign process TZ, Asia/Shanghai +08:00, America/New_York), outages,
restarts mid-run, manual runs between ticks, runs longer than the fire-claim
TTL and two contending processes. Only run_job and the platform send are
faked.

Oracle, independent of cron/jobs.py: croniter over wall time in the job's
zone. After every tick the fired-occurrence set equals the model (no early,
late, duplicate or skipped fire beyond the catch-up policy), every stored
next_run_at equals the model and is in the future; at the end every
execution is terminal and truthful (delivered <=> the fake platform got
exactly that execution's message once). Red-proven against re-injected
#105690 manual-run re-stamp, the fire-claim heartbeat deadlock shape, UTC
next_run, skipped boot recovery, missing cross-process exclusion and the
DST fold due-compare bug (#120314; newyork_on_shanghai_fall carries a strict
xfail 'fixed by #120314' until it lands). #119969 is pinned as a strict xfail.
2026-09-23 16:21:19 -07:00
ethernet
c58744e59e fix(release): link the draft at cut time, not a page that 404s until green
The stable cut already creates the draft on the claim ref before it
dispatches the gate, but the output pointed at releases/tag/<claim> and
then at releases/tag/v<version> "when it is green". GitHub serves a draft
only at an untagged-* URL, so neither link showed it. Operators read that
as "the draft appears after the build".

Take the URL gh release create prints (the release's html_url) and print
it up front, with a note that notes edited during the build survive:
edit_draft_release keeps the body and strips only the warning fence. The
v<version> URL stays, labelled as where the release lives once published.

The canary resume path had the same broken releases/tag/<tag> link; it
now uses the create output or the url field of gh release view.
2026-09-23 19:20:27 -04:00
hermes-seaeye[bot]
9822170b19 fmt(js): npm run fix on merge (#120737)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 23:07:57 +00:00
brooklyn!
a88121ac8c fix(desktop): keep the custom-model action out of the model id namespace
Radix Select items for models carry a `model:` prefix, so a slug like
`__custom__` selects the model instead of entering typing mode.
2026-09-23 18:02:19 -05:00
brooklyn!
e1295e33d9 fix(desktop): offer a custom model only when the query matches nothing
A trailing custom-model section under live catalog matches was noise, and
made Enter pick a slug the user was still narrowing toward. Offer it once
the query matches no row, or when asked via the new "Add custom model…"
button the Switch model dialog now shares with the composer menu.
2026-09-23 18:02:19 -05:00
brooklyn!
8a67351872 feat(desktop): custom model entry in the Settings model selects
Every model Select (main, auxiliary, MoA slots, fallbacks) gets a
"Custom model…" row that swaps the control for an inline text field;
Enter or blur remembers the id. withActive moves next to the new
ModelSelect.
2026-09-23 18:02:19 -05:00
brooklyn!
7fe3318ed5 feat(desktop): add a custom model from the composer menu and pickers
Typing an id no provider lists offers one row per configured provider
(current first); Enter switches to it and remembers it. The composer
menu also gets an explicit "Add custom model…" footer row that turns
the search box into slug entry, and Edit models can add or remove
custom rows. Only mount the catalog list when it has rows so sections
below it no longer sit under two separators.
2026-09-23 18:02:19 -05:00
brooklyn!
d24d5556ed feat(desktop): remember typed model slugs as custom models
The provider catalog is a hint list; a slug it lacks (a newer release, a
custom endpoint) is still a model to the backend. Store typed ids per
provider in localStorage and merge them into the catalog rows so every
picker offers them as normal entries. Strings for all locales.
2026-09-23 18:02:19 -05:00