_row() forwarded provider, provider_label and provider_profiles from the
catalog metadata but dropped provider_primary, so a provider card's own
index-0 credential (ALIBABA_CODING_PLAN_CN_API_KEY) reached the Keys tab
with provider_primary unset. buildProviderKeyGroups then fell through to
the first non-advanced key var — which, for the (Coding Plan, China) and
(Token Plan, China) cards, is a shared fallback alias (DASHSCOPE_API_KEY /
ALIBABA_TOKEN_PLAN_API_KEY) contributed by peer profiles with primary:
false. The card's Paste key wrote a foreign tier's credential and the
CN-specific var was rendered nowhere when unset.
Regression test: test_get_api_env_passes_provider_primary_through pins
the CN card's own key as provider_primary=True and the shared DASHSCOPE
alias as primary=False from that provider's profile.
The Desktop switch flipped the panel before the toolset PUT, so a rejected
write left panel on / tools off (or the reverse) with only a toast. Write the
toolset first and move the panel only to what the backend holds: on a
rejection re-read the toolsets list (a timeout may still have committed) and
follow it; if the re-read also fails, leave the panel untouched. The switch is
disabled while the write is in flight and live again for a retry.
The Plugins page's Desktop-relevant toggle was UI-only: flipping it never
told the backend, so kanban stayed active regardless. The switch now also
calls setToolsetEnabled('kanban', on, profile) — the same
PUT /api/tools/toolsets/kanban route the Toolsets tab uses — invalidates
the toolsets query cache, and confirms with a toast (en/fr/de/es).
Fixes#96969
Co-authored-by: previous contributors' direction via #101582
RunClock ticked from tasks.started_at (the task's first-ever start), so after
a review timeout + retry a healthy current run showed cumulative card age
(e.g. working · 2h for a minutes-old run).
_task_dict() now emits current_run_started_at from the run row
tasks.current_run_id points at (batched single query, same as latest
summaries), and RunClock prefers it, falling back to started_at for older
backends. Fixes#99819.
Help advertises /clear as start a new session, but desktop treated it
as a TUI-only screen clear so history and context usage never reset.
Alias it onto the existing /new action. Stop/interrupt is unchanged.
Existing directories were reported as success:true while the preview
pane opened nothing. Fail closed with an explicit error and do not
emit preview.open. HTTP(S) URLs and regular files are unchanged.
test_openssl_installs_once_and_rejects_damaged_shared_install and
test_native_build_command_preserves_failures_and_spaces each carried an
inline copy of the _powershell invocation with the pre-330ff28d44 30s
timeout. 330ff28d44 already established that a real powershell.exe child
on a cold CI runner can exceed 30s and raised the shared helper to 180s,
but these duplicated call sites kept the stale budget and timed out
intermittently on the windows-latest-32-arm-core lane.
Both now call _powershell(script, HELPER, tmp_path): identical argv and
environment construction (the inline setdefault calls were no-ops on
Windows), with the helper's documented cold-runner budget. No product
script changes.
Scope the #49645 initial-dial retry to LOCAL connections only: a remote
dial failure must stay a retryable boot failure (stage 'dialing') so the
bounded boot-retry loop re-runs the whole handshake, including a fresh
getConnection() against a possibly-rebuilt tunnel.
The local retry's first attempt reuses the WS URL minted at the boot
boundary and only re-mints on later attempts, keeping the mint count
observable for the reconnect-path contract tests.
The renderer made a single gateway.connect() call during boot, so a
freshly spawned backend that was still initializing (event loop blocked
15-30s by MCP connects and plugin discovery) made the one dial lose and
boot ended in the 'Could not connect to Hermes gateway' modal even
though the backend became healthy moments later. Retry the initial dial
with bounded attempts, re-minting the WS URL on every attempt (OAuth
tickets are single-use), and propagate reauth failures immediately.
Co-authored-by: Mani Saint-Victor, MD <drmani215@gmail.com>
Simple mode shadows fileBrowserOpen, so a flip on this row lands in the
session reveal layer and never persists. Show the same Simple-mode note the
other shadowed Appearance rows carry, so the row doesn't promise a standing
default it can't keep.
Settings > Appearance > Window & layout gets a File Browser toggle bound to the same persisted state as the titlebar toggle and Cmd+J, so the open/closed default is a visible, searchable preference.
A model switch persists display_kind=model_switch with role=user
(tui_gateway/server.py). Hydration renders that row as a system message
("model changed"), so the authoritative latest page holds one more message
than the window looking at the same chat. The stale-transcript guard measured
staleness as `remoteChat.length > localMessages.length`, so a session that had
switched models reported "This window was behind another view of the same
chat", refused the send, and repeated the refusal on every retry. No second
window existed, and nothing in the session was damaged.
Measure authored content instead of array length:
* hydration.ts marks a converted backend notice with ChatMessage.systemNotice,
via one NOTICE_DISPLAY_KINDS predicate that now also drives the existing
system-role decision.
* stale-transcript-guard.ts compares authoredMessageCount on both sides, so a
notice never counts as another view's work.
* use-session-actions/utils.ts classifies the new field in IGNORED_FIELDS: the
transcript paints the row from role + parts, and role is already COMPARED.
Notices still render unchanged. Tool rows folded into an assistant bubble keep
the existing behavior. A genuinely forked chat is still refused, covered by
tests.
Trade-off: a difference consisting only of notices no longer installs the page,
so a "model changed" row can wait for the next natural hydrate. The send
proceeds, which is the point of the guard.
Tests: new apps/desktop/src/lib/stale-transcript-guard.test.ts, 5 cases (the
notice regression, both real-fork cases, the identical-page case, and the
empty-page contracts). The guard had no test before. Verified with
`vitest run --project ui` (8944 passed; the 2 failures in voice-prefs.test.ts
are pre-existing and reproduce with these edits stashed) and `npm run
typecheck` (clean).
Regression case carried over from #124414 (@Yun-0000), the same
non-idempotent prefix graft reached from a page whose first row is an
orphan tool fold. The fix in this branch already covers it; the case
pins the exact shape reported in #124311 next to the module's other
graft tests.
A transcript hydrated from the newest page can open on a page-local tool
fold: `toChatMessages` flushes a tool batch with no active assistant into a
synthetic message whose id is not durable, so its `rowId` is undefined.
`graftRefreshedTailOntoBackfill` treated any unstored row in front of the
anchor as proof of earlier history and re-prepended it to the refreshed page.
When the window itself was hydrated from that same page, the page already
carried the row, so the graft returned `previous.slice(0, anchor)` plus the
page: one row longer than the window it was given, on every read.
`messagesIfTranscriptBehind` compares lengths, so a graft that is not
idempotent reports "behind" forever. Both submit paths then `return false`
before `prompt.submit` runs, the toast promises a refresh that cannot reach
the compared length, and the transcript gains another duplicate per retry.
Keep the prefix branch, but drop only the prefix copies the refreshed page
already carries. Durable prefix rows still travel in front of the refreshed
tail, and so does an unstored row the page has no copy of.
An owned path like skills/research/web-search/scripts has no SKILL.md of
its own, so the category rule merged it and left retired files behind.
Look for SKILL.md in the dir and its parents below skills/.
The owned-category merge (#123646) treated a shipped dir as a category only
when it held nothing but DESCRIPTION.md and dotfiles. A category that also
ships README.md, LICENSE or any other metadata file was classified as a root
and replaced wholesale on update, deleting skills that `hermes skills install`
or the agent had added to skills/<category>/.
Under skills/, classify by the skill marker instead: a dir without SKILL.md
is a category and is merged per skill; a dir with SKILL.md is an authored
skill and is still replaced whole. Outside skills/ the old rule stands. The
merge loop, the symlinked-container walk and the pre-write guard all share
_is_container, so the guard still covers exactly what the copy merges.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
A shell with `set -x` (user rc, BASH_ENV) traces `+ echo <sentinel>` into
the merged output. That line is an extra separator for _split_segments, so
the segment count mismatched and read_file_raw (the V4A/replace write-back
source) failed with "Failed to read file".
_fenced_read now turns xtrace off before the fence; `set +x`'s own trace
goes to the group's discarded stderr.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Extends the archived_at test with a legacy archive flattened under its folder name (accelerate) whose record is keyed by the SKILL.md name (huggingface-accelerate). Red on 52a6835e2b^.
_archived_ts called get_record per archive dir, and each call re-read and
re-parsed .usage.json (N+1 reads). Load the map once before the scan and
index it; the state/archived_at checks read keys that _backfilled never
changes, so behaviour is identical. Also drop the separate _archive_dir
import in favour of the existing skill_usage module import and split the
~130-char conditional.
The legacy-archive rescue looked up the usage record by the archive dir
name. Older archives were flattened under the DIRECTORY name (`accelerate`
for skill `huggingface-accelerate`, as restore_skill already documents), so
the lookup found no archived record, fell back to the stale dir mtime and
purged a skill the record says was archived today.
Resolve the record key with skill_usage._read_skill_name (the same
frontmatter reader restore_skill uses), falling back to the dir name when
SKILL.md is missing or has no name.
Preferring the usage record's archived_at outright let a stale archived_at
beat a fresh dir mtime: set_state only rewrites archived_at on a state
change, so a skill moved out of .archive by hand (record still "archived")
and archived again keeps the old timestamp and is purged at once, even
though archive_skill just stamped a fresh mtime.
Take the newer of the two signals. Legacy archives (stale mtime, fresh
archived_at) still survive, the stale-record case now survives too, and any
disagreement errs toward keeping the archive, which is the no-data-loss
direction for an irreversible purge.
archive_skill now stamps the archive dir mtime, but archives created before
that fix still carry the skill's last-edit mtime, so the first
`hermes curator purge` after upgrading deletes an idle skill archived
yesterday. The usage record already stores archived_at (set_state), so key
the TTL on it and fall back to the dir mtime only when the record is not in
the archived state (collision-suffixed dirs, forgotten records).
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
archive_skill moves the skill dir into .archive/ with rename (or shutil.move, which copies the mtime), so the archive keeps the skill's last-edit mtime. `hermes curator purge` ages archives by that mtime, so a long-idle skill archived today was already older than any archive_ttl_days and got purged at once. Stamp the archive dir's mtime when it is archived.
(cherry picked from commit 1f48db8d3727677eac01dcfae70c7e3ee29e1bcd)
Memory, context files and plugin prose precede the Model:/Provider:/Platform:
trailer and may contain lines of their own with those labels. When the live
value is empty the trailer omits the line, so a prose line stood in for it:
with the identity check now failing closed on an emptied value, that read as
a mismatch on every turn and rebuilt the prompt each time.
Route commits no longer NULL the stored prompt, so _stored_prompt_matches_runtime is the
only rebuild trigger. A model-only request that clears agent.provider skipped the compare
and replayed the previous route's prompt verbatim. The builder omits empty trailer lines,
so the rebuilt prompt matches from the next turn (no rebuild loop); prompts without
identity lines keep reusing.
* fix(desktop): keep portal sign-in alive across redirect aborts
(cherry picked from commit 3de10973b8e234853ddd0a8ed15bdb66722f5728)
* refactor(desktop): share the ERR_ABORTED predicate between both cookie windows
The remote OAuth window (main.ts, #110308) and the portal window
(portal-session.ts) each carried the same inline `code === -3 ||
/ERR_ABORTED/` test. One resolver per policy: extract it to
oauth-navigation.ts so the two windows cannot drift on what counts as a
superseded navigation. Helper and its unit test come from #89377.
Co-authored-by: Matt Earls <matt@marketplacevelocity.com>
---------
Co-authored-by: Austin Pickett <pickett.austin@gmail.com>
Co-authored-by: Matt Earls <matt@marketplacevelocity.com>
The summary-interrupt path added a second copy of the string the loop's
interrupt path builds; turn_explainers matches its prefix, so both now
share interrupted_during_api_call_reason().
An InterruptedError from the now-interruptible max-iteration summary was
swallowed by handle_max_iterations' broad except and delivered as the
max_iterations_no_summary fallback with interrupted=False, so finalize_turn
cleared the pending interrupt message and CLI/gateway had nothing to requeue.
Propagate the cancellation, drop the unanswered summary nudge, and surface it
in the finalizer exactly like an interrupted loop API call: interrupted=True,
interrupted_during_api_call exit reason, INTERRUPT_WAITING_FOR_MODEL_PREFIX
text, and result["interrupt_message"] preserved.
Invariant for the #123987 fix: whether profile_build is "ask" (default,
profile-build offer) or "off" (plain intro), the first-contact note must
tell the model to do a real first-message task before the intro/offer.
_hmwa_first_contact_notes re-implemented the branch logic of
agent.onboarding.first_contact_turn_note (profile_build mode check,
is_seen, mark_seen, plain-intro fallback) that the TUI already uses, so
the gateway and TUI paths could drift apart. #123987 deduplicated only
the note literal. Call the shared helper instead; it already falls back
to PLAIN_INTRO_NOTE on error, so the local try/except goes away. The
has_any_sessions() gate stays.
Suggested in review of #123987 by jonpol01.
#123987 added the "do the task first" carve-out only to PLAIN_INTRO_NOTE,
which is used only when onboarding.profile_build is "off". The default is
"ask", so a fresh default install still received profile_build_directive,
which opens with "After a one-sentence introduction ... OFFER" and lets
the intro/profile offer replace a real first-message task.
Factor the carve-out into TASK_FIRST_CLAUSE and lead both first-contact
notes with it. The consent-gated profile-build steps are unchanged.
The zero-session first-contact sidecar note (PLAIN_INTRO_NOTE /
_hmwa_first_contact_notes) unconditionally told the model to just
introduce itself, with no carve-out for the case where the user's
first-ever message IS a real task. On a fresh tenant whose voice task
was the very first message, the model followed the note verbatim and
replied with a static 'I'm Hermes. /help shows the available commands.'
- no tool call, no attempt at the task at all. This is a third shape of
the first-turn-onboarding-hijack class (turn replaced outright, not
just augmented with the known profile-build pitch).
Fix: PLAIN_INTRO_NOTE now instructs the model to do the task first
(call whatever tools it needs) and fold the one-line intro into the
close of that same reply; only a message with no real request gets the
old bare intro behavior. Also de-duplicated the literal note text in
gateway/run_turn.py, which had drifted into an inline copy instead of
importing agent.onboarding.PLAIN_INTRO_NOTE.
(cherry picked from commit fc7c3839b0b6774133d4fe8df59a9e98d23bdd8b)
CI diagnostics (tmux 3.4 on ubuntu-24.04) show the /exit hang is tmux's:
after /exit the pane process is a single-threaded zombie of the tmux
server (Threads:1, PPid = tmux server), the server is running with
SIGCHLD caught, not blocked and not pending, and pane_dead_status never
fills for 60s. The TUI had already exited (its epilogue is in the PTY
transcript). A zombie tmux has not reaped after 2s now gets a SIGCHLD
nudge so tmux runs its waitpid loop and reports the real exit status; a
process that is still running is never touched and still times out.
Main now appends the install's first-contact onboarding note to the first
user message on the wire only (per-turn sidecar, never persisted). The
clarify cell's invariant is that the answer is not a second user turn, so
compare what the user typed. An un-reaped pane zombie on CI now reports
the pane process threads (state/wchan) and the tmux server's signal masks.
tmux marks a pane dead on pty EOF, which the kernel delivers when the exiting
process closes its last tty fd -- before SIGCHLD lets tmux reap it and fill
pane_dead_status. On a loaded CI runner the harness read the status in that
gap and failed first_exits_clean with an empty '/exit status '. Poll until
tmux reports the exit status or signal, and make every exit failure report
status/signal, raw tmux answer, pane pid state, elapsed time, the frame
before /exit and the tail of a pipe-pane PTY transcript.
- wait for the startup session (status bar 'ready') before the first submit;
PHASE names in harness errors
- known_failure pins (merge-order safe) instead of strict xfail; new cell for
/resume typed during startup being undone (#121456)
- verbose tool progress so tool output is rendered and checked exactly once
- width cells check the paragraph layout (whole words, rows read on) so a
stale-width frame goes red on shrink
- tmux socket under the test root (-S), removed on close
- persisted/screen, summary, exit and raw interrupt-partial checks