Commit Graph

37658 Commits

Author SHA1 Message Date
teknium1
3eb9712180 Revert "feat(catalog): restore website install links to Desktop"
This reverts commit cf59980e2c.
2026-09-19 00:11:36 -07:00
teknium1
007e80ee0f Revert "docs(catalog): restore native browsing and install guidance"
This reverts commit 8b551a1ab3.
2026-09-19 00:11:36 -07:00
teknium1
086628ad8a Revert "feat(ui): support icon-only segmented controls"
This reverts commit 2e05fcf524.
2026-09-19 00:11:36 -07:00
teknium1
8bd0da2b8c Revert "feat(desktop): restore native skill and plugin catalogs"
This reverts commit cbd76e4ea3.
2026-09-19 00:11:36 -07:00
teknium1
f923398fee Revert "feat(desktop): browse catalogs as cards with a saved list option"
This reverts commit 4e9d3c713a.
2026-09-19 00:11:36 -07:00
kshitijk4poor
25a43ddfb3 fix(tui_gateway): probe foreign rows by keyset and re-hydrate after a remote compaction
The review round on the final head found two things. The unconditional full
transcript decode on every prompt.submit cost ~155 ms at 3k rows even when
nothing was foreign; SessionDB.get_messages(after_id=) answers the same
question in ~0.04 ms, so the full read now runs only when the probe finds
rows below this turn's submit row.

_strip_in_memory_overlap was built on a wrong premise: archive_and_compact
re-stamps _row_id on the in-memory dicts, so a local /compact leaves nothing
above the boundary and the helper never ran. The epoch that does exist is a
compaction by another surface, whose summary row leads the foreign rows and
can never align with the in-memory tail. That case is now detected on the
probe rows (_compressed_summary) and the history is re-hydrated from the DB
the way session.resume does, below the submit row, instead of appended to.

Test 2 covers both epochs: local compaction is a no-op, remote compaction
replaces the history with summary markers intact and the submit row excluded.
2026-09-19 12:33:57 +05:30
kshitijk4poor
83f6c77ae8 fix(tui_gateway): next prompt sees turns another surface appended to the session (#42962)
A Telegram reply or cron run appended to a live desktop/TUI session repainted
in the transcript (change_watcher -> sessions.changed, #86588) but the next
prompt.submit still snapshotted the in-memory session["history"], which only
the agent's own results ever wrote. The model answered from a stale history
(live: agent.turn_context history=2 while state.db held 4 active rows).

Before the turn snapshot, _adopt_out_of_band_turns appends the active rows
between the highest _row_id the agent's own flushes stamped onto the in-memory
messages (sync_flushed_message_markers) and this turn's own user row, which
prompt.submit already wrote at send (#111868), and bumps history_version.
Rows a compaction re-inserted under fresh ids while the in-memory dicts keep
the archived originals' ids are recognised as the durable in-memory tail and
skipped, so /compact does not double the history on the next turn. A row-id
boundary rather than a (role, text) anchor: a foreign reply that repeats an
earlier one verbatim ("OK") must not clip the tail.
2026-09-19 12:33:57 +05:30
kshitijk4poor
603007ead3 refactor(desktop): folder-derived project name lives in the input only
Drop the derived resolvedName fallback: pickFolder already writes the folder
name into the input, so the fallback only fired when the user cleared the
field after picking, and then Create stayed enabled with a name the dialog
never showed. Submit and the button gate go back to name.trim().

The dialog test asserts the fields this flow owns (folders, name) instead of
the exact createProject argument object, which would break on any harmless
extra kwarg (dropPlacement already rides along); beforeEach sits next to the
mocks it resets.
2026-09-19 12:19:30 +05:30
Anton Vykhovanets
c63730f565 test(desktop): cover folder-derived project creation 2026-09-19 12:19:30 +05:30
Anton Vykhovanets
e1866bf7a6 fix(desktop): create a project from a folder in one step
When a folder is picked in the New Project dialog and no name was
typed, auto-fill it from the folder's last path segment. The submit
guard is relaxed so creation works with just a folder — the name is
derived as a fallback. This removes the redundant naming step for the
common case where the folder already has a meaningful name, and does
so entirely within the left-sidebar Projects flow (no right-rail or
per-session cwd involved).

Creating with `use: true` set the active-project pointer but left the
view scope at the overview, so a new chat still started detached. The
dialog now calls `enterProject` on success, matching what clicking the
project row does, so `pick folder → Create → working in it` completes
in one gesture.

The derivation imports the sidebar's existing `baseName` helper directly
from its pure workspace-groups module rather than hand-rolling a
POSIX-only split, so Windows paths resolve correctly without pulling the
Projects component barrel into the dialog test. Both the submit path and
the footer button's disabled state read one `resolvedName`, so the button
can never render enabled but do nothing.
2026-09-19 12:19:30 +05:30
kshitijk4poor
cb9cddc738 chore: map Anton Vykhovanets to @vykhovanets for #68621 salvage 2026-09-19 12:19:30 +05:30
kshitijk4poor
ded0789f9a docs(agent): the finish explainer's terminal set is a sibling heuristic, not the same one 2026-09-19 12:12:04 +05:30
kshitijk4poor
112446afd1 fix(agent): judge "wrong script" against the user's own message; one continuation kind per stop
Review findings on the first cut:

* The letters-but-no-ASCII arm flagged every terse non-Latin answer
  ("是", "Готово", "はい", "تم") after tool work, so a Chinese/Russian
  conversation on the Responses transport would pay a re-prompt on most
  one-word answers. "Wrong script" now means wrong relative to the
  conversation: when the user's message carries non-ASCII letters, a
  non-Latin reply is an answer. "пар" after an English prompt still counts.
* The leading-punctuation arm caught ":8080", ":)", ";;", ":=", "}". It now
  requires a letter right after the punctuation ("?warming", "!warming").
* finalize_turn chose the continuation kind twice (log ladder and nudge
  pick); it is computed once, the tool-row count too.
* tool_results_this_turn documents the load-bearing invariant: any user
  row (the nudges included) closes the window, which is what bounds the
  guard to one re-prompt per collapse.

Tests: user-script negatives in Chinese/Russian/Japanese/Arabic incl. a
multi-part user message; ":8080"/":)"/";;"/"}" negatives; both red when
the script check or the letter-after rule is dropped.
2026-09-19 12:12:04 +05:30
kshitijk4poor
80f0d1b52f fix(agent): re-prompt once when a turn that did tool work ends on a collapsed fragment
A Responses-wire collapse (#103483): the tool work runs correctly, then the
final text stop is a fragment — a stray wrong-script word ("пар"), a token
starting mid-punctuation ("?warming up") — and the loop accepted it as the
answer, so the turn reported completed and an unattended job abandoned the
task.

The guard rides the existing ack-continuation path in finalize_turn: same
scope knob (agent.intent_ack_continuation, default auto = Responses
transports), the same bounded per-turn counter, the same durable interim +
nudge rows, and the same synthetic-user recognition in the compressor. It
fires only when the turn has tool results after the last user row and the
answer matches a deliberately narrow predicate: <= 24 chars, no sentence
terminal, and either leading punctuation no answer begins with or letters
with no ASCII letter/digit at all. "42", "SQLite", "report.csv", "€12.50",
"你好。", "Done." never match. Because shape cannot prove a collapse, the
nudge asks for the same answer again when it WAS complete, so a false
positive costs one call, never the answer. The nudge row closes the
tool-work window, so a second fragment ends the turn as before.

Rebuilt from PR #111472 by dankkush against current main; the mid-task
"stall note" arm and the ephemeral-row popping are not taken (the former
false-positives on declarative answers, the latter buries flagged rows once
a recovery response calls a tool).

Co-authored-by: dankkush <brandan@wrengineers.com>
2026-09-19 12:12:04 +05:30
dankkush
f0b88402dc test(agent): degenerate-final recovery — a collapsed fragment after tool work is re-prompted, not accepted
Loop tests through run_conversation with a mocked client (harness mirrors
test_dropped_tool_call_recovery): fragment after a tool round is re-prompted
with the degenerate-final nudge and the real answer lands; a second fragment
ends the turn (bounded, never a loop); a chat-only terse reply is untouched.
Predicate tests pin the fragments the guard exists for and the terse
legitimate answers a shape-only guard was shown to re-prompt.

Adapted from PR #111472 to the reused ack-continuation path.
2026-09-19 12:12:04 +05:30
kshitijk4poor
ad28e3a558 chore(contributors): map dankkush (PR #111472 salvage) 2026-09-19 12:12:04 +05:30
kshitijk4poor
6338e988bd refactor(shared): type the handshake close as CloseEvent; drop the dead error-detail branch
Review follow-up: the close listener is a CloseEvent by construction (code is always a number), so the
`as { code?: unknown }` cast and the 'unknown' fallback guarded nothing. Renderer 'error' events carry no
message, and every consumer of this client is renderer-side, so the `: detail` suffix was never produced —
the class alone is the detail. Test: the fake socket loses its unused OPEN/static-last bookkeeping and the
prefix test asserts the real contract (prefix kept, not equal to the bare message) instead of prose.
2026-09-19 12:11:50 +05:30
kshitijk4poor
efe042abe9 fix(shared): a failed gateway dial says which failure it hit
JsonRpcGatewayClient.connect() rejected every failure with the bare connectErrorMessage, so the
Desktop boot overlay showed "Could not connect to Hermes gateway" whether the server closed the
handshake with 4403 token_mismatch, the socket errored before opening (TLS, DNS, refused) or nothing
answered within the connect timeout. Reporters on #41566 verified their gateways with curl and a
plain WebSocket client and still could not say what the app had hit.

The rejection now carries the failure class: "WebSocket closed during handshake: code 4403
token_mismatch", "WebSocket error before open[: detail]" or "no WebSocket open within N ms". The
base message stays as the prefix, so the web sidebar's includes()-based transport matchers and the
desktop recovery tests are unaffected.

Refs #41566
2026-09-19 12:11:50 +05:30
kshitijk4poor
4626656c74 refactor(tui_gateway): evaluate the quiescence predicate outside _sessions_lock
Re-gate follow-up: _session_is_lru_evictable may read state.db (_session_has_active_delegations) and the
predicate now also runs on the turn thread at every turn end; the repo already bounds state.db work under
_sessions_lock (prompt_turn.py routing heal). Snapshot the other sessions under the lock, evaluate outside —
the answer is advisory either way.
2026-09-19 12:11:44 +05:30
kshitijk4poor
e0e54981e2 fix(tui_gateway): the turn-completion trim waits for the other sessions too
Review follow-up: prompt_turn.py::_finish_turn ran the non-forced "tui turn completion" trim on every turn
end from the turn thread, while sessions B..N in the same process could be mid-turn — the same GIL +
arena-lock stall the reaper gate closes, and it fires more often than the 300 s scan. The quiescence
predicate moves to session_reaper.py::_sessions_quiescent(exclude=) so both callers share it; the finishing
session is excluded because it is still marked running at that point, so a sole TUI session keeps its
post-turn trim.

Tests: one contract for the sibling path; the two pre-existing trim tests adopt the shared stub helper, the
"every scan" docstring that the gate made false is folded into the quiescent case, and an inert age
override is dropped from the busy case.
2026-09-19 12:11:44 +05:30
kshitijk4poor
11b7a1e7fe fix(tui_gateway): the periodic memory trim waits until no session is busy or attached
The idle reaper ran hermes_cli.mem_trim.trim_memory on every 300 s scan regardless of what the
sessions were doing. The trim's gc.collect() holds the GIL and glibc malloc_trim(0) takes every arena
lock, so on multi-GB heaps the event loop stalled 20-50 s each scan — long enough to blow the 10 s WS
write deadline, drop the Desktop client and interrupt its running turn (#58576, Ufonik88's
duration_ms 21065/49632 matching the stall lengths 1:1).

The periodic trim now runs only on a quiescent scan: no session mid-turn, building, awaiting input,
holding live delegations, or on a live transport (the LRU reaper's existing exemption predicate).
Forced trims (agent close, cache pressure) are unchanged.

Refs #58576
2026-09-19 12:11:44 +05:30
kshitijk4poor
4f766f8fa5 fix(desktop): profile-only secondaries count their live turn too
Re-gate follow-up: liveSessionScopes() holds only registry-scoped `conn:…::profile` keys, so a
local/legacy secondary (connectionId null, scope = bare profile) never registered a turn in flight and
one missed wake ping closed it mid tool call. The boot hook now derives one liveWorkScopes() — registry
scopes plus the bare profile of every working/needs-input local session — that feeds both the pruner's
keep-set (unchanged behaviour) and the registry's liveScopes hook; the probe matches the same way
foregroundPinned does (composite key, or bare profile for a connectionId-less entry). The hook is
referenced lazily because the registry is configured earlier in the effect than the helper's const.

Also restores the misplaced doc comment on LIVENESS_PROBE_FAILURE_STREAK.
2026-09-19 12:11:26 +05:30
kshitijk4poor
e796a29866 fix(desktop): wake probe counts the foreground turn as in-flight work; sibling pins and prune follow-through
Review follow-ups on the #112834 salvage:

- The secondary's wake probe fed `activeRequests` into decideLivenessForceClose, but prompt.submit
  returns before the turn ends, so a foreground turn mid tool call showed 0 and ONE missed ping
  closed the socket — the #95327 false-kill the streak policy exists to prevent. The registry now
  exposes `liveScopes` (session-states.ts::liveSessionScopes, already used by the pruner's keep-set)
  and the probe counts a live scope as work in flight. The lifecycle test drives that production
  shape instead of a parked prompt.submit.
- A relay-retained secondary (bot relay pin) was still blind-closed on a forced wake; it joins the
  probe-instead-of-close predicate the pruner and redial drain already use.
- The pruner is event-driven, so a socket spared by the 30 s grace with no store change afterwards
  held its pool slot forever; the existing 60 s keepalive tick now recomputes the keep-set.
- A second wake inside the 15 s holdoff skipped reconnectNow entirely (primary probe, redial of
  already-closed secondaries); only the destructive forceOpenSockets half is coalesced now.
- Shape: LIVENESS_PROBE_TIMEOUT_MS lives once in gateway-liveness-policy.ts (was mirrored in two
  files); the -32601 carve-out uses lib/gateway-rpc.ts::isMissingRpcMethod; the always-true
  `typeof gateway.request` guard, the `openedOnce` twin of `lastOpenedAt` and the dead `> 0` guard
  are gone; SECONDARY_MIN_LIFETIME_MS is exported so the tests stop hard-coding 31_000.
2026-09-19 12:11:26 +05:30
psam
ad6c7a5a4c fix(desktop): defer secondary wake-probe force-close behind a failure streak
One unanswered 5s ping on a live-in-use secondary could fire while the
backend's event loop is starved by a long tool call: force-closing it fed
the backend's ws_orphan_reap and interrupted the valid turn (#94769
review). The secondary wake probe now applies the SAME policy the
primary's reconnectNow probe uses (decideLivenessForceClose,
gateway-liveness-policy): while a request is in flight, the first
failure defers behind a bounded re-probe (LIVENESS_REPROBE_DELAY_MS) and
only an exhausted streak (2 unanswered probes) closes the socket; an
idle-but-pinned socket still closes on the first failure. The streak
resets on any answered ping (including the -32601 version-skew
carve-out) and on every fresh socket open; a deferred re-probe is
cleared on dispose and re-dial.

Tests: the wake-probe regression test now holds an in-flight request and
proves the first timeout keeps the socket while the unanswered re-probe
closes it (red on the previous commit, green here).
2026-09-19 12:11:26 +05:30
psam
a915f0b3b2 fix(desktop): probe live secondaries on forced wake instead of skipping them
Review follow-up for #94769: a live-in-use socket cannot simply be SKIPPED
by the forced wake path either — a half-open TCP connection reports 'open'
forever and fires no close event, so an in-flight request would ride the
dead transport until its per-call timeout (30 minutes for prompt.submit).
On backends that predate ping/heartbeat there is no other healing path.

The wake path now runs the same bounded ping probe the primary's
reconnectNow uses: a healthy-but-busy backend answers and keeps its socket;
a dead transport is closed and healed by the entry's ordinary reconnect
backoff. -32601 (method not found) keeps the socket — a version-skewed but
healthy backend, the same carve-out the primary's probe makes.

Also stop burning the 15s forced-wake holdoff while boot is incomplete or a
gateway switch is in flight — reconnectNow no-ops there, so stamping the
holdoff would drop the next 'online' (often the one with the network
actually back) and leave recovery to the backoff loops.
2026-09-19 12:11:26 +05:30
psam
69c51bb008 fix(desktop): stop forced wake reconnects from churning live secondary sockets
Multi-profile desktop setups flicker continuously: every wake signal
(power resume, network 'online', focus) forced a reconnectNow({forceOpenSockets:
true}) that closed EVERY open secondary socket and redialed it, and the
live-work pruner could dispose a freshly opened idle socket before its
consumer registered in the keep-set. Closing a socket a mounted surface is
bound to detaches its runtime, the backend orphan-reaps it, session.reclaimed
unbinds the surface, and the re-resume lands on a fresh socket the same
signal closes again — a reconnect/remount loop every 2-5s (#94769).

Three guards, all on the renderer side:

- Coalesce forced wake reconnects: one forced reconnect per 15s holdoff
  (renderer twin of the main process's POWER_RESUME_REVALIDATION_HOLDOFF_MS).
  Windows fires 'online' on any interface change; a socket dropped in the
  holdoff window is still healed by the ordinary close/reconnect backoff and
  the non-forced focus/visibility nudges.
- reconnectSecondaryGateways({forceOpenSockets}) spares sockets with in-flight
  requests or a foreground-pinned surface — the same pins the live-work
  pruner already honors (#93892).
- pruneSecondaryGateways grants freshly opened secondaries a 30s
  min-lifetime grace so the prune <-> on-demand-dial race cannot close a
  socket its consumer has not registered yet. Aged sockets reap exactly as
  before.

Fixes #94769
2026-09-19 12:11:26 +05:30
kshitijk4poor
8600cb0744 chore: map psam-717 contributor email for #112834 salvage 2026-09-19 12:11:26 +05:30
kshitijk4poor
afaa53e5fe fix(agent_init): judge the served window only for local endpoints; clamp after the floor as before
Review findings on the first cut:

* A hosted provider with a stale model.ollama_num_ctx passed the floor for
  a 40K model although nothing ever raises a hosted window. The served
  window now counts only when the endpoint is local (is_local_endpoint),
  the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
  compressor clamp ahead of it, which flipped two observables: a
  model.context_length above a sub-64K num_ctx was rejected instead of
  constructed-and-clamped, and the floor's message reported the clamped
  value with advice (set model.context_length) that could not help. The
  phase is split: resolution runs before the floor, the clamp
  (_clamp_compressor_to_ollama_num_ctx) stays at its original position,
  so every case main constructed still constructs with identical
  compressor numbers and the floor's message is unchanged.

Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
2026-09-19 11:44:15 +05:30
kshitijk4poor
3de7140bcd fix(agent_init): the 64K context floor judges the window Ollama serves, not the GGUF metadata
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".

num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.

Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-19 11:44:15 +05:30
kshitijk4poor
14a3463454 fix(anthropic): mirror only the Keychain item that held the spent pair; parse both security attribute encodings
Review findings on the mirror, all verified live against throwaway items:

* Identity gate. The service name is fixed but the file path honours
  CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
  item under the service can hold a different login's pair or a newer
  rotation. Overwriting it would be the bug in the other direction. The
  refresh token that was just POSTed is threaded through
  `_write_claude_code_credentials(spent_refresh_token=...)` from both
  callers (the singleton refresher and the pool commit) and the mirror
  only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
  is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
  as `0x<HEX>  "<echo>"` otherwise. The old regex modelled `\"` escaping
  that never happens: a `"` in the account truncated it and the mirror
  created a second item; any non-ASCII byte made it return "" and the
  mirror silently no-oped. One `find-generic-password -g` call now yields
  account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
  succeeded when the mirror runs, and a raise here made the refresher
  mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
  `-w` returns non-ASCII payloads as hex.

Two pre-existing test doubles for the writer accept the new keyword.
2026-09-19 11:25:23 +05:30
kshitijk4poor
bf5a6f6ae6 fix(anthropic): mirror the Keychain item through security -i with a hex payload, under its own account
Two defects in the cherry-picked mechanism, both found live on macOS:

* `add-generic-password -w` with no value prompts on /dev/tty when a
  terminal exists, so the CLI refresh path hung for the 10 s timeout and
  wrote nothing; with no terminal it read one line from stdin, hit EOF on
  the confirmation read and stored an EMPTY password — bricking the very
  item this exists to keep fresh. The command line now goes to
  `security -i` on stdin with the payload hex-encoded (`-X`): no argv, no
  tty prompt, no quoting of the JSON.
* `-a getpass.getuser()` assumed the item's account is the login user;
  `-U` matches on account + service, so a mismatch would have created a
  second item Claude Code never reads. The account is read from the
  existing item's `acct` attribute and the mirror is skipped without one.

The command builder is a pure function so the no-argv / hex / account
contract is tested on every lane; the Darwin-gated no-entry no-op test is
`macos_only` instead of faking `platform.system`. Tests trimmed to the
invariant bar; the merge test pins that `mcpOAuth` siblings survive.

Live E2E on a throwaway Keychain item seeded with 24 mcpOAuth entries:
triple rotated, scopes/subscriptionType and all 24 siblings preserved,
one item, updated under the seeded account.
2026-09-19 11:25:23 +05:30
salch-cred
6d4ca15d38 fix(agent): mirror Claude Code OAuth refresh into the macOS Keychain (#98334)
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").

_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).

Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
2026-09-19 11:25:23 +05:30
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
29c98eecc6 refactor(gateway): one call shape for the re-stamp seam
`_profile_name_for_source` already accepts `adapter_profile=None`, so the
branch in `_stamp_routed_profile` was two spellings of the same call.
Collapse it and give the two test doubles that replace the method the
same keyword so they keep matching the production signature.

Salvage follow-up to #106019.
2026-09-18 22:36:24 -07:00
Italo Fernandes
032b91d9b9 fix(gateway): reject inbound events when route matching fails
A matcher exception fell through to the default profile, so a transient
failure while resolving a sender route silently served the message from the
default profile's runtime. Raise the existing rejection instead, which the
ingress gate already drops on.

Only the failure path changes: a source that simply matches no route keeps
falling through to the default/active profile as before.
2026-09-18 22:36:24 -07:00
Italo Fernandes
e864bedf95 fix(gateway): close sender-routing gaps found in review
Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.

Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.

An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
2026-09-18 22:36:24 -07:00
Italo Fernandes
d6f3285b81 feat(gateway): route inbound messages to profiles by sender user_id
Closes #33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.

Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.

Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
2026-09-18 22:36:24 -07:00
kshitijk4poor
9dee8863e1 docs(model_switch): say what the fallback actually relies on
The switched-provider comment implied the guard catches the resolver raise;
it is suppress() leaving st.* at the session values that does. The
stay-on-provider comment now states the one semantic widening of moving to
the shared helper: an empty resolver result keeps the session endpoint on
any host, not only off openrouter.ai.
2026-09-19 10:43:50 +05:30
kshitijk4poor
0cb3760383 fix(model_switch): treat a fall-through to OpenRouter's default as "no custom endpoint configured"
Gate finding: resolve_runtime_provider(requested="custom") never returns an
empty base_url — with no model.base_url but an OpenRouter key present it
lands on OpenRouter's default host, so the "keep the current endpoint"
fallback was dead and an Anthropic session adopting `provider: custom`
hopped to openrouter.ai. The guard _creds_for_current_provider already
carries for this (#74143) is now a shared helper used by both branches;
the branch assigns once instead of assign-then-undo; the import sits with
the other from-imports; the new test's write_text carries an encoding.
Second invariant test pins the no-endpoint case.
2026-09-19 10:43:50 +05:30
kshitijk4poor
72749b2482 test: read config with encoding=utf-8 in the custom-provider switch tests
Pre-existing windows-footgun hits in the file this PR touches; the lint
lane wants the whole file clean.
2026-09-19 10:43:50 +05:30
kshitijk4poor
daf882e3c2 fix(model_switch): switching to a bare custom endpoint from another provider resolves the configured one
The bare-custom credential branch kept the session's current base_url and
key whenever the target was `custom`, regardless of where the session was.
That is right for a bare-custom session picking another model on its own
endpoint, but when the TUI's per-turn config sync adopts `provider: custom`
from an OpenRouter (or any other) session, the new model was paired with
OpenRouter's URL and key and every request 400/404'd (Bug 2 in issue 73680).

Arriving from another provider now resolves the configured custom endpoint
through resolve_runtime_provider; with none configured the current endpoint
is kept, so direct aliases still supply their own URL afterwards and the
bare-custom same-endpoint case is unchanged.
2026-09-19 10:43:50 +05:30
kshitijk4poor
cf13448580 refactor(agent): name the tool-call XML namespace prefix once
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
2026-09-19 10:41:08 +05:30
DoGMaTiiC
60735cbf81 fix(agent): strip namespace-prefixed text-channel tool-call XML from visible text
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).

Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).

With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.

Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).

Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
2026-09-19 10:41:08 +05:30
kshitijk4poor
956dadf0ad chore: map DoGMaTiiC's contributor email for the #114136 salvage 2026-09-19 10:41:08 +05:30
teknium1
c07708671d fix(gateway): every adapter session key goes through one seam (+ lint)
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.

Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.

Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.

Phase 2 of #88715.
2026-09-18 22:04:43 -07:00
teknium1
ad651b8250 feat(gateway): RoutingIdentity — one frozen identity per inbound event
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.

Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).

Phase 1 of #88715; the gateway rows of #90142 / #93943.
2026-09-18 22:04:26 -07:00
teknium1
03fee43ca3 chore: map contributor email for @Caelier (#100423 salvage) 2026-09-18 20:57:07 -07:00
teknium1
c1ad9daa8f test: two invariant tests for Codex singleton adoption, replacing the salvaged suites
The independent-account test drives the real load_pool() -> _refresh_entry() path
and asserts the second account POSTs its own refresh token and keeps its principal
(red on base: the row silently became account A). The alias test pins both halves
of the same-account rule: never fall back onto an older singleton, always follow a
newer re-auth.
2026-09-18 20:57:07 -07:00
Tranquil-Flow
20512a47a6 fix(agent): refuse stale singleton adoption over rotated manual:device_code entries (#106705)
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).

Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
2026-09-18 20:57:07 -07:00
Marc Caelier
46503f1672 fix(auth): isolate Codex singleton sync by principal 2026-09-18 20:57:07 -07:00