Commit Graph

408 Commits

Author SHA1 Message Date
Teknium
4fda3bbdca Merge pull request #115846 from NousResearch/fix/boa-tts-stt-voice-tts-config-minlen
feat(tts): streaming TTS speaks a short first sentence sooner via tts.streaming.min_len on every surface (#96927, salvage #96933)
2026-09-19 11:33:15 -07:00
Teknium
0d2cb2cacc Merge pull request #115890 from NousResearch/fix/boa-res-R6-streaming-compaction-astra-oauth-gate
fix(compression): gpt-6-astra on Codex OAuth gets native server-side compaction when opted in (#103720, salvage #103718)
2026-09-19 11:28:10 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
Teknium
7feb1af028 Merge pull request #115872 from NousResearch/fix/boa-res-R4-routing-catalog-azure-codex-900k-cap
fix(codex): opted-in -900k aliases cap at the live catalog max_context_window (#105443, salvage #105445)
2026-09-19 11:12:31 -07:00
teknium1
ec2c4eeb80 chore: merge origin/main (resolve apps/desktop/src/lib/voice-client-direct.ts, hermes_cli/config_defaults.py, tools/voice_client_config.py, website/docs/user-guide/features/tts.md) 2026-09-19 10:54:38 -07:00
teknium1
e85cb94da1 chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:40 -07:00
teknium1
8e3a833365 chore: merge origin/main (resolve website/docs/developer-guide/context-compression-and-caching.md) 2026-09-19 10:47:48 -07:00
teknium1
a516db44e7 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-19 10:46:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
4c4b0b748e docs: describe the summary-provider-overload abort in the compression failure ladder 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
26675fbcb7 docs: state that codex_responses_compact_threshold applies only under codex_responses_native
The compression config table and the yaml example described
`codex_responses_compact_threshold` as "the server compaction trigger"
without saying it is read only while `codex_responses_native: true`
(agent/native_compaction.py::native_compaction_context_management returns
early otherwise). Users set it expecting local compaction to move and
filed #101867. Name the gate in the table row, the yaml comment and the
native-compaction prose, and point at `threshold` / `threshold_tokens`
as the local trigger.

Fixes #101867
2026-09-19 10:27:28 -07:00
teknium1
fbbc7dfe5d fix: strip stale Codex reasoning on 401 token_expired before the credential pool
With a multi-entry openai-codex pool, a 401 `token_expired` caused by a stale
replayed `encrypted_content` blob went pool-first: each healthy entry was
force-refreshed (single-use refresh token) or benched STATUS_EXHAUSTED before
the strip in `_recover_format_errors` finally ran, so pooled users lost every
account for the bench window over a session-state problem. The single-credential
path likewise burned a forced OAuth refresh on a bearer that was fine.

`recover_after_classification` now takes the Codex stale-reasoning strip first
(after the Nous welcome-tier repair) when the 401 carries `token_expired` and
the transcript still holds `codex_reasoning_items`; the same one-shot latch
bounds it, so a real expiry pays one extra round-trip and then takes the pool /
refresh path exactly as before. The strip body moved into
`_recover_stale_codex_reasoning`, shared with the 400 `invalid_encrypted_content`
branch. Docs: one line in the OpenAI Codex path section describing the
self-heal.

Tests: pooled control with a real two-entry CredentialPool (red on the previous
ordering: pool rotated, entry benched; green now: strip runs, nobody benched);
the single-credential test now asserts no refresh is burned before the strip.
2026-09-19 09:32:00 -07:00
teknium1
0c24dd56cb fix: pin the auxiliary wire's route-scoped reasoning_details strip; document OpenRouter/Nous-only replay
Replaces the direct transport.build_kwargs test with one driving prepare_chat_messages
(groq client strips, OpenRouter client keeps), so dropping the base_url kwarg goes red.
2026-09-19 09:28:45 -07:00
kshitijk4poor
7c6f21a5e1 docs(compression): examples and the delegation trigger note follow the default threshold_tokens cap
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
2026-09-19 16:28:57 +05:30
kshitijk4poor
cfbc6e51f4 docs(compression): developer guide lists the threshold_tokens cap next to threshold 2026-09-19 16:28:57 +05:30
teknium1
d899d890f8 fix(gateway): a restored lane whose receiving bot is offline delivers nowhere
Composing PR-4 (intake/delivery split) with PR-5 (persisted transport_profile): the delivery
fallback to the runtime profile's unique adapter is only for sources with NO identity. A pinned
identity names the receiving bot; if that bot has no adapter it is offline and the lane fails
closed, never answering from the runtime profile's bot. Also: tests and docs reference the split
helper, not the removed _adapter_for_source.
2026-09-19 02:28:50 -07:00
teknium1
04a151c3e0 docs(gateway): identity survives restore, relay, callbacks and thread hops 2026-09-19 02:28:50 -07:00
teknium1
e09cd0d2bc fix(relay): echo the routed profile on every outbound frame and follow_up
The connector stamps `profile` on inbound and passthrough_forward frames but the
gateway never sent it back, so the connector had nothing to stamp on the NEXT
interaction of a routed chat. `_capture_scope` now remembers the routed profile
per chat, `_with_scope` echoes it as `metadata.profile` on chat-addressed frames,
and `send_follow_up` derives it from the `agent:<profile>:` key namespace. A
single-profile gateway emits no key — frames stay byte-identical. Contract §4
documents the round-trip.
2026-09-19 02:28:50 -07:00
teknium1
19ca2e505a docs(gateway): document the intake vs delivery transport matrix
Five topology rows: per-credential, shared credential → satellite, shared bot → profile with its own bot (reply via the receiving bot), secondary bot → default, and restored/synthetic (intake fails closed; delivery via the unique owner). Part of #88715 (phase 4).
2026-09-19 01:28:23 -07:00
teknium1
2c4d1f79e2 fix: document the reasoning-only stall failover under Fallback Model
Configured fallback_providers now also engage on a Codex reasoning-only stall
(incomplete_response) with one bounded grace call; the developer guide listed only the
429/5xx/401/403 triggers.
2026-09-19 01:15:11 -07:00
GodsBoy
33b2461be4 fix(compression): enable Astra native compaction on official Codex OAuth 2026-09-19 00:53:26 -07:00
teknium1
a585ad83b2 fix: keep the Codex catalog fetch shape while capping -900k at max_context_window
Store the catalog's max_context_window in a parallel per-token dict populated
by the same fetch instead of widening the fetch's return tuple and versioning
the cache key ("v2:"), so the one external caller and every existing cache
reset keep working. The cap moves into _apply_verified_bump as
min(900K, live max): the bump still fires only for an opted-in -900k alias
whose advertised window is exactly 272K, 900K stays the offline/absent
fallback and a catalog max above 900K does not raise it.

Trim the salvaged tests to two invariants (parametrized alias cap incl. the
absent and above-cap controls; base slug never inherits the max) and document
the live-catalog cap next to the -900k opt-in.
2026-09-19 00:40:24 -07:00
teknium1
b408544510 test(gateway): identity-canonical ingress invariants (#88715 rows A/B) + docs
Two red-on-base invariants with the real GatewayRunner resolvers and real
BasePlatformAdapter ingress against a temp home with three served profiles:
per-credential collision (two bots, same chat.id == user.id: distinct lanes,
/stop on A cannot touch B's run, clarify reply on A resolves A's pending) and
shared-credential routed chat (runtime satellite, transport the receiving bot,
busy path keyed on the satellite lane; an unserved route dropped at fresh,
batched, busy and control ingress alike).
2026-09-19 00:30:29 -07:00
liuhao1024
9823ce4ff1 feat: make the streaming TTS first-sentence threshold configurable (tts.streaming.min_len)
SentenceChunker hard-coded min_len=20 and all three construction sites
(gateway StreamingTTSConsumer, CLI/TUI stream_tts_to_speaker, dashboard
/api/audio/speak-stream) called SentenceChunker() with no arguments, so a
short CJK opener ("记得,叫团团。", 7 chars) was always buffered behind the
second sentence, delaying the first audible audio by roughly one LLM
sentence in voice setups. Add tts.streaming.min_len (default 20, unchanged
behaviour) to DEFAULT_CONFIG and a SentenceChunker.from_config(tts_config)
constructor that every construction site now uses, so the knob applies
identically on every speech surface. Invalid values keep the default; 0
floors to 1 so chunking cannot be disabled by accident.

Slim redo of PR #96933 by @liuhao1024 (added a second
first_sentence_min_chars key and per-emission state), credited as author.

Fixes #96927
2026-09-19 00:16:47 -07:00
Italo Fernandes
d6f3285b81 feat(gateway): route inbound messages to profiles by sender user_id
Closes #33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.

Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.

Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
2026-09-18 22:36:24 -07:00
teknium1
ad651b8250 feat(gateway): RoutingIdentity — one frozen identity per inbound event
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.

Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).

Phase 1 of #88715; the gateway rows of #90142 / #93943.
2026-09-18 22:04:26 -07:00
teknium1
b71359066c fix: /stop halts background subagents and returns their partial results as interrupted completions
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:

- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
  already did this in `_handle_reset_command`, the second request is
  idempotent) and the idle `_handle_stop_command` tail, which replied
  "No active task to stop." while a background child was running; it now
  stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
  spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.

The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.

An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.

Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.

Part of #114456
2026-09-18 20:34:42 -07:00
teknium1
e56ec8c9f7 fix(chronos): a 403 invalid_client from NAS hands cron fires to the built-in ticker (#97494)
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.

NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.

Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
2026-09-18 20:07:21 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
Konstantin Khlopkov
ece262150a docs(memory-provider): state that no bundled provider advertises checkpoint API v2
Operators enabling compression.checkpoint_required need to know up front that
the contract is opt-in for third-party archiving providers; otherwise the flag
blocks every compression attempt and the reason is only visible after the fact.

Docs hunk of #106882 (the code change was consolidated with #106879).
2026-09-18 13:58:52 -07:00
kshitijk4poor
9e6f02538e docs(memory-provider): agent_context carries cron/subagent, not always primary 2026-09-19 00:55:05 +05:30
teknium1
aee7de4db5 fix(desktop): point the remote-backend desktop-half tooltip at the install path
The honest "unavailable (remote backend)" state (previous commit) tells the
user the copy will never happen; the tooltip now also says what does work
against a remote backend — Install from Git with the Desktop target checked
clones the desktop half onto this machine (the install modal already takes
that branch for connection.mode === 'remote'). Docs note the new state next to
the existing remote-backend paragraph.
2026-09-18 10:53:41 -07:00
teknium1
66ba4e5114 test(tui_gateway): pin that a live-turn rewind is not redirected; document the 4009 contract
The contributor test covers the queue fall-through; the Desktop's normal case is
a redirect-capable AIAgent, where busy_input_mode=interrupt turned the edit into
a mid-turn redirect and left the un-edited transcript in place. Pin that branch
too, and document in the rewind section that a truncating submit refuses with
4009 while a turn runs so hosts interrupt + retry (the Desktop already does).
2026-09-18 10:52:27 -07:00
teknium1
14d862384b fix(desktop): page-owned header controls get their own area; titleBar.center stays permanent
`444c75c10a` rendered `titleBar.center` in the workspace panel header on
full pages so the kanban board switcher would sit beside the page title
instead of colliding with the sidebar tab strip. That made the area migrate
between two React subtrees on chat <-> page navigation: the old instance's
effect cleanup ran after the new instance's setup and wiped every global
side effect a third-party plugin had just re-created (#114290).

Keep both intents: `titleBar.center` is a permanent titlebar slot again (the
salvaged commit), and page-owned controls move to a dedicated
`WORKSPACE_PAGE_HEADER_AREA` (`workspace.pageHeader`) that the workspace
pane projects into its vetoed tab row while `$workspaceIsPage` holds —
exactly the placement `444c75c10a` introduced, just not through the
plugin-facing titlebar area. Kanban's board switcher contributes there.

- SDK exports `WORKSPACE_PAGE_HEADER_AREA`; `TITLEBAR_AREAS` documents the
  permanent-mount contract.
- Test: a `titleBar.center` component's effect runs setup once and cleanup
  never across a chat -> /skills -> chat round trip (red on base).
- Docs: plugin SDK guide covers the lifecycle contract and the page-header
  area.
2026-09-18 10:45:44 -07:00
teknium1
b941f41ead docs(acp): every tool call reaches a terminal status (event bridge) 2026-09-18 10:33:54 -07:00
teknium1
d9ca304047 docs(acp): note the failed-turn transcript boundary in the ACP lifecycle 2026-09-18 10:33:21 -07:00
teknium1
aefa503479 fix(browser): gave-up supervisor unregisters itself; tests trimmed to two invariants
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:

- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
  which logs the single final warning AND drops the supervisor from
  ``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
  (``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
  attached" instead of a dead entry, and the next browser call for the task
  starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
  (two invariants: bounded + unregistered after attach; first-dial failure
  still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
  main — that file is the opt-in real-Chrome E2E suite and its module-level
  skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
  reconnect.

Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.

Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
2026-09-18 10:28:40 -07:00
teknium1
f39ae9000c docs(sessions): say what user_id holds for desktop and dashboard sessions 2026-09-18 10:16:46 -07:00
teknium1
eafed27cf0 fix(cron): a completed run keeps its result when a fire-claim heartbeat sample misses
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.

Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
  heartbeat miss is re-checked against the store, and a claim that still
  validates keeps the run's real outcome. The owner-fenced mark_job_run and
  fire_claim_fence remain the authority, so a genuinely re-owned claim still
  yields (no delivery, no terminal write over the new owner, ledger discard).
  An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
  _FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
  agent run and ends the lease refresh, so a single sample must not do that; a
  genuinely re-owned claim misses twice and latches ~1 s later.

The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.

Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:57:02 -07:00
teknium1
c61bbc2461 fix(gateway): persist the authored text, not the Discord triggering note, as the user row
The `[Triggering message id: …]` note that `_prepend_inbound_reply_context`
adds for Discord turns is a model instruction (which id to pass to the
discord tools), not something the user wrote. `_hmwa_apply_message_timestamp`
derived `persist_user_message` from the already-wrapped text, so every
Discord-origin user row stored the note as the first line of `content` —
the desktop transcript (and FTS, memory providers, exports) showed the
envelope instead of the message.

- `run_inbound.py`: the note is now the OUTERMOST prefix (after the reply
  pointer) and rendered by one function, `discord_triggering_note`;
  `strip_discord_triggering_note` peels exactly that prefix for THIS event
  off the persisted text, so the `[Replying to: "…"]` pointer survives.
- `run_turn.py::_hmwa_apply_message_timestamp`: persist the stripped text.
  The wrapper keeps riding `message_text`; when the durable row differs, the
  live bytes land in the replay-only `api_content` sidecar by the existing
  persist-override contract (same as timestamps and per-turn sidecar notes).
- `run_turn.py::_run_agent_queued_followup`: the in-band queued follow-up
  runs the same inbound prep but passed no persist override; it now carries
  the authored text too.

Live: real inbound prep → persist seam → SessionStore.append_to_transcript on
a temp state.db. Before: content='[Triggering message id: `1550…` — use as
`message_id` …]\n\nCreate a project plan for Q4'. After: content='Create a
project plan for Q4', api_content=<wrapped bytes>; reply-pointer and
no-message_id (desktop relay) controls unchanged.

Slimmer redo of #71309 by @JonthanaHanh (regex strip on a `gateway/run.py`
site that no longer exists; tests exercised a copied regex); #71619
(@calvinnwq, +1303/-40 over 10 files) is over the salvage bar.

Fixes #71304
Fixes #114719
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-09-18 09:53:09 -07:00
teknium1
9f5b7ea02e fix(compression): keep the aux ceiling and re-probe feasibility on every main-runtime change
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).

- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
  _apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
  only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
  every runtime change: switch_model (outside the rollback guard), fallback activation
  and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
  fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
  rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:49:14 -07:00
teknium1
d7865939af fix(tts): xAI TTS availability probe prefers XAI_API_KEY like the synthesis paths
`tools/tts_tool.py::_xai_requirements` (the provider check text_to_speech_tool
dispatch consults) still resolved xAI credentials OAuth-first, while both TTS
synthesis paths now pass `prefer_api_key=True`. Truthiness is unchanged, but a
configured key no longer routes the availability probe through the OAuth pool
(pool select / refresh) for a bearer the metered `/v1/tts` endpoint rejects.

Docs: the streaming-TTS provider table now states the ordering.

Part of the #113727 salvage of #113728 (@beardthelion); sibling site the PR
missed.

Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
2026-09-18 09:47:06 -07:00
teknium1
e81d4b8d96 docs(relay): scale-to-zero obligation 7 — relay-fronted only, re-checked at suspend time 2026-09-18 09:34:42 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
Yagna Vudathu
87cd8a3c84 fix(gateway): /stop, /new and /reset keep a parked internal wake instead of discarding it
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).

Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).

Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 09:27:02 -07:00
teknium1
5811653914 fix(api): a run admitted but not yet started settles interrupted at shutdown
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.

Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.

Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
2026-09-18 09:24:21 -07:00
teknium1
3e408dcc2a fix(gateway): Telegram topic-binding heal cannot re-pin a route /new moved during its lookups
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
2026-09-18 09:20:39 -07:00
teknium1
d177b119e9 docs: setup UX a standalone memory provider keeps; catalog migration note for users 2026-09-17 20:35:22 -07:00
Victor Kyriazakos
541b20290c fix(bedrock): restore Grok context with provider-confirmed cache provenance 2026-09-17 18:36:10 -07:00