The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.
- agent/background_review.py: extract the review-fork construction into
build_cache_parity_fork() — same runtime/credentials as the parent,
byte-identical system prompt / tools[] / reasoning config on the
same-model path, shared session_id for prefix warmth, full persistence
detachment (no state.db writes, no rotation, no external memory,
in-place-only compaction). The review thread now calls the helper;
behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
AIAgent exists — replays the untruncated snapshot as warm cache reads,
denies every tool at dispatch via an empty thread whitelist (tools[]
stays byte-identical for cache parity), attributes usage to the parent,
and trims a mid-turn snapshot tail so role alternation holds. The
one-shot digest remains as fallback (no live agent = cold cache anyway,
and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
the chat's cached agent (parity with how turns reuse it).
Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).