Commit Graph

432 Commits

Author SHA1 Message Date
teknium1
8f0322da5b fix(agent): close a failed turn's durable user tail so the next prompt is not merged into it
Terminal-failure paths (HTTP-200 content-policy refusal, ``_Trunc.end_turn``, retry
exhaustion, interrupt before any assistant text) persist the accepted user row and return
before ``finalize_turn``, so ``user`` stays the durable conversation tail. The next prompt
appends a second user row, ``repair_message_sequence`` merges the pair, and the provider
is asked to act on the failed request again. The gateway compensates with
``_hmwa_close_failed_turn`` (#108033); standalone ACP, the CLI and the TUI/Desktop hand
``result["messages"]`` straight back as history and had no closer.

Close it once, at ``agent/conversation_loop.py::run_conversation`` — the seam every
envelope leaves through — with a Hermes-authored assistant boundary
(``agent/turn_failure_copy.py::FAILED_TURN_NOTICE`` / ``PARTIAL_FAILED_TURN_NOTICE``,
which the gateway now aliases instead of keeping its own copy). Idempotence is keyed on
``SessionDB.latest_conversation_role`` (durable state, not content), so a redelivery or a
tail another writer already closed is a no-op and the gateway's closer no-ops in turn.
The context-pressure classes (``compression_exhausted``, ``compression_deferred``,
``failure_reason == "context_overflow"``) are excluded: appending to an oversized session
is the #1630 growth loop; their repair is rotation.

Adjacent defect from the same report: ``acp_adapter/server.py::_finish_turn`` called
``final_response.startswith`` on ``None`` for an interrupted turn — the same one-line fix
PR #64471 by @israellot filed first (its wider prompt()-restructure is superseded by the
current ``_finish_turn`` shape).

Slimmer redo of #114168 by @kendrickkester (same seam and invariants; the +1023-line
PR carried a new copy module, an accepted-turn re-anchoring scan and an 859-line suite).
Two invariant tests: the real ACP path (loopback provider, refusal then a new prompt) and
the durable-tail idempotence / overflow exclusion.

Co-authored-by: Kendrick Kester <kendrick.kester@gmail.com>
Co-authored-by: Israel Lot <israel.lot@gmail.com>
2026-09-18 10:33:21 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
55a26c9983 fix: only drop an interrupted partial for the runaway repetition shape
is_repetition_dominated is tuned for the truncated-continuation nudge, where a
false positive merely skips a continuation. At the interrupt checkpoints the
same verdict DROPS the partial from history and tells the model the reply
degenerated, so a legitimately repetitive but correct reply - twelve distinct
INSERT rows sharing a long prefix trip the 60-char window scan - was erased and
mislabelled. Gate the two checkpoints on is_runaway_repetition: dominated AND,
when the text has line structure, at most half of its non-empty lines distinct.
Byte-identical repeated lines (the #112764 shape) still qualify; distinct batch
rows no longer do. The continuation path keeps the looser predicate.

The plain interrupt site now mirrors the redirect placeholder (empty content,
display_kind=hidden, api_content=[response interrupted]) so the bracketed
placeholder no longer surfaces as an assistant bubble in transcript replays.
2026-09-16 17:10:13 -07:00
teknium1
2f3c3cfbc0 fix: keep looped partials out of every interrupt checkpoint, tell the model why
Follow-up to the #112774 salvage (@KoNit-K): the redirect path now drops a
repetition-dominated partial from the replayed correction. This commit
finishes the class:

- The correction carries a one-line descriptive checkpoint
  (REPETITION_LOOP_INTERRUPTED, agent/repetition_guard.py) instead of silently
  omitting the partial, so the model knows the reply degenerated and was cut
  off without seeing the bytes that re-seed the loop. The alternation
  placeholder takes the existing hidden shape (no bytes, neutral api_content).
- Sibling site agent/turn_api_call.py::handle_api_interrupt (plain Esc/stop
  interrupt, no redirect) appended the looped partial verbatim as the
  interrupted assistant row, which is replayed on the next turn exactly like
  the redirect's api_content. It now keeps the sanitizer's neutral
  "[response interrupted]" row and returns the descriptive line as the
  user-visible final response.
- Tests: the salvaged redirect test also pins the placeholder; one new
  invariant pair for handle_api_interrupt (looped vs ordinary partial).

Why: a partial that is >50% one repeated fragment is the signature of a
repetition loop; replaying it makes the model re-latch and the corruption is
persisted in the conversation (survives a gateway restart, #112764). The
truncated-continuation path already applied this guard; the two interrupt
paths did not.
2026-09-16 17:10:13 -07:00
KoNit-K
3493695092 fix(agent): skip repeated interrupt partial replay 2026-09-16 17:10:13 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Siddharth Balyan
cbcf7b72f7 feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop

* feat(gateway): /signin signs the free tier into a Nous account from a DM

* feat(cli): chat surfaces name /signin as the sign-in verb

* fix(auth): review follow-ups for the shared sign-in flow and /signin

* fix(i18n): carry the /status free-tier line in every locale catalog

* refactor(cli): the chat sign-in command is /login

* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
2026-09-11 03:45:33 +05:30
Teknium
94f77dfa0d fix(agent): thread turn_author through conversation_loop.run_conversation
The facade forwarded turn_author= to the loop's public entry point, which did not
declare it: every real AIAgent.run_conversation() turn raised TypeError (16 CI
failures across provider, sidecar, cron and finite-chat suites). The PR's tests
only exercised build_turn_context directly, so the missing hop was invisible.
Adds one facade-through-loop test that goes red when the kwarg is dropped.
2026-09-10 10:27:07 -07:00
Erosika
70b1ff6930 feat(memory): carry the turn's author into the memory-provider contract
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include:
remaining_tokens, model, platform, tool_count" — and `MemoryManager`
forwards whatever it receives. Its only caller passed nothing, so a memory
provider had no way to learn who wrote the turn it was being told about.

Providers that key durable state on identity resolve one identity when the
session is created. A shared session does not work that way: threads are
shared by default (`thread_sessions_per_user` is False), so alice, bob, and
another agent all write turns into a session whose peer is whoever spoke
first. The gateway's answer today is the `[name]` prefix it prepends to the
message text, which the model reads and a provider cannot.

`turn_author` now travels from the gateway through `run_conversation` into
`build_turn_context`, which forwards `author_id`, `author_name`, and
`author_is_bot` to every provider. It stops there — the trio never reaches
the model, and providers that ignore the kwargs are unaffected.

The bot flag is sent on every transport, not only shared sessions: a
provider deciding whether a turn may write to durable memory needs it in a
DM too.

`SessionSource.is_bot` is only as good as its producers. `build_source`
defaults it to False and 3 of 32 adapter call sites pass it, so most
platforms still report every author as human. Populating the rest is
follow-up work; nothing here depends on the flag being right yet.
2026-09-10 10:27:07 -07:00
fangliquanflq
feb03196a2 fix(agent): isolate detached forks from lifecycle hooks 2026-09-10 13:23:00 +02:00
yoyodine-industries
e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
nftpoetrist
511633be90 fix(agent): stop the run-budget wrap-up notice from mutating a persisted tool row
_maybe_inject_run_budget_wrapup() appends its wrap-up notice to the newest
role:"tool" message in place, with no _DB_PERSISTED_MARKER check. Its sibling,
_maybe_inject_iteration_budget_warning(), got exactly this guard added in the
same recent saga (turn_iteration_prep.py), with the comment "an older turn may
already be cached."

The reachability is structural, not an edge case: _maybe_inject_run_budget_wrapup
is only ever called from prepare_iteration(), at the START of the next iteration
-- strictly after tool_executor.py's _flush_session_db_after_tool_progress has
already flushed and marked the previous iteration's tool row persisted. So every
successful injection was mutating an already-persisted row: the wire request for
that turn carried the notice, but the durable transcript never did, diverging
replay from the live bytes and invalidating the provider's prompt-cache prefix
from that row onward.

Fix:
- Add the same _DB_PERSISTED_MARKER guard to _maybe_inject_run_budget_wrapup,
  scoped to the specific tool row the reversed scan lands on (not just
  messages[-1], since this function -- unlike its sibling -- scans backward for
  the newest tool row rather than only checking the tail).
- Wire _maybe_inject_run_budget_wrapup into _flush_session_db_after_tool_progress
  (pre-flush), mirroring exactly how _maybe_inject_iteration_budget_warning is
  wired in both places. Without this, the guard alone would make the notice stop
  firing in the common case, since prepare_iteration's call site almost always
  hits an already-persisted row -- the pre-flush call site is what actually lets
  it land in durable bytes.

Verified empirically: read the real call graph (tool_executor.py's three
_flush_session_db_after_tool_progress call sites cover every tool-completion
path) to confirm the guard's premise, then added an end-to-end test using a real
AIAgent + SessionDB that flushes and checks the persisted row for the notice
text. Mutation-verified: reverting the two production files drops exactly the 2
new/updated assertions (28 pass, 2 fail); reapplying restores green (30 passed).
Also ran the sibling iteration-budget-warning and /steer suites (71 passed) to
check for interaction regressions -- none.
2026-09-09 17:04:30 +05:30
Felipe Portavales
37f42713ef feat(loop): export {turn_id, current_turn_user_idx} on every result envelope
Hosts that settle their own transcript by index (hermes-webui) cannot prove which
row of result["messages"] is the current user turn once this loop rewrote history
(alternation repair, compaction, post-turn micro-compaction): the instance-side
_persist_user_message_idx predates those rewrites, and a text match relabels an
identical historical prompt and claims its old answer. Only the producer can
assert the coordinate against the exact list it returns.

run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair
through export_current_turn_boundary on every envelope that leaves the loop
(success, partial/error, interrupt, retry-exhausted, tool-limit, preflight
timeout, codex runtime), computed on the final messages after finalize_turn and
micro-compaction. The pair is exported only when the addressed row is this turn's
user message verbatim (reanchor's last-match rule); a rewritten row exports
nothing so hosts fail closed. The final index is mirrored into
_persist_user_message_idx for the persist override.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
kshitijk4poor
defdf64790 simplify(agent): surface switch — reuse flatten_message_text / agent_tool_names / one runtime-boundary split
- _transcript_row_texts re-implemented agent.message_content.flatten_message_text
  and the api_content sidecar rule; the note can only land on a user row,
  so the transcript scan now skips assistant/tool rows (the bulk of the bytes).
- Three sites computed "names of agent.tools"; tools.mcp_tool_agent gains
  agent_tool_names() used by the switch note and conversation_loop, which
  also stops importing the private _def_name across modules. The name list
  is only captured when a switch was announced.
- split_runtime_boundary() is the single owner of the runtime-block
  rpartition/END check for both identity_line_value and
  _stored_prompt_matches_runtime.
- platform_surface_hint was a public alias of _platform_hint; the function is
  now platform_hint (its docstring pointed at the pre-move module).
- consume_gateway_turn_context_notes and consume_surface_switch_note share
  _pop_turn_note so the two one-shot channels have identical semantics.
- platform check hoisted above the transcript scan.
2026-09-09 10:31:26 +05:30
kshitijk4poor
6cc177a76c refactor(agent): surface-switch note lives in its own sibling; skip it where no sidecar exists
Move the six surface-switch helpers out of the conversation_loop facade
into agent/surface_switch.py (AGENTS.md: new behaviour goes in a topical
sibling), and fold the review findings on #104494:

- MoA and codex_app_server turns never stamp the api_content sidecar, so
  the staged note could not be read back from the transcript and was
  re-sent on every turn after a switch. Those modes now skip the note
  (stored prompt still reused).
- The announced surface was parsed with split(".") — a plugin platform
  with a dot in its name would never compare equal and re-stage the note
  every turn. The note now closes the name with a fixed terminator.
- One identity-line parser (identity_line_value) shared by
  _stored_prompt_matches_runtime and the switch detector instead of two
  copies of the runtime-boundary/rpartition logic; tool names via the
  existing tools.mcp_tool_agent._def_name; the transcript scan is bounded
  to the last 200 rows (it ran every turn over the whole history).
- consume_surface_switch_note reduced to a plain pop; developer-guide
  prompt-assembly.md updated (Platform is no longer an identity field);
  17 new tests trimmed to 10 (same-shape pin/retire variants folded).

Restoring Platform as an identity field still turns 5 tests red.
2026-09-09 10:31:26 +05:30
joaomarcos
4e7a49d182 fix(agent): retire stale surface notes on bot-chat refresh, isolate platform from decoys
When Bot Chat capability refresh rebuilds the system prompt for the current
surface, call _stage_surface_switch_note() so any earlier switch note sitting in
the transcript is retired instead of overriding the rebuilt prompt.

Also isolate _stored_prompt_platform() to parse only the authoritative identity
portion before '# Hermes runtime environment' (with legacy fallback for prompts
without the boundary), preventing embedder prose or HERMES_ENVIRONMENT_HINT decoys
from shadowing the real platform and falsely suppressing surface switch announcements.

Credit to @ehz0ah, who identified both correctness gaps on current main and
verified the regression scenarios.
2026-09-09 10:31:26 +05:30
joaomarcos
4a96311503 fix(agent): hold the tools pin through a surface switch, name what it carried
The announcing turn used to skip the tools freeze and re-persist the array the new
surface had just built. That is the one mutation this fix cannot afford: tools[] is
serialized ahead of the system prompt, so rebuilding it moves the request at token 0
and re-prefills everything behind it — the exact cost #104414 measured (1% cache hit
on a 220K session), spent on the very turn the fix exists to make cheap. On a
`desktop -> tui` switch with a configured toolset selection (`_gui_surface_toolsets`
gives desktop `desktop_ui`, the TUI nothing), skipping the pin dropped ~a dozen tools
and bought back the whole miss.

The pin now holds. `_merge_preserving_prefix` still appends what the new surface
brought, so a `tui -> desktop` switch pays a break no freeze could have avoided, and
the tools it carries FORWARD are named at the end of the surface note instead of being
silently advertised: a `focus_pane` a terminal turn can only answer with
`tool_error("desktop only")` now reads as unavailable rather than as live capability.
The toolset converges at the next real rebuild boundary, where the break is already
paid.

Credit to @StanleyStetson, who caught that the tool array is evaluated ahead of the
system prompt and that the bypass reintroduced the miss this PR is about.
2026-09-09 10:31:26 +05:30
joaomarcos
a020050c8c fix(agent): the surface note must not outlive its own truth
Two holes the first cut left open, both created by the note itself.

Switching BACK to the surface the prompt was built for (desktop -> tui -> desktop) left the
`Platform:` trailer agreeing with the runtime, so nothing was staged — while the newest note
in the transcript still told the model it was on tui. And a rebuild for an unrelated reason
(a model switch) refreshed the prompt but not that note, leaving the same contradiction from
the other side.

Compare the runtime surface against what the model was last TOLD — the newest surface note
when one exists, else the prompt's own trailer — and stage from the rebuild path too. The
full surface guidance rides along only when the prompt itself is out of date; when the prompt
already describes the current surface the note just retires the stale one and points at it.
2026-09-09 10:31:26 +05:30
joaomarcos
aa40dbe765 fix(agent): skip the tools freeze once on a surface switch, not for the session
The first cut gated the saved tool_names pin on "the surface drifted", which stays true
for as long as the stored prompt names the old surface — i.e. until the next compaction.
On the gateway path, where a fresh AIAgent is built per turn, that left the tools freeze
off for every remaining turn, so a check_fn that flaps could reorder `tools[]` and break
the tool cache block on its own.

Gate it on the turn that actually ANNOUNCES the switch instead, and persist the fresh,
toolset-correct names there. The next turn's row already holds this surface's tools, so
the pin resumes immediately: skipped once, not disabled.
2026-09-09 10:31:26 +05:30
joaomarcos
80d6bda144 fix(agent): a surface switch must not re-prefill the whole request (#104414)
`_stored_prompt_matches_runtime` treated `Platform` as a runtime-identity field, so
answering a live session from another surface — desktop -> TUI, or a resume after a
dashboard restart whose chat is a PTY TUI child — declared the stored prompt stale and
rebuilt it. The system prompt is the first thing in the request, so changing any byte of
it moves the first divergent byte to the head of a 220K-token request and the entire
conversation behind it re-prefills: a session that was hitting 240000/240287 came back
at 1536/219861.

The guard was not wrong about correctness — a desktop-built prompt on a terminal session
advertises inline widgets and a MEDIA: channel the TUI does not have — but the surface is
advisory metadata about the renderer, not a cache domain. Model/provider and cwd drift
change what the prompt should SAY; the surface changes only one paragraph.

Reuse the stored bytes across a surface switch and correct the paragraph where it costs
nothing to cache: `_stage_surface_switch_note` stages a one-shot note carrying the CURRENT
surface's guidance on the same per-turn user-message channel the gateway's must-deliver
notes use. It lands after the cached prefix and is stamped into the byte-stable
`api_content` sidecar, so later turns replay it instead of re-prefilling, and the prompt
converges at the next compaction — a boundary that already breaks the cache.

The saved tool_names prefix is not pinned across a switch: the tool registry is
process-global, so `_merge_preserving_prefix` would carry a saved-but-unloaded tool
forward (under `coding_context: focus` desktop gets a desktop_ui toolset the TUI cannot
run). On the same surface the tools freeze is untouched.
2026-09-09 10:31:26 +05:30
Teknium
9da8df8d26 fix(prompt): preserve shared project prefixes across worktrees
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.

Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-07 04:45:48 -07:00
686f6c61
c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium
89fbd5d4d3 simplify(compat): conversation_loop — drop 9 re-exports, repoint 7 callers, 35 test sites 2026-09-03 13:12:23 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Justin Wilson
9f12121206 fix(compression): do not let prune rearm lock out over-threshold sessions
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.

Fixes #101889
2026-09-03 12:23:55 +05:30
Teknium
22ae371664 refactor(agent/conversation_loop): single-line helper signatures and guard expressions 2026-09-02 19:37:48 -07:00
Teknium
ac77ef36df refactor(agent/conversation_loop): seed _LoopState from TurnContext by name, compact result dict builders and guard ladders 2026-09-02 19:23:46 -07:00
Teknium
762ff705e1 refactor(agent/conversation_loop): compact billing message assembly, _LoopState comments, prompt-identity helpers 2026-09-02 19:18:17 -07:00
Teknium
f76fe598a3 refactor(agent/conversation_loop): lift the API retry loop into _run_api_retry_loop; drop blank lines after lazy imports 2026-09-02 19:11:11 -07:00
Teknium
83f8de7eab refactor(agent): flatten interrupt() claim closures, group lazy-origin imports, reflow literals/log calls byte-identically 2026-09-02 19:05:00 -07:00
Teknium
b3d75f78f3 refactor(agent/conversation_loop): compact per-turn agent resets and finalize_turn call 2026-09-02 18:57:46 -07:00
Teknium
7fa7b58b3c refactor(agent): compact conversation_loop docstrings/comments, unify partial-turn result shape, tidy empty_response_guard/interrupt_compat/iteration_budget 2026-09-02 18:53:53 -07:00
Teknium
351d0444dc refactor(agent/conversation_loop): collapse boolean ladders, unify context-engine hook gate and tool-call copy-on-write 2026-09-02 18:42:39 -07:00
Teknium
004cd68e20 refactor(agent/conversation_loop,deadline): extract bot-chat staleness + prompt persist helpers, unify guidance printing, compact deadline docstrings 2026-09-02 18:34:28 -07:00
Teknium
31025f210c refactor(agent/conversation_loop): drive phase helpers through a _LoopState + _run_phase (run_conversation 597 -> ~250 LOC) 2026-09-02 18:18:15 -07:00
Teknium
2769937936 refactor(turn): lift iteration entry/announce, Nous rate guard, API interrupt, retry-restart consumer and preflight-timeout result out of run_conversation 2026-09-02 16:37:14 -07:00
Teknium
a56f731ac6 refactor(turn): extract preflight gate + per-iteration transcript prep into agent/turn_preflight_gate.py, agent/turn_iteration_prep.py 2026-09-02 16:27:44 -07:00
Christopher
cf86c47624 fix(agent): evict stale screenshot payloads before send
Call the existing keep-newest vision retirement on the per-call
api_messages clone after sanitization so OpenAI-style tool-result
screenshots are not re-uploaded on every later turn.
2026-09-03 04:51:51 +05:30
Teknium
9780739e12 refactor(turn): extract per-iteration request assembly (api_messages/MoA/cache plan/pressure) into agent/turn_request_assembly.py 2026-09-02 16:10:07 -07:00
Teknium
9fda4e5bac refactor(turn): extract retry-loop API error handler, request build, provider call and response check into agent/turn_api_*.py + agent/turn_response_check.py 2026-09-02 16:07:51 -07:00
Teknium
0dc36e0934 refactor(turn): extract tool round + final text response branches into agent/turn_tool_round.py, agent/turn_final_response.py 2026-09-02 15:40:17 -07:00
Teknium
67ef2e50fe refactor(turn): extract response intake (normalize/hooks/scratchpad/codex-incomplete) into agent/turn_response_intake.py 2026-09-02 15:37:47 -07:00
Teknium
d42282ae9e refactor(turn): extract outer-loop exception handler into agent/turn_loop_errors.py 2026-09-02 15:29:00 -07:00
kshitijk4poor
a1d5a976b3 fix(agent): never floor an anchored pressure figure; keep the estimator total on lone surrogates
Follow-ups from review of the two salvaged #87490 commits:

- _pressure_with_real_floor now applies only on the rough fallback branch.
  A valid usage anchor is provider-exact and wins as-is: on MoA turns the
  anchor deliberately uses the pre-fold aggregator usage while
  last_real_prompt_tokens holds the folded figure, so flooring the anchored
  value would re-add fan-out tokens the anchor exists to exclude. Docstring
  rewritten to describe the real path split (anchor since d3a1c46510).
- estimate_tokens_rough: encode with errors="replace". main's estimator
  never raised; text.encode() on a lone surrogate (routine in tool output,
  see message_sanitization) raised UnicodeEncodeError and would abort a
  turn where main produced a slightly-off number.
- Record the cl100k/o200k/Qwen2.5 calibration for the bytes/4 rule.
- tests: accented Latin within +10% of the ASCII rule; mixed Cyrillic/ASCII
  counts ASCII at one byte; lone surrogates don't raise; anchored pressure
  is never floored (wiring shape).
2026-09-03 03:09:06 +05:30
Darafei Praliaskouski
73b8ec3cea fix(agent): floor pre-API compaction pressure at the last real prompt size
The chars/4 rough estimate under-counts Cyrillic and other non-ASCII scripts
by up to ~2x, so a session can ride the provider's real context ceiling while
the rough pressure stays under the compaction threshold. On providers that
silently clip over-window prompts (ollama /v1) the reactive overflow handler
never fires either, and the length-continuation retry path re-enters the API
call without passing the post-response gate — reproducing the truncation
death spiral this branch already addresses (observed live after the first
commit: real prompts 64,842 -> 64,995 against a 55,705 threshold, output
room shrinking 694 -> 541 tokens).

Floor the pre-API pressure figure at the provider's last reported
prompt_tokens — authoritative, script-independent — except for the one turn
after a compaction when that value is known-stale (#36718's
awaiting_real_usage_after_compression window).
2026-09-03 03:09:06 +05:30
Teknium
ab48b1ddc5 refactor(agent): extract per-call API message build into turn_context.build_api_messages 2026-09-02 13:30:23 -07:00