Commit Graph

390 Commits

Author SHA1 Message Date
teknium1
3f59b5c594 refactor(agent): /context breakdown and native-compaction retention use the canonical token estimator
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
2026-09-13 05:09:43 -07:00
Teknium
fc3d60af09 feat(compression): provider-scoped model_thresholds keys ("<provider>:<substr>")
A bare `astra: 0.85` in compression.model_thresholds was written for the Codex
OAuth route, where Astra is capped at 272K and 50% would compact at ~136K. The
key is substring-matched on the model name alone, so it also fired on
openai/gpt-6-astra via OpenRouter and Nous, where the window is 1.1M: the user's
0.5 global threshold was silently replaced by 0.85 and the session sat at 620K
(~59%) without compacting.

Keys may now carry a provider prefix: `"openai-codex:astra": 0.85` applies only
when the session's provider is openai-codex; bare keys keep their route-agnostic
behaviour. Ranking is by model-substring length with scope as the tie-break, so
`astra-900k` still outranks `openai-codex:astra` for the 900K picker. The
provider flows through ContextCompressor (ctor + update_model), the ContextEngine
base class and the TUI hot-reload path, so a /model switch between routes
re-scopes the override.
2026-09-11 02:07:48 -07:00
gaoanze888
8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
kshitijk4poor
26f4a674e0 fix(agent): a /steer row is human input for every user-turn predicate
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.

Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
2026-09-09 13:08:25 +05:30
Teknium
e9313f6458 fix: let provider evidence adjudicate past-window preflight estimates 2026-09-07 14:11:41 -07:00
kshitijk4poor
ed063034ad test(compressor): pin custom_providers threading at the get_model_context_length seam
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
2026-09-08 02:27:01 +05:30
dhruv kejriwal
71516214c3 fix(compressor): thread custom_providers into context-length resolution
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
2026-09-08 02:27:01 +05:30
Teknium
77086b1439 refactor: isolate compression summary dispatch 2026-09-07 05:59:08 -07:00
Teknium
be58c276ee feat(compression): per-image token cost learned from the provider's own usage (#70328, supersedes #70463)
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).

The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.

- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
  anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
  ~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
  read the same bound value, so trigger and walk agree; the per-message memo now caches text
  tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.

evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.

Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
2026-09-06 14:19:42 -07:00
Joey
820106d4a5 fix(compaction): count api_content in tail budget 2026-09-06 14:04:05 -07:00
Teknium
562e6e4824 fix(compression): opaque encrypted_content costs 0 in every local token estimate
Codex Responses reasoning and compaction items carry ciphertext the provider prices by its own
token count, never by bytes; a single native compaction checkpoint is ~5M chars, which the
bytes/4 estimator turned into ~1.29M "tokens" against a 204K threshold (#100611). #104192 deferred
that decision for one request; this removes the mis-pricing at the source so the preflight
estimator and the tail-budget walk agree (a mismatched size class protects blob-heavy rows as
"small" and compaction re-fires). Only real usage prices these items, and the usage anchor carries
that price forward.

evals/native_compaction/ab_checkpoint_preflight.py: preflight estimate after checkpoint
1,292,413 -> 58 rough tokens; the over-threshold negative arm still compresses.
2026-09-06 13:21:17 -07:00
Teknium
0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
Benjamin Brumbaugh
cd71ee0708 fix(compression): defer local preflight after native checkpoint
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.

Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.

Fixes #100611
2026-09-06 09:09:00 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
kshitijk4poor
c06f7cf924 test(compression): pin rearm-reset wiring through a public boundary
Post-review cleanup on the salvage stack:
- The re-warn-after-reset guard now drives on_session_reset() instead of
  the private helper, so a site regressing to a bare rearm-zero fails it.
- One _over_threshold_warnings() helper replaces four inline caplog filters.
- Comment the redundant None check that narrows current_tokens for ty.
2026-09-03 12:26:22 +05:30
kshitijk4poor
58a4a11727 fix(compression): release the reclamation no-op dedup key on every rearm reset
Follow-up to the salvaged #101894 (@jwilson411). The over-threshold
"reclamation did not run" warning is deduped on (reason, rearm mark), and
the key was only cleared when a prune committed. Every other path that
zeroes the rearm mark — compress(), on_session_reset/on_session_end,
bind_session_state, update_model — left the key in place, so a lockout
that warned at rearm=0, then a full compaction, then the same lockout
again was silent, contradicting the helper's own "warns again" contract
(and leaking the key across sessions on a rebound compressor).

- ContextCompressor._reset_proactive_prune_rearm(): one helper for the
  five rearm-to-zero sites; clears the dedup key alongside the mark.
- _warn_reclamation_no_op(): dropping back under threshold releases the
  key (mirrors _clear_context_overflow_warn semantics on the agent side).
- test_proactive_prune_loop_wiring: the attempts_exhausted fixture now
  models the only state the real engine can produce for that branch
  (should_compress() is should_compress_info()[0]) — budget spent
  (max_compression_attempts=0) with the engine saying RUN, instead of
  should_compress=False paired with (True, None).
- Two guards: lockout warns again after a rearm reset; dropping under
  threshold releases the key. Both fail with the clears removed.
2026-09-03 12:23:55 +05:30
Justin Wilson
9f12121206 fix(compression): do not let prune rearm lock out over-threshold sessions
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.

Fixes #101889
2026-09-03 12:23:55 +05:30
Teknium
0d27aaa247 refactor(agent/context_compressor): fold cooldown/pressure/handoff-strip branches; drop lazy-import blank lines 2026-09-02 23:19:04 -07:00
Teknium
e43eb87b4e refactor(agent/context_compressor): template-driven tool-result summarizers; compact docstrings/comments (every rule kept) 2026-09-02 23:16:11 -07:00
Teknium
72bbbf8f75 refactor(agent/context_compressor): fold gate/backoff loops, content-shape helpers, echo/bridge index math 2026-09-02 23:09:25 -07:00
Teknium
2539771553 refactor(agent/context_compressor): AST-identical bracket hug pass 2026-09-02 23:01:55 -07:00
Teknium
ef34ee566b refactor(agent/context_compressor): extract compress() summary/assembly phases; flatten boundary, handoff-strip, and placement helpers 2026-09-02 22:58:16 -07:00
kshitijk4poor
d6de054bac refactor(compressor): one helper for the last template-visible role
The re-append slot check duplicated the summary-role scan at the head/summary
boundary; both now call _last_template_visible_role.
2026-09-03 11:27:40 +05:30
kshitijk4poor
ab73aeff05 fix(compressor): keep pending tool_calls, avoid double anchors, no header stacking
Review follow-ups on the in-flight replay:

- Run _sanitize_tool_pairs BEFORE the re-append. Its trailing-in-flight
  exemption (#79278) walks back from the list end; with the replay user row
  there, a genuinely pending assistant(tool_calls) looked orphaned and had its
  calls stripped, so the executor's late tool result was dropped.
- When the restatement is merged onto a user-pinned summary carrier, flag the
  carrier (_inflight_replay_merged). The carrier's metadata marks it synthetic,
  so conversation_compression._ensure_compressed_has_user_turn inserted a second
  copy of the same request; it now treats the flag as intent-present, and the
  next cycle recognises the carrier as the task instead of losing it.
- Header idempotency: restate the text after the last header so a task that
  survives several compactions carries one header and one copy.
- Exclude metadata-flagged scaffolding (_todo_snapshot_synthetic, recovery
  nudges) from the in-flight scan via the shared _is_real_user_message.

Tests cover all four; each fails with its fix reverted.
2026-09-03 11:27:40 +05:30
kshitijk4poor
1d260f7683 fix(compressor): judge the re-append slot on template-visible roles
The in-flight replay checked only compressed[-1] before choosing between
appending a user row and merging onto the summary carrier. A user-pinned
summary followed by an exempt tool_calls/tool tail therefore gained a second
visible user turn and broke the Mistral-style alternation pre-flight
(tests/agent/test_summary_role_template_alternation.py::test_zero_user_guard_still_forces_user).
Look through the exempt tail to the last template-visible role instead.
2026-09-03 11:27:40 +05:30
Justin Wilson
0ed3f9b45b fix(cron): re-append in-flight job prompt after compaction
Compaction of a single-prompt (cron) session left no user message after
the handoff, so SUMMARY_PREFIX ordered the model to do nothing and the
scheduler recorded success. Re-append the in-flight task after the
summary.

Fixes #100818

(cherry picked from commit c0ea50ac02fd7553b1ef7deb0162483cb9a1b6ac)
2026-09-03 11:27:40 +05:30
Teknium
45a4672428 refactor(agent/context_compressor): split _build_summary_prompt into template/temporal helpers (prompt text byte-identical) 2026-09-02 22:35:45 -07:00
Teknium
29aa1bdf3d refactor(agent/context_compressor): share micro-cursor reset, breaker predicate, real-usage verdict; collapse session-state boilerplate 2026-09-02 22:34:33 -07:00
Teknium
081484e259 refactor(agent/context_compressor): unify tool-call field access, JSON-arg parsing, content rewrite; compact module helpers 2026-09-02 22:27:25 -07:00
Teknium
c14ae92fea refactor(agent/context_compressor): fold derived-budget cache resets 2026-09-02 20:44:33 -07:00
Teknium
bed463381f refactor(agent/context_compressor): class-qualify helper calls in _clear_compression_failure_cooldown (bare-stub binding) 2026-09-02 20:31:31 -07:00
Teknium
71bfa94ff9 refactor(agent/context_compressor): use the failure-kind dataclass directly; compact threshold/arg helpers 2026-09-02 20:23:47 -07:00
Teknium
932f7532ca refactor(agent/context_compressor): third AST-identical reflow pass at 118 cols 2026-09-02 20:19:10 -07:00
Teknium
5f7284e86e refactor(agent/context_compressor): repack historical-prefix and abort-message literals (values byte-identical) 2026-09-02 20:15:48 -07:00
Teknium
684fa87ad4 refactor(agent/context_compressor): fold parent-row rotation reads 2026-09-02 20:10:37 -07:00
Teknium
91cf271093 refactor(agent/context_compressor): second AST-identical reflow pass 2026-09-02 20:08:10 -07:00
Teknium
322f04f8fb refactor(agent/context_compressor): reflow docstrings/comments to width (every word preserved) 2026-09-02 20:04:00 -07:00
Teknium
321e8a79b0 refactor(agent/context_compressor): one newest-first token walk shared by tail-cut and prune boundary 2026-09-02 19:53:45 -07:00
Teknium
d3b93c45c9 refactor(agent/context_compressor): flatten historical-media strip, summarizer part/tool-call rendering, orphan sanitizer 2026-09-02 19:49:48 -07:00
Teknium
adbac330c7 refactor(agent/context_compressor): repack implicitly-concatenated string literals (values byte-identical) 2026-09-02 19:46:07 -07:00
Teknium
7b8d547ffc refactor(agent/context_compressor): normalize blank lines (ruff E302/E303/E305) 2026-09-02 19:43:18 -07:00
Teknium
939305f186 refactor(agent/context_compressor): pack multi-line def signatures (AST-identical) 2026-09-02 19:42:06 -07:00
Teknium
b233420b3d refactor(agent/context_compressor): shared real-user index helper, folded guard ladders, telemetry loop cleanup 2026-09-02 19:39:24 -07:00
Teknium
1d7fb737cd refactor(agent/context_compressor): prompt scaffolding helpers, shared real-usage reset, compact text/length helpers 2026-09-02 19:25:44 -07:00
Teknium
5f716407b9 refactor(agent/context_compressor): move per-section summarizer instructions into a keyed table (text byte-identical) 2026-09-02 19:21:58 -07:00
Teknium
a7e616ee5f refactor(agent/context_compressor): flatten _scan_window_handoffs around an early return 2026-09-02 19:17:16 -07:00
Teknium
e468577b13 refactor(agent/context_compressor): shared content-part text helpers for handoff projection/unwrapping 2026-09-02 19:13:23 -07:00