Commit Graph

454 Commits

Author SHA1 Message Date
teknium1
607dff390a feat(telemetry): v5 efficiency metrics — turn cost, wasted tokens, tool output truncation, tool overhead, prompt-cache breaks
Six opt-in, bucketed shared-metrics counters (schema + contract + docs in `v5 efficiency` blocks):

- hermes.task_cost.count {provider, model, tokens_bucket, tool_calls_bucket, api_calls_bucket,
  outcome}: one row per user turn the user saw end (pre_llm_call-started, attended; forks, delegated
  children and session-close aborts excluded). Tokens = prompt+completion over the turn's primary
  calls. Tool/API call buckets reach gte_101 (TURN_ACTIVITY_BUCKETS; COUNT_BUCKETS untouched).
- hermes.wasted_tokens.count {provider, model, reason in undo/retry/interrupt, tokens_bucket}:
  emitted from the v4 friction sites (no re-detection). /undo N counts N turns; an interrupted turn
  later undone counts once; turns this process never saw read unknown.
- hermes.tool_output_truncation.count {tool, truncated, original_size_bucket}: one row per tool
  result, judged after the per-turn budget; covers the per-result spill, the turn budget, and the
  shared head/tail notice tools write when they cut their own output (original size reported).
- hermes.tool_overhead.count {enabled_tool_count_bucket, tool_schema_tokens_bucket,
  execution_surface} + hermes.tool_enabled_unused.count {toolset (shipped TOOLSETS else custom),
  used}: once per closed interactive conversation (merged across compression lineage).
- hermes.cache_break.count {provider, model, cause}: compression (committed), model_switch /
  system_prompt_rebuild (continuing conversation rebuilt its prompt), toolset_change (tool array
  changed mid-conversation, Bot Chat capability rebuild), provider_reported_miss (cold read after a
  warm read on the same route with no Hermes-known cause), cache_expired (same after >=5 min idle).
  A Hermes-known break suppresses the miss it causes. No prompt hashing.

Live-proven against a fake OpenAI-compatible server (chat -q, --resume -m, TUI gateway JSON-RPC
undo/retry/interrupt); disabled => zero telemetry files.
2026-09-28 12:43:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
10708bc117 fix(bot-chat): rebuild a stale Bot Chat's tools through the surface builder (#124211, salvage #124307)
Trim the salvaged detector/delivery pair to the salvage bar and fix the
delivery half. The PR resolved the refresh selection as
platform_toolsets.<surface> (desktop/tui), a key no config path writes,
so _get_platform_tools fell back to the constructed hermes-<surface>
composite that resolves to 0 tools (a cold resume went 30 -> 1 tool).

Delivery now goes through the builder new desktop/TUI sessions use
(tui_gateway.server._load_enabled_toolsets(platform) +
_load_disabled_toolsets) as an explicit refresh_agent_mcp_tools
override, so the refreshed set equals a fresh session's on that surface;
only desktop/tui agents are long-lived, every other surface builds a
fresh agent per process/turn and is left alone. Drops the
toolsets.py/tui_gateway policy relocation and the raw-value fallback in
the fingerprint; keeps the detector (platform_toolsets +
agent.disabled_toolsets replace the dead tools.enabled_toolsets key)
and the tools re-pin after the refresh.

Tests: 2 invariants in tests/agent/test_bot_chat_toolset_refresh.py
(desktop -> builder consulted, disabled override passed, re-pinned;
cli -> untouched) and the existing fingerprint axis test now edits
the key `hermes tools enable/disable` writes.
2026-09-28 06:06:17 -07:00
finn763
fcda937c5d fix(bot-chat): deliver toolset changes to the canonical Bot Chat
The capability epoch watched tools.enabled_toolsets, a key no surface
writes; real hermes tools enable/disable edits (platform_toolsets.* plus
agent.disabled_toolsets) never flipped it. Even when stale, the refresh
rebuilt only the prompt and never tools[]. Watch the real keys and
rebuild + re-pin the tool snapshot on Bot Chat capability refresh.
Closes #124211
2026-09-28 06:06:17 -07:00
kshitijk4poor
a4db03cee9 fix(prompt): stored Model/Provider line with an empty live value is a stale route
Route commits no longer NULL the stored prompt, so _stored_prompt_matches_runtime is the
only rebuild trigger. A model-only request that clears agent.provider skipped the compare
and replayed the previous route's prompt verbatim. The builder omits empty trailer lines,
so the rebuilt prompt matches from the next turn (no rebuild loop); prompts without
identity lines keep reusing.
2026-09-27 02:22:23 +05:30
Richard McKinney
f762b96bfd fix(agent): keep prompt-cache redecoration returns consistent
_redecorate_prompt_cache_for_provider returned a 2-tuple when
tools_for_api was None but a 3-tuple otherwise, while its only production
caller (agent/turn_api_request.py) always unpacks three values. Passing
None raised "not enough values to unpack (expected 3, got 2)". The path is
latent today (agent.tools is always a list), but the union return type
invited exactly this mismatch.

Always return (messages, prepared, planned_tools); planned_tools falls
back to agent.tools when tools_for_api is None, as it already did
internally. Update the two tests that unpacked two values and add one
parametrized invariant over None / [] / list x cache-off / native.

Salvaged from #122441, rebuilt on current main with only the redecorate
hunks: the original commit was made on a stale conversation_loop.py and
would have reverted the continuation-overlap dedupe, the legacy
length-continuation stub, the tools-available stub wording and the Bot
Chat NULL guard.
2026-09-27 00:33:45 +05:30
kshitijk4poor
b7f91258b9 fix(agent): tools-available note on dropped-tools continuation; keep legacy stub recognized (#74990)
The dropped-tools partial-stream continuation now also states the cut was a
transport interruption and tools remain available. The pre-rewording
network-stub text stays in the compressor's synthetic-turn set so
crash-persisted nudges from older sessions are not mistaken for user turns.
2026-09-24 22:27:18 +05:30
Ugo Enyioha
eaf0fd95cb fix(agent): re-assert tool availability on partial-stream continuation
Re-port of 6f01f7476b (PR #72273) onto the _LENGTH_CONTINUATION_NETWORK_STUB
constant that _get_continuation_prompt now returns: state the cut was a
transport interruption, that tools remain available, and drop 'Finish the
answer directly' which read as a text-only instruction.
2026-09-24 22:27:18 +05:30
kshitijk4poor
6979cb9bab refactor(gateway): apply shutdown-notice interim marker once in _send_notice_logged
Also type continuation parts as (text, is_stub) tuples end-to-end and bound
the overlap scan to the tail of the joined text.
2026-09-24 22:26:11 +05:30
fangliquanflq
d98deef030 fix(agent): scope continuation dedupe to stream stubs
(cherry picked from commit 82acba16e08cce9f613598325d7c566838fb002e)
2026-09-24 22:26:11 +05:30
fangliquanflq
3814aab269 fix(agent): dedupe repeated stream continuation tails
(cherry picked from commit 87c93960cd9e8dd502ffe82a10febdefa3cf64c4)
2026-09-24 22:26:11 +05:30
brooklyn!
50ff26bdc0 fix(tools): rebuild Bot Chat prompt when model capability overrides flip
capability_fingerprint ignored model.supports_vision and model.context_length,
so an eternal Bot Chat kept its stored prompt after those overrides changed.
A NULL stored prompt already rebuilds and no longer takes the stale probe.
2026-09-24 05:29:27 -05:00
teknium1
027809f88e fix(agent): don't restore a foreign Session ID or pin an empty workspace snapshot
- _stored_prompt_matches_runtime: with the Session ID trailer on
  (--pass-session-id / HERMES_TUI_PASS_SESSION_ID), a stored prompt whose
  Session ID line is not this session's is a mismatch. A /branch child now
  copies its parent's prompt bytes, and without this it told the model the
  parent's id. Gated on the flag so a "Session ID:" line in project text
  cannot force a rebuild every turn when the trailer is off.
- _seed_workspace_pin: adopt only a real workspace block. A stored prompt
  that names this cwd but carries no block (built on a messaging surface,
  resumed on the CLI in the same repo) pinned "no workspace" and dropped the
  git snapshot for the rest of the session; now the next build captures one.
2026-09-23 17:06:48 -07:00
teknium1
921ab7a163 fix(agent): keep the session-start workspace snapshot across prompt rebuilds
The workspace git snapshot is pinned per session so a rebuild (compaction,
/compress) replays it instead of re-probing a repo that moved. Two gaps let
the rebuild re-probe anyway and rewrite the system prompt mid-session:

- The pin was keyed by resolve_context_cwd(), which is None when no cwd is
  bound (CLI launch dir) and the path once the TUI /compress binds the
  session cwd: the same directory under two keys. Key by the directory the
  probe actually inspects (resolve_context_cwd() or resolve_agent_cwd()).
- An agent that did not build the session's prompt (resumed, a fresh
  gateway/TUI agent whose first act is /compress, a kill -9 restart) had
  no pin, so its first rebuild probed git now instead of replaying the
  session-start bytes. Seed the pin from the prompt the session already
  sends (cached copy, else its persisted row), only when that prompt names
  this cwd and its snapshot's Root covers it.

reset_session_state still drops the pin, so /new, /resume and /branch
re-snapshot at their own session start.
2026-09-23 17:06:48 -07:00
teknium1
3dbb417809 fix(agent): type the failed-turn boundary row so clients never read it as the model's reply
8f0322da5b closes a failed turn with a Hermes-authored assistant row
(FAILED_TURN_NOTICE / PARTIAL_FAILED_TURN_NOTICE). It carried no marker, so
the only way a client could tell it from a real answer was matching the
English copy.

Both writers (agent/conversation_loop.py::_close_durable_failed_turn and
gateway/run_turn.py::_hmwa_close_failed_turn) now stamp
display_kind="failed_turn" (agent/turn_failure_copy.py::FAILED_TURN_DISPLAY_KIND).
display_kind is a DB/display column already stripped from every provider
request (agent/turn_context.py), so the wire bytes and prompt cache are
unchanged; the ACP loopback test asserts the replayed row is exactly
{"role": "assistant", "content": FAILED_TURN_NOTICE}.

session.resume (tui_gateway/session_history.py::_legacy_display_kind) also
types untyped rows already on disk from the last five days, matching the
Python constants in-process, so clients key on the type alone.
2026-09-23 16:58:39 -07:00
teknium1
d956f0ae57 fix: key the tools[] pin by code version; never re-add config-excluded tools
Review follow-up on the byte-identical tools[] pin.

- The pin records the code identity that built it (checkout/build sha, else
  the release version). Written by the same code, every pinned tool that is
  still available keeps its pinned bytes, including tools whose parameters
  are derived per surface (delegate_task, text_to_speech, memory, patch).
  The per-tool "parameters differ -> take current" rule replaced those bytes
  on every surface hop and rewrote the ~44KB pin each time. A pin from other
  code (`hermes update`, legacy name lists) takes the current definitions
  once and is re-pinned.
- A pinned tool this process did not build is carried forward only while
  this agent's toolset selection allows it (enabled minus disabled toolsets
  and role reservations, before check_fn). It must also pass the session
  schema gates on the merged array, so browser_exec never comes back once
  terminal is gone. Client-surface toolsets (desktop_ui, project) still
  carry across hops: no config choice removed them there.
- The rotation compaction child inherits the parent's pin in the publish
  transaction.
- `hermes sessions recover` keeps pin rows in its system_prompts sweep and
  clears dangling pin hashes, as lost-and-found now does too. Profile moves
  carry the pin like the prompt. A continuing session whose pin is missing
  or unreadable (a row swept by an older build) pins the tools it sends on
  that turn, so later hops stay stable.
2026-09-23 15:43:51 -07:00
teknium1
7a31c365c3 fix: one session sends byte-identical tools[] across TUI, oneshot and gateway hops
The session tools pin (sessions.tool_names) stored names only, so every fresh
process re-materialized the bytes from its own surface and every surface hop
of one durable session was a full prompt-cache miss:

* tool_search's deferred catalog is built per process ("Search 6 additional
  tools" in the TUI gateway vs 5 in -q);
* a pinned tool missing from the fresh build (skill_manage under the -q
  footprint) came back from the static registry schema, without its
  dynamic_schema_overrides;
* a -q --resume that rebuilt the stored prompt (model switch, cwd drift)
  persisted its own pruned array over the pin.

The pin now stores the full definitions and restore replays a pinned tool that
is still available byte-for-byte (deregistered tools drop, new ones append at
the tail, legacy name-only pins still work). A continuing session whose prompt
is rebuilt applies the pin before building it, matching the freeze policy
(tools[] only changes on /new, /reload-mcp, compaction). The array is
content-addressed in the existing system_prompts store like the prompt itself,
so identical arrays across sessions are stored once; get_session resolves it.
2026-09-23 15:43:51 -07:00
teknium1
503d818ddf fix: vision_analyze does not re-embed an image already attached natively to the current turn
When a surface (Telegram gateway, CLI, TUI, delegated child) attaches an
image natively to the user turn and the model then calls vision_analyze on
the same path, the native fast path embedded the identical pixels a second
time as a multimodal tool result in the same request (#76411).

conversation_loop.run_conversation now scopes the turn's native image
handles (the [Image attached at: ...] hints build_native_content_parts
writes) in a ContextVar for the duration of the turn; the tool answers with
a short text result instead. Region crops and other images still embed;
the scope ends with the turn so later turns can re-load after compression.
2026-09-19 09:52:18 -07:00
kingrubic
8535b7f196 fix(agent): fail over from codex app-server quota errors (#71642 re-port)
Re-applies kingrubic's #71642 logic on the _LoopState loop shape: classify the
codex app-server result["error"] text after the turn, guard on
_has_pending_fallback(), fail over on billing / rate_limit / upstream_rate_limit,
carry the failed API call in api_call_count and resync the failover system
message before retrying the same user turn on the generic loop.

Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
2026-09-19 09:29:49 -07:00
teknium1
f7d1965f78 fix: unwind re-port of #71642 for attribution
Temporarily removes the hunks that 08ad795b9db carried over line-for-line from
kingrubic's PR #71642 (codex_result error + _has_pending_fallback guard,
classify_api_error(RuntimeError(str(err)), ...), the eligible-reason frozenset,
api_call_count accounting, _sync_failover_system_message) so the next commit can
re-apply them under the contributor's authorship. Tree is intentionally
non-functional between this commit and the next.
2026-09-19 09:29:49 -07:00
teknium1
d7e405ef79 fix(agent): codex app-server quota and rate-limit failures fail over to fallback_providers
The codex_app_server dispatch returned the turn result unconditionally, so a
billing/usage-limit/rate-limit error carried as result["error"] text never reached
the classify_api_error -> _try_activate_fallback chain every other runtime uses.
Classify that text after the codex turn; on a fallback-eligible verdict activate
the configured fallback and continue the same user turn on the generic loop,
keeping the projected codex rows and the failed API call in the accounting.

Re-port of PR #71642 (@kingrubic) onto the _LoopState loop shape; the invariant
test drives the real activate -> retry path against a local fake OpenAI server.

Fixes #71633
2026-09-19 09:29:49 -07:00
kshitijk4poor
80f0d1b52f fix(agent): re-prompt once when a turn that did tool work ends on a collapsed fragment
A Responses-wire collapse (#103483): the tool work runs correctly, then the
final text stop is a fragment — a stray wrong-script word ("пар"), a token
starting mid-punctuation ("?warming up") — and the loop accepted it as the
answer, so the turn reported completed and an unattended job abandoned the
task.

The guard rides the existing ack-continuation path in finalize_turn: same
scope knob (agent.intent_ack_continuation, default auto = Responses
transports), the same bounded per-turn counter, the same durable interim +
nudge rows, and the same synthetic-user recognition in the compressor. It
fires only when the turn has tool results after the last user row and the
answer matches a deliberately narrow predicate: <= 24 chars, no sentence
terminal, and either leading punctuation no answer begins with or letters
with no ASCII letter/digit at all. "42", "SQLite", "report.csv", "€12.50",
"你好。", "Done." never match. Because shape cannot prove a collapse, the
nudge asks for the same answer again when it WAS complete, so a false
positive costs one call, never the answer. The nudge row closes the
tool-work window, so a second fragment ends the turn as before.

Rebuilt from PR #111472 by dankkush against current main; the mid-task
"stall note" arm and the ephemeral-row popping are not taken (the former
false-positives on declarative answers, the latter buries flagged rows once
a recovery response calls a tool).

Co-authored-by: dankkush <brandan@wrengineers.com>
2026-09-19 12:12:04 +05:30
teknium1
8f0322da5b fix(agent): close a failed turn's durable user tail so the next prompt is not merged into it
Terminal-failure paths (HTTP-200 content-policy refusal, ``_Trunc.end_turn``, retry
exhaustion, interrupt before any assistant text) persist the accepted user row and return
before ``finalize_turn``, so ``user`` stays the durable conversation tail. The next prompt
appends a second user row, ``repair_message_sequence`` merges the pair, and the provider
is asked to act on the failed request again. The gateway compensates with
``_hmwa_close_failed_turn`` (#108033); standalone ACP, the CLI and the TUI/Desktop hand
``result["messages"]`` straight back as history and had no closer.

Close it once, at ``agent/conversation_loop.py::run_conversation`` — the seam every
envelope leaves through — with a Hermes-authored assistant boundary
(``agent/turn_failure_copy.py::FAILED_TURN_NOTICE`` / ``PARTIAL_FAILED_TURN_NOTICE``,
which the gateway now aliases instead of keeping its own copy). Idempotence is keyed on
``SessionDB.latest_conversation_role`` (durable state, not content), so a redelivery or a
tail another writer already closed is a no-op and the gateway's closer no-ops in turn.
The context-pressure classes (``compression_exhausted``, ``compression_deferred``,
``failure_reason == "context_overflow"``) are excluded: appending to an oversized session
is the #1630 growth loop; their repair is rotation.

Adjacent defect from the same report: ``acp_adapter/server.py::_finish_turn`` called
``final_response.startswith`` on ``None`` for an interrupted turn — the same one-line fix
PR #64471 by @israellot filed first (its wider prompt()-restructure is superseded by the
current ``_finish_turn`` shape).

Slimmer redo of #114168 by @kendrickkester (same seam and invariants; the +1023-line
PR carried a new copy module, an accepted-turn re-anchoring scan and an 859-line suite).
Two invariant tests: the real ACP path (loopback provider, refusal then a new prompt) and
the durable-tail idempotence / overflow exclusion.

Co-authored-by: Kendrick Kester <kendrick.kester@gmail.com>
Co-authored-by: Israel Lot <israel.lot@gmail.com>
2026-09-18 10:33:21 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
55a26c9983 fix: only drop an interrupted partial for the runaway repetition shape
is_repetition_dominated is tuned for the truncated-continuation nudge, where a
false positive merely skips a continuation. At the interrupt checkpoints the
same verdict DROPS the partial from history and tells the model the reply
degenerated, so a legitimately repetitive but correct reply - twelve distinct
INSERT rows sharing a long prefix trip the 60-char window scan - was erased and
mislabelled. Gate the two checkpoints on is_runaway_repetition: dominated AND,
when the text has line structure, at most half of its non-empty lines distinct.
Byte-identical repeated lines (the #112764 shape) still qualify; distinct batch
rows no longer do. The continuation path keeps the looser predicate.

The plain interrupt site now mirrors the redirect placeholder (empty content,
display_kind=hidden, api_content=[response interrupted]) so the bracketed
placeholder no longer surfaces as an assistant bubble in transcript replays.
2026-09-16 17:10:13 -07:00
teknium1
2f3c3cfbc0 fix: keep looped partials out of every interrupt checkpoint, tell the model why
Follow-up to the #112774 salvage (@KoNit-K): the redirect path now drops a
repetition-dominated partial from the replayed correction. This commit
finishes the class:

- The correction carries a one-line descriptive checkpoint
  (REPETITION_LOOP_INTERRUPTED, agent/repetition_guard.py) instead of silently
  omitting the partial, so the model knows the reply degenerated and was cut
  off without seeing the bytes that re-seed the loop. The alternation
  placeholder takes the existing hidden shape (no bytes, neutral api_content).
- Sibling site agent/turn_api_call.py::handle_api_interrupt (plain Esc/stop
  interrupt, no redirect) appended the looped partial verbatim as the
  interrupted assistant row, which is replayed on the next turn exactly like
  the redirect's api_content. It now keeps the sanitizer's neutral
  "[response interrupted]" row and returns the descriptive line as the
  user-visible final response.
- Tests: the salvaged redirect test also pins the placeholder; one new
  invariant pair for handle_api_interrupt (looped vs ordinary partial).

Why: a partial that is >50% one repeated fragment is the signature of a
repetition loop; replaying it makes the model re-latch and the corruption is
persisted in the conversation (survives a gateway restart, #112764). The
truncated-continuation path already applied this guard; the two interrupt
paths did not.
2026-09-16 17:10:13 -07:00
KoNit-K
3493695092 fix(agent): skip repeated interrupt partial replay 2026-09-16 17:10:13 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Siddharth Balyan
cbcf7b72f7 feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop

* feat(gateway): /signin signs the free tier into a Nous account from a DM

* feat(cli): chat surfaces name /signin as the sign-in verb

* fix(auth): review follow-ups for the shared sign-in flow and /signin

* fix(i18n): carry the /status free-tier line in every locale catalog

* refactor(cli): the chat sign-in command is /login

* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
2026-09-11 03:45:33 +05:30
Teknium
94f77dfa0d fix(agent): thread turn_author through conversation_loop.run_conversation
The facade forwarded turn_author= to the loop's public entry point, which did not
declare it: every real AIAgent.run_conversation() turn raised TypeError (16 CI
failures across provider, sidecar, cron and finite-chat suites). The PR's tests
only exercised build_turn_context directly, so the missing hop was invisible.
Adds one facade-through-loop test that goes red when the kwarg is dropped.
2026-09-10 10:27:07 -07:00
Erosika
70b1ff6930 feat(memory): carry the turn's author into the memory-provider contract
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include:
remaining_tokens, model, platform, tool_count" — and `MemoryManager`
forwards whatever it receives. Its only caller passed nothing, so a memory
provider had no way to learn who wrote the turn it was being told about.

Providers that key durable state on identity resolve one identity when the
session is created. A shared session does not work that way: threads are
shared by default (`thread_sessions_per_user` is False), so alice, bob, and
another agent all write turns into a session whose peer is whoever spoke
first. The gateway's answer today is the `[name]` prefix it prepends to the
message text, which the model reads and a provider cannot.

`turn_author` now travels from the gateway through `run_conversation` into
`build_turn_context`, which forwards `author_id`, `author_name`, and
`author_is_bot` to every provider. It stops there — the trio never reaches
the model, and providers that ignore the kwargs are unaffected.

The bot flag is sent on every transport, not only shared sessions: a
provider deciding whether a turn may write to durable memory needs it in a
DM too.

`SessionSource.is_bot` is only as good as its producers. `build_source`
defaults it to False and 3 of 32 adapter call sites pass it, so most
platforms still report every author as human. Populating the rest is
follow-up work; nothing here depends on the flag being right yet.
2026-09-10 10:27:07 -07:00
fangliquanflq
feb03196a2 fix(agent): isolate detached forks from lifecycle hooks 2026-09-10 13:23:00 +02:00
yoyodine-industries
e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
nftpoetrist
511633be90 fix(agent): stop the run-budget wrap-up notice from mutating a persisted tool row
_maybe_inject_run_budget_wrapup() appends its wrap-up notice to the newest
role:"tool" message in place, with no _DB_PERSISTED_MARKER check. Its sibling,
_maybe_inject_iteration_budget_warning(), got exactly this guard added in the
same recent saga (turn_iteration_prep.py), with the comment "an older turn may
already be cached."

The reachability is structural, not an edge case: _maybe_inject_run_budget_wrapup
is only ever called from prepare_iteration(), at the START of the next iteration
-- strictly after tool_executor.py's _flush_session_db_after_tool_progress has
already flushed and marked the previous iteration's tool row persisted. So every
successful injection was mutating an already-persisted row: the wire request for
that turn carried the notice, but the durable transcript never did, diverging
replay from the live bytes and invalidating the provider's prompt-cache prefix
from that row onward.

Fix:
- Add the same _DB_PERSISTED_MARKER guard to _maybe_inject_run_budget_wrapup,
  scoped to the specific tool row the reversed scan lands on (not just
  messages[-1], since this function -- unlike its sibling -- scans backward for
  the newest tool row rather than only checking the tail).
- Wire _maybe_inject_run_budget_wrapup into _flush_session_db_after_tool_progress
  (pre-flush), mirroring exactly how _maybe_inject_iteration_budget_warning is
  wired in both places. Without this, the guard alone would make the notice stop
  firing in the common case, since prepare_iteration's call site almost always
  hits an already-persisted row -- the pre-flush call site is what actually lets
  it land in durable bytes.

Verified empirically: read the real call graph (tool_executor.py's three
_flush_session_db_after_tool_progress call sites cover every tool-completion
path) to confirm the guard's premise, then added an end-to-end test using a real
AIAgent + SessionDB that flushes and checks the persisted row for the notice
text. Mutation-verified: reverting the two production files drops exactly the 2
new/updated assertions (28 pass, 2 fail); reapplying restores green (30 passed).
Also ran the sibling iteration-budget-warning and /steer suites (71 passed) to
check for interaction regressions -- none.
2026-09-09 17:04:30 +05:30
Felipe Portavales
37f42713ef feat(loop): export {turn_id, current_turn_user_idx} on every result envelope
Hosts that settle their own transcript by index (hermes-webui) cannot prove which
row of result["messages"] is the current user turn once this loop rewrote history
(alternation repair, compaction, post-turn micro-compaction): the instance-side
_persist_user_message_idx predates those rewrites, and a text match relabels an
identical historical prompt and claims its old answer. Only the producer can
assert the coordinate against the exact list it returns.

run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair
through export_current_turn_boundary on every envelope that leaves the loop
(success, partial/error, interrupt, retry-exhausted, tool-limit, preflight
timeout, codex runtime), computed on the final messages after finalize_turn and
micro-compaction. The pair is exported only when the addressed row is this turn's
user message verbatim (reanchor's last-match rule); a rewritten row exports
nothing so hosts fail closed. The final index is mirrored into
_persist_user_message_idx for the persist override.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
kshitijk4poor
defdf64790 simplify(agent): surface switch — reuse flatten_message_text / agent_tool_names / one runtime-boundary split
- _transcript_row_texts re-implemented agent.message_content.flatten_message_text
  and the api_content sidecar rule; the note can only land on a user row,
  so the transcript scan now skips assistant/tool rows (the bulk of the bytes).
- Three sites computed "names of agent.tools"; tools.mcp_tool_agent gains
  agent_tool_names() used by the switch note and conversation_loop, which
  also stops importing the private _def_name across modules. The name list
  is only captured when a switch was announced.
- split_runtime_boundary() is the single owner of the runtime-block
  rpartition/END check for both identity_line_value and
  _stored_prompt_matches_runtime.
- platform_surface_hint was a public alias of _platform_hint; the function is
  now platform_hint (its docstring pointed at the pre-move module).
- consume_gateway_turn_context_notes and consume_surface_switch_note share
  _pop_turn_note so the two one-shot channels have identical semantics.
- platform check hoisted above the transcript scan.
2026-09-09 10:31:26 +05:30
kshitijk4poor
6cc177a76c refactor(agent): surface-switch note lives in its own sibling; skip it where no sidecar exists
Move the six surface-switch helpers out of the conversation_loop facade
into agent/surface_switch.py (AGENTS.md: new behaviour goes in a topical
sibling), and fold the review findings on #104494:

- MoA and codex_app_server turns never stamp the api_content sidecar, so
  the staged note could not be read back from the transcript and was
  re-sent on every turn after a switch. Those modes now skip the note
  (stored prompt still reused).
- The announced surface was parsed with split(".") — a plugin platform
  with a dot in its name would never compare equal and re-stage the note
  every turn. The note now closes the name with a fixed terminator.
- One identity-line parser (identity_line_value) shared by
  _stored_prompt_matches_runtime and the switch detector instead of two
  copies of the runtime-boundary/rpartition logic; tool names via the
  existing tools.mcp_tool_agent._def_name; the transcript scan is bounded
  to the last 200 rows (it ran every turn over the whole history).
- consume_surface_switch_note reduced to a plain pop; developer-guide
  prompt-assembly.md updated (Platform is no longer an identity field);
  17 new tests trimmed to 10 (same-shape pin/retire variants folded).

Restoring Platform as an identity field still turns 5 tests red.
2026-09-09 10:31:26 +05:30
joaomarcos
4e7a49d182 fix(agent): retire stale surface notes on bot-chat refresh, isolate platform from decoys
When Bot Chat capability refresh rebuilds the system prompt for the current
surface, call _stage_surface_switch_note() so any earlier switch note sitting in
the transcript is retired instead of overriding the rebuilt prompt.

Also isolate _stored_prompt_platform() to parse only the authoritative identity
portion before '# Hermes runtime environment' (with legacy fallback for prompts
without the boundary), preventing embedder prose or HERMES_ENVIRONMENT_HINT decoys
from shadowing the real platform and falsely suppressing surface switch announcements.

Credit to @ehz0ah, who identified both correctness gaps on current main and
verified the regression scenarios.
2026-09-09 10:31:26 +05:30
joaomarcos
4a96311503 fix(agent): hold the tools pin through a surface switch, name what it carried
The announcing turn used to skip the tools freeze and re-persist the array the new
surface had just built. That is the one mutation this fix cannot afford: tools[] is
serialized ahead of the system prompt, so rebuilding it moves the request at token 0
and re-prefills everything behind it — the exact cost #104414 measured (1% cache hit
on a 220K session), spent on the very turn the fix exists to make cheap. On a
`desktop -> tui` switch with a configured toolset selection (`_gui_surface_toolsets`
gives desktop `desktop_ui`, the TUI nothing), skipping the pin dropped ~a dozen tools
and bought back the whole miss.

The pin now holds. `_merge_preserving_prefix` still appends what the new surface
brought, so a `tui -> desktop` switch pays a break no freeze could have avoided, and
the tools it carries FORWARD are named at the end of the surface note instead of being
silently advertised: a `focus_pane` a terminal turn can only answer with
`tool_error("desktop only")` now reads as unavailable rather than as live capability.
The toolset converges at the next real rebuild boundary, where the break is already
paid.

Credit to @StanleyStetson, who caught that the tool array is evaluated ahead of the
system prompt and that the bypass reintroduced the miss this PR is about.
2026-09-09 10:31:26 +05:30
joaomarcos
a020050c8c fix(agent): the surface note must not outlive its own truth
Two holes the first cut left open, both created by the note itself.

Switching BACK to the surface the prompt was built for (desktop -> tui -> desktop) left the
`Platform:` trailer agreeing with the runtime, so nothing was staged — while the newest note
in the transcript still told the model it was on tui. And a rebuild for an unrelated reason
(a model switch) refreshed the prompt but not that note, leaving the same contradiction from
the other side.

Compare the runtime surface against what the model was last TOLD — the newest surface note
when one exists, else the prompt's own trailer — and stage from the rebuild path too. The
full surface guidance rides along only when the prompt itself is out of date; when the prompt
already describes the current surface the note just retires the stale one and points at it.
2026-09-09 10:31:26 +05:30
joaomarcos
aa40dbe765 fix(agent): skip the tools freeze once on a surface switch, not for the session
The first cut gated the saved tool_names pin on "the surface drifted", which stays true
for as long as the stored prompt names the old surface — i.e. until the next compaction.
On the gateway path, where a fresh AIAgent is built per turn, that left the tools freeze
off for every remaining turn, so a check_fn that flaps could reorder `tools[]` and break
the tool cache block on its own.

Gate it on the turn that actually ANNOUNCES the switch instead, and persist the fresh,
toolset-correct names there. The next turn's row already holds this surface's tools, so
the pin resumes immediately: skipped once, not disabled.
2026-09-09 10:31:26 +05:30
joaomarcos
80d6bda144 fix(agent): a surface switch must not re-prefill the whole request (#104414)
`_stored_prompt_matches_runtime` treated `Platform` as a runtime-identity field, so
answering a live session from another surface — desktop -> TUI, or a resume after a
dashboard restart whose chat is a PTY TUI child — declared the stored prompt stale and
rebuilt it. The system prompt is the first thing in the request, so changing any byte of
it moves the first divergent byte to the head of a 220K-token request and the entire
conversation behind it re-prefills: a session that was hitting 240000/240287 came back
at 1536/219861.

The guard was not wrong about correctness — a desktop-built prompt on a terminal session
advertises inline widgets and a MEDIA: channel the TUI does not have — but the surface is
advisory metadata about the renderer, not a cache domain. Model/provider and cwd drift
change what the prompt should SAY; the surface changes only one paragraph.

Reuse the stored bytes across a surface switch and correct the paragraph where it costs
nothing to cache: `_stage_surface_switch_note` stages a one-shot note carrying the CURRENT
surface's guidance on the same per-turn user-message channel the gateway's must-deliver
notes use. It lands after the cached prefix and is stamped into the byte-stable
`api_content` sidecar, so later turns replay it instead of re-prefilling, and the prompt
converges at the next compaction — a boundary that already breaks the cache.

The saved tool_names prefix is not pinned across a switch: the tool registry is
process-global, so `_merge_preserving_prefix` would carry a saved-but-unloaded tool
forward (under `coding_context: focus` desktop gets a desktop_ui toolset the TUI cannot
run). On the same surface the tools freeze is untouched.
2026-09-09 10:31:26 +05:30
Teknium
9da8df8d26 fix(prompt): preserve shared project prefixes across worktrees
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.

Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-07 04:45:48 -07:00
686f6c61
c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium
89fbd5d4d3 simplify(compat): conversation_loop — drop 9 re-exports, repoint 7 callers, 35 test sites 2026-09-03 13:12:23 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00