The gate now defers for custom:<name>, lmstudio and local (vllm/llama.cpp)
routes, not only bare `custom`, so "custom endpoint" misdescribes why the
upgrade waited.
A keyed `providers:` entry runs as custom:<key>, but its display name does
not alias that id ("GPTOSS Local" -> custom:gptoss-local vs custom:gptoss),
so a pin spelled as the display name fired concurrently with the turn on
the same single-slot server (#120558 gate W2), even though the auxiliary
resolver accepts it as that endpoint. When the alias match misses, resolve
the pin through the side-effect-free config lookup
(resolve_custom_provider over get_compatible_custom_providers) and treat it
as shared when its base_url equals the main route's.
normalize_provider maps vllm/llamacpp/llama.cpp/llama-cpp to `local`, and
the auxiliary resolver sends those pins to the same local server as the
turn (#106010), but _is_self_hosted_provider only knew custom/lmstudio, so
a `vllm` title pin still fired concurrently with the main request (#120558
gate W1). Treat `local` as self-hosted and normalise inside the helper so
the main route and the pin resolve aliases (ollama, lm-studio) the same
way. The pin check now runs inside the gate's existing fail-open block
instead of its own defensive try/except around the providers import.
title_upgrade_must_wait_for_turn returned False as soon as
auxiliary.title_generation.provider was anything other than
""/auto/custom, so a pin of `custom:<name>`, the bare config name or the
display name of the very route the turn is running on fired the
json_schema title request at turn start against the same single-slot
server (#117296 race, #120558 follow-up raised in review).
A pin is now only treated as "elsewhere" when it is a hosted provider, or
when its own base_url differs. Self-hosted pins (custom / custom:<name> /
lmstudio, incl. aliases via normalize_provider) fall through to the
existing base_url comparison; a bare/display-name pin matches the main
`custom:<name>` route through hermes_cli.providers.custom_provider_aliases,
the same identity set the resolver uses.
Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: Brian Fernstrom <otstructures@gmail.com>
Final-gate reuse/quality: the refusal message re-typed the marker prefix
the new compression_marker leaf exists to own. Also pin that the bare
prefix alone is not an artifact (a prefix-only matcher now fails a test).
One block decision travelled as block_message + block_error_type +
block_payload, with _blocked_tool_result re-deriving the message from the
payload it was also handed, and _pruned_tool_arguments_block existing only to
package two module constants for its single caller. Replace message+payload
with a single block_body dict: scope/plugin blocks become {"error": msg}
(byte-identical to before), the pruned-arg block is built inline, and
_blocked_tool_result serialises whatever it is given. What the model sees is
unchanged.
Also: the _dispatch_authorized_once docstring now lists the pruned-arg check
between pre-hooks and guardrails where it actually runs, and the refusal text
gains one static sentence telling the model how to legitimately remove a
marker that already landed on disk (match the HERMES-CONTEXT-COMPRESSION
prefix instead of quoting the full marker, which the guard would refuse).
The dispatch-boundary detector re-typed the marker's wording next to the
producer's template and hid the import behind functools.lru_cache, citing a
context_compressor -> prompt_builder -> tool_dispatch_helpers cycle that does
not exist (prompt_builder only mentions this module in a comment; both import
orders succeed). Two copies of one string means a template edit silently
disables the guard.
Move the prefix/template into a dependency-free leaf, agent/compression_marker,
and build the regex from the template's first sentence with the count
placeholders swapped for \d[\d,]*. context_compressor re-exports the names, so
its callers and tests are unchanged. The matcher keeps the intended behaviour:
prefix-only mentions do not match; a minted marker, flat or nested, does. The
old \s+ tolerance for multiple spaces is dropped because the producer never
emits one, so only a template-shaped copy counts.
tool_dispatch_helpers stays light: importing it alone still does not load
context_compressor (auxiliary_client, context_engine, ...).
Two identical checks (before and after pre_tool_call hooks) doubled the
traversal for no extra safety: whatever reaches dispatch is the
post-hook argument set, so a single check placed after _pre_tool_block
covers both copied markers in the model's original arguments and
markers introduced by plugin modify hooks.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
The legacy tail check was a bare substring classifier with no producer
anchoring: the compressor on main only emits ' ...[truncated]' for
user-message/summary text, never for tool-call arguments, so any
effectful call whose content legitimately ends in '...[truncated]'
(e.g. writing docs or tests about truncation) would be refused. Keep
only the producer-shaped ⟪HERMES-CONTEXT-COMPRESSION⟫ marker regex.
Compile the regex once via a cached helper instead of per call. The
import stays lazy because agent.context_compressor -> prompt_builder ->
tool_dispatch_helpers would make a module-level import circular.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
capability_fingerprint ignored model.supports_vision and model.context_length,
so an eternal Bot Chat kept its stored prompt after those overrides changed.
A NULL stored prompt already rebuilds and no longer takes the stale probe.
`HIGH_EFFORT_SILENCE_FLOOR_SECONDS` is documented as the reasoning-effort
>= high floor and gated on effort by `_high_effort_silence_floor()`, yet
the first-progress budget applied it to every large official-Codex
request regardless of effort. Someone tuning the high-effort floor would
silently retune the first-progress fuse. `CODEX_FIRST_PROGRESS_TIMEOUT_SECONDS`
carries the same 300s, so no behaviour changes; it just gets its own knob
and rationale.
The new lifecycle-only test re-implemented `_shorten_implicit_idle_watchdog`
inline. The helper now accepts field overrides, and the test also inherits
its env cleanup so a polluted runner env cannot push the resolver off the
implicit branch.
The notice builder and the kill loop each hand-computed
`retry_started_ts or call_start` and the "stream open, no progress yet"
predicate. Two copies that must agree or the notice countdown and the
actual kill diverge. `_codex_watchdog_snapshot()` now returns the
attempt origin under the same lock, and `_pre_progress()` is the single
phase predicate for both sites.
Also restores the base preference for `retry_started_ts` over
`last_event_ts` as the non-pre-progress activity anchor. On head a
reconnect marker only coexists with events during pre-progress (retries
reset `last_event_ts`; first progress clears the marker), so the
ordering is behaviour-neutral in production and keeps the existing
`(0.0, 0.0, 1.0)` wait-notice case green.
The attempt-local first-progress deadline added `progress_timeout` to
`_NonStreamWatchdogs`, but the SimpleNamespace stubs in
test_nonstream_wait_notice.py and test_wait_notice_cadence.py were not
updated. `_emit_wait_notice` reads `wd.progress_timeout` inside its
blanket `except Exception`, so the AttributeError silently produced no
notice and no liveness touch: 11 tests red.
The new "Codex stream opened" debug line also called `time.time()`
eagerly, consuming a tick of the 3-tick stub in
test_codex_first_event_timing.py and shifting the first-event stamp.
The first-event log a few lines below already carries the attempt's
timeline anchor, so the open marker logs without an epoch.
The per-attempt "first parsed event" / "first substantive progress" /
"physical retry" markers are what an operator needs to see where a
lifecycle-only Codex attempt stalled, so they stay at info. The
"stream opened" line fires on every attempt and carries no diagnostic
value on its own, so it goes to debug to keep the default log quiet.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
The previous commit replaced the positional write with _replace_entry, leaving
the enumerate index unused (ruff B007; flagged by the pre-arm simplify gate).
`_credential_token_pair` already maps a non-dict row to (None, None), so the
separate isinstance guard and second early return in
`_merge_pool_row_generation` were the same branch. Adopting a peer generation
now goes through the existing `_replace_entry` swap primitive.
The profile-claims-its-own-credential branch of add_entry wrote the
profile store directly, so the newly owned rows had no recorded token
base until the next pre-refresh sync, and a stale peer could not be told
apart from an unknown one on the first later flush. Pass the current bases
and reseed from the written rows, mirroring _persist().
_persist() may swap a peer's newer token generation into self._entries,
but _adopt() returned the pre-persist object. Callers that rebind the live
client from that return value (try_refresh_matching -> _swap_credential,
auxiliary_client, account_usage) would keep using the stale pair for one
guaranteed 401 while the pool already held the rotated one. Re-look the
entry up by id after persisting so the return value matches the pool.
The "adopt only if the written pair differs from the live pair" guard is
kept: without it every ordinary flush would replace the live object and
pull peer cooldown state merged into the written row into memory. The
kept generation test now asserts an unchanged-pair flush preserves object
identity and that a stale writer's _adopt returns the adopted entry;
reverting either the guard or the re-lookup fails both parametrizations.
Also move the reference-only-rows comment onto the blank-pair guard it
describes instead of the rehydration statement.
The id->token-pair base map was built three ways: _persist dropped
(None, None) pairs, load_pool and _sync_entry_from_pool_store kept them.
_merge_pool_row_generation treats a missing base as "unknown" (plain
recency merge) but a (None, None) base as a known generation, so a
token-less row got the generation override on the first flush after
load_pool and the plain merge on every later one.
Policy chosen: KEEP blank bases everywhere (new auth_mod._token_pairs_by_id,
built on the existing _entry_ids). A blank base means "no pair when we last
looked", which is exactly the CAS witness the boundary needs: when a peer
lands a pair on that row, every flush of ours keeps the peer's generation
instead of writing our blank tokens back. Dropping blank bases would make
the first flush keep the peer's pair and later flushes overwrite it.
_sync_entry_from_pool_store now records the base before its no-token-
material bail-out so all three sites agree. _update_root_pool_rows reuses
_entry_ids for its incoming map as well.
Also collapse the three copy-from-disk-else-pop loops in
_merge_pool_row_generation into a local _take_from_disk helper.
failure_reason is passed alongside _POOL_STATUS_FIELDS rather than added
to the shared constant, which auth_codex/auth_oauth_grants and
_merge_disk_cooldown_state also consume.
When a stale pool's persist finds the disk pair moved, #120943 copied every
status field from disk, so the stale writer's NEWER account-wide verdict
(402 billing / 429 throttle -> EXHAUSTED) was erased and the rotated pair
re-entered selection immediately. Only a terminal auth death is tied to the
pair it was observed on: discard the stale writer's DEAD onto the peer's
pair, but leave any other status for _merge_disk_cooldown_state's ordinary
recency merge.
Also travel `scope` and `inference_base_url` with the pair (a Nous refresh
rewrites them together with the tokens, so a stale writer must not stamp an
old scope/route onto the newer pair); use the always-initialised
_persisted_token_pairs directly; and re-hydrate an in-memory entry from the
written row only when the store overrode its pair, so unchanged rows keep
runtime-only fields to_dict() omits.
Tests: fold the stale-later-402 witness (rt-1 kept AND last_status ==
exhausted) into the stale-terminal-verdict test and drop the separate 429
rollback test it subsumes.
Fixes#120815
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
handle_api_interrupt exits via break into finalize_turn, whose
_close_transcript_tail already closes the tail with the same final_response, so
its own close was a duplicate; the scaffold strip stays (the partial-text row
must follow the tool row). The scaffolding drop no longer returns a flag nobody
reads, and the give-up test relies on the autouse backoff stub in
tests/agent/conftest.py instead of re-stubbing it.
`_persist_session` closed the tool tail it uncovered after popping
empty-response scaffolding with the generic "Operation interrupted." row.
That close was dead on every path whose owner already shapes the tail
(finalizer, abort_turn_on_interrupt) and wrong where it did fire: a
non-interrupt terminal exit (retries exhausted, invalid response, rate
guard) and even the give-up path got an "interrupted" row the finalizer
then followed with a second closing row, and an in-flight Stop during the
nudge request lost its own "Operation interrupted: waiting for model
response (...)" reason because the finalizer's close no-ops on an
assistant tail.
The persist layer now only pops the scaffolding. `handle_api_interrupt`
strips it before appending its partial/placeholder row and, when it has no
partial text, closes the tail with its own reason, mirroring
`abort_turn_on_interrupt`. The give-up path is closed by
`_close_transcript_tail` with the delivered final_response, as before the
stack.
abort_turn_on_interrupt closes an open tool sequence with the caller's
specific interrupt text and then persists. When a Stop lands during
empty-response recovery, the synthetic assistant+nudge pair still sits
after the executed tool result, so the close sees no exposed tool tail
and the generic close in _persist_session wins with "Operation
interrupted.". Strip only the request-local scaffold first, so the exit
owner closes the tail with its own reason.
(taken from 413ad9647b72b27e39c49d8e0daa05bac5411ec1 in #120883,
agent/turn_recovery.py hunk only)
After an empty-response give-up or a Stop mid-recovery,
_drop_trailing_empty_response_scaffolding popped the flagged scaffolding
and then also rewound the trailing tool results and their
assistant(tool_calls) row. Those rows are saved before the tools run, so
the rewind only removed them from the live history: result["messages"]
(which CLI, TUI and ACP reuse as the next turn's history) lost a tool call
that had already executed, and a "continue" made the model run it again.
On a Stop, the finalizer then found no tool tail to close, so the saved
transcript ended on an unanswered tool result.
Drop only the flagged scaffolding rows. _persist_session closes a tool
tail the drop uncovers with close_interrupted_tool_sequence, which covers
the exits that return without the finalizer; _terminal_empty and the
finalizer already close the tail themselves.
(cherry picked from commit 9fe02fc7f27742a28ece5f789be78cfa4b4052ef)
A Journey memory edit or delete (Desktop panel, `hermes journey`, TUI
`learning.*`) rewrote the whole MEMORY.md / USER.md from an unlocked
snapshot via `_read_file` + `_write_file`: a memory the agent stored in
between was dropped, and hand-edited content that does not round-trip
through the § parser was reformatted with no .bak.
Both mutations now go through `MemoryStore._mutate` — the memory tool's
cross-process lock, re-read under lock and drift guard. The node id is
resolved to its entry text inside the lock and matched by exact text
against the re-read entries; a vanished target, drift or an unreadable
file is refused with the store's own message instead of written over.
Router/CLI/TUI response shapes are unchanged.
Fixes#119668
`/compress here N` left two live copies of a resume-merged user row when
the merged row was the newest tail row. Resume repair merges consecutive
user rows into the first dict (keeps its _row_id, drops the persist
marker; the second row's id leaves the held set). _held_watermark
exempted a verbatim tail from the marker check because its copies are
marker-swept by construction, assuming their ids were exact and newest.
That was false for the merged case: the cap landed on the merged row's
id and the absorbed durable row above it was cloned as a concurrent
append beside the row that already carries its text (main archives it).
Put the exactness decision where the marker is still visible: compress_now
keeps a tail copy's _row_id only when the source dict still carries the
marker. _held_watermark then applies one rule to every row (an id is
exact when its dict is unchanged since load: marker on a held row, a tail
copy's id trusted as given) and the newest row must have one, with no
tail exemption and max(held) computed once. `here 2` joins the existing
merge test; it is red without this change.
Resume repair merges consecutive user (and assistant) rows into the first
one's dict: that dict keeps its _row_id, drops the persisted marker, and
the later row's id leaves the held history. _held_watermark capped the
archive at max(held), so the later row sat above the cap, was cloned as a
concurrent append, and stayed live beside the merged row that already
carries its content. A context engine that rewrites the held list in
place reaches the same state.
A dict loaded from the DB is born carrying both the id and the marker, so
a newest held row with an id but no marker is exactly the rewritten case:
keep the lease watermark there, as main does. A here-N tail is exempt; it
is copied verbatim without the marker, its ids are exact, and being
newest they lift the cap above anything a merge earlier in the history
absorbed.
The new test reproduces the merge through get_resume_conversations, the
real resume path, rather than a hand-built dict. The existing
exact-duplicate check cannot see this case (the merged and cloned rows
are different strings), so it asserts the absorbed prompt appears in one
live row. Red on the previous head, green here.
(cherry picked from commit c41b9df15d72765d7d4fbb690077de463d37beed)
_held_watermark skipped trailing dicts that carried neither `_row_id` nor the persisted marker
when picking the "newest held row", assuming such a row is not durable. The TUI model-switch
marker is: server.py appends the bare dict to session["history"] and writes it with a plain
append_message, stamping nothing. With that row last in history a plain /compress capped the
archive at the previous stamped row, so the marker's durable row sat above the cap, was cloned
as a "concurrent append" AND inserted from the compacted set: two live copies (probe on the PR
head: rows 32 and 33 both the marker; origin/main keeps one).
Any trailing dict of unknown provenance now disables the cap (lease watermark, today's
behaviour) instead of being looked past. The regression case rides the existing
never-leaves-two-live-copies test as a third parameter; red on the previous predicate.
Review-fix on #120156.
(cherry picked from commit b114f564f0eddb3ccfe16123e364097747562928)
The in-place commit archives every active row up to the lease watermark, the newest row in
state.db. But a surface compacts the history it holds: the Desktop/TUI session.compress RPC
compacts session["history"], the CLI's /compress its conversation_history. When another surface
appended turns to the same session since (a Desktop session continued from Telegram after
/handoff, #42962, or a CLI resumed elsewhere), those rows sat under the watermark, took the
positional rewind slots and became active=0, compacted=0: gone from every surface's model and
display history, REST, the gateway's next replay and session_search. The summarizer never saw
them, and nothing re-adopts them. `/compress here N` (#119962) loses them the same way.
The watermark is now capped at the newest durable row the compressor was handed (its messages
plus the kept here-N tail). Rows above it take the existing concurrent-append path, cloned after
the compacted set. That is the premise _adopt_out_of_band_turns already rests on, so the next
prompt adopts them as before. The cap applies only while the held history is a live prefix of
the session: its newest durable row must carry its _row_id and still be active. Otherwise (a
history another surface already compacted, or rows held without ids, as the gateway's replay
dicts are) the commit falls back to the lease watermark, unchanged.
(cherry picked from commit 930866a2bbc2ec460898a0174c3600d97d8ed088)
_recover_from_stall now samples both the total wait and seconds-since-progress
itself, before the fallback retry, so the "reached its total ceiling after Ns"
warning reports the stall rather than stall plus retry time (as before the
ladder refactor) and the three callers stop passing the same value. The
snapshot check trusts the worker's Tuple[list, str] contract like the rest of
the module, and the teardown test goes back to a 0.3s idle budget: the 2.0s
idle made the 0.3s ceiling dead and stretched the join grace to 2s.
The settled-at-deadline no-op path, the post-cancel unchanged-commit path
and the idle-timeout path each carried their own copy of the same
retry-chain -> on_timeout -> degraded-prompt ladder (three copies after
the deterministic-fallback fix). Fold them into one closure,
_recover_from_stall, plus _is_unchanged_snapshot for the identity check.
Behavioural alignment for the two newer paths: when the retry chain yields
None they now return (messages, fallback prompt) like the idle path did,
instead of the worker's own tuple with its unset prompt, and they log the
same "no progress" warning when no on_timeout callback is installed.
Co-authored-by: Mark DiPietro <mark@svavnc.com>
Edit tool rows now carry display_metadata.tool_result_metadata.inline_diff
(~9KB of ANSI per edit) on the live message dict. build_api_messages strips
display_kind/display_metadata/_row_id before the wire, but the local
estimator's wire shadow only dropped PERSISTENCE_ONLY_MESSAGE_FIELDS
({"timestamp"}), so estimate_messages_tokens_rough and
estimate_request_tokens_rough priced the diff: one edit row went from 32 to
2527 estimated tokens (100 edits: 3100 -> 252600). That inflates compaction
preflight, post-tool checks and turn-overflow scoring, compacts early and
breaks the prompt cache.
PERSISTENCE_ONLY_MESSAGE_FIELDS (agent/message_metadata.py) is now the single
set of local-only fields: timestamp, display_kind, display_metadata, _row_id.
build_api_messages pops exactly that set and the estimator shadow drops it,
so the two can no longer drift. The iteration-limit summary path
(_iteration_summary_api_messages) hand-builds its wire messages and stripped
timestamp but not display_*; its key set now unions the same constant, so
the inline diff never reaches the provider there either (strict gateways
reject unknown keys). The tail-budget walk (_estimate_msg_budget_tokens) is
allowlist-based and already ignored these fields.
Gemini rejects an unsupported schema by naming its own generationConfig
keys ("Unknown name response_json_schema", "Invalid value at
generation_config.response_schema...", "response mime type ... is
unsupported"), never OpenAI's response_format, so such 400s bypassed the
one-retry-without-format rung and surfaced as hard failures now that the
native adapter actually sends the schema. Teach _is_structured_output_rejection
the three Gemini phrasings; the status gate (400/422 only) is unchanged.
Wiring response_format into generationConfig exposed a regression main never
had: any aux call that sends tools with tool_choice=auto AND a json_schema
format (a proxy/relay shape, or a caller that declares tools and asks for
JSON) now 400s on gemini-2.5-* with "Function calling with a response mime
type: 'application/json' is unsupported", where main silently ignored the
format and answered. Only Gemini 3+ combines function declarations with
structured output (ai.google.dev/gemini-api/docs/structured-output), so the
JSON keys are dropped whenever tools are declared on a pre-Gemini-3 model,
reusing the is_gemini3 flag build_gemini_request already computes.
Tests: the seven salvaged tests are trimmed to three invariants. The
former "keeps json output when tool_choice is auto" case asserted the
regressed behaviour (model="" is pre-Gemini-3) and is replaced by the
parametrized gemini-2.5 (drops) / gemini-3 (keeps) pair; the two schema
surfaces collapse into one parametrized test; the json_object-only and
missing-schema-key fallbacks were change-detectors on the translator.
Live (gemini-2.5-flash): tools + json_schema -> HTTP 400 on the PR head,
finish=stop on this commit; json_schema without tools still returns JSON.
Translates OpenAI-style response_format into Gemini's generationConfig as
responseMimeType plus responseJsonSchema on v1beta or responseSchema elsewhere,
reusing the existing tool-schema preparation so $ref inlining and key stripping
stay consistent. The translation is skipped when tool_choice forces function
calling, since Gemini rejects mode ANY combined with a JSON response type.
`_usage_from_metadata` (agent/gemini_native_adapter.py:476) read
`candidatesTokenCount` for `completion_tokens` and never read
`thoughtsTokenCount`; a repo-wide grep found that field nowhere. Gemini
reports hidden thinking in its own counter — `candidatesTokenCount` covers
visible output only, while `totalTokenCount` already includes thoughts. So on
any thinking turn the emitted usage contradicted itself: prompt + completion
did not add up to total.
Both the non-streaming assembly (translate_gemini_response) and the streaming
one (the finish chunk in translate_stream_event) share this helper, so both
under-billed identically. `normalize_usage` takes `completion_tokens` for
`output_tokens` and looks for reasoning under
`completion_tokens_details.reasoning_tokens`, which the adapter never set, so
the reasoning column read 0 for models whose spend is mostly reasoning.
Thoughts are now folded into `completion_tokens` (OpenAI's counter includes
reasoning) and surfaced under `completion_tokens_details.reasoning_tokens`,
matching the nesting the adapter already uses for
`prompt_tokens_details.cached_tokens`. The field is absent on non-thinking and
older responses; those count 0 and their numbers do not move.
Live: for usageMetadata {prompt 10, candidates 200, thoughts 5000, total 5210}
the adapter emitted completion_tokens=200 (10 + 200 != 5210) and
normalize_usage returned output_tokens=200, reasoning_tokens=0. It now emits
completion_tokens=5200 (10 + 5200 == 5210) with
completion_tokens_details.reasoning_tokens=5000, and normalize_usage returns
output_tokens=5200, reasoning_tokens=5000.
codex_runtime, codex_app_server and the runtime migration import this module
for HERMES_TOOLS_MCP_SERVER_NAME, so a module-level import of hermes_bootstrap
exported TMPDIR/TMP/TEMP/HERMES_SCRATCH_DIR into every library importer
(gateway.relay's read-only relay_fronted_platforms included; caught by
tests/gateway/relay/test_cold_opt_out.py).