A failed restart-safe handoff in run_one_job() recorded the failure on the
job and in the executions ledger, then returned before any incident or
delivery path ran: no cron_incidents row, no failure-lane notice. Route the
dispatch-failure branch through _deliver_crash_failure() so it opens the
same job+signature incident and delivers the same failure notice as any
other job failure, with the existing alerted-cooldown withholding repeats.
A notice-path exception no longer loses the bookkeeping: mark_job_run and
finish_execution still run with a "failed" delivery outcome (#123401).
(cherry picked from commit e5b5969df1b7ca212c5a6e27d30f4778fb835c52)
The comment claimed the notice goes out immediately. In production
will_retry runs before mark_job_run, so the first failure (5m rung beats
the 10m run) is held; only the attempt-1 failure asserted here yields.
plan_retry's exhausted-ladder warning re-listed _ladder_instant's gates
(recurring, not paused, retry enabled) by hand, so the two could drift and
the warning fire for a job the ladder never applied to. Both now call
_ladder_applies(job). _ladder_instant checks the attempt count first, so the
exhausted path loads config exactly once (in the warning gate).
Behaviour unchanged: old-vs-new equivalence over 1,728 job states
(will_retry, plan_retry result, mutated job, log levels) shows 0 diffs.
WHY: the ten-line docstring restated the commit narrative. Keep the
invariant and note why calling will_retry after mark_job_run is valid: the
predictor reads only persisted job state.
WHY: will_retry hand-copied plan_retry's decision (recurring/paused, enabled,
ladder exhausted, next-rung instant, yield to the natural run) - the exact
drift that let fast jobs hold failure notices forever. Both now call a pure
_ladder_instant(job, natural_next, now); will_retry keeps only its own checks
(final finite repeat, uncomputable natural run) and reads the clock once.
The final-repeat check is spelled `times is not None and times > 0` as in
_advance_after_run. Behaviour is unchanged: an old-vs-new differential over
1728 job states matches on will_retry, plan_retry result, job state and logs.
Co-authored-by: Yuan Li <dskwelmcy@163.com>
WHY: on the last run of a finite repeat, _advance_after_run completes the
job and mark_job_run skips plan_retry (is_terminal_job), but will_retry still
predicted a re-run, so that final failure notice was held forever. Mirror
the terminal guard. Folds in the edge reported by #109991 (Liuzikaii).
WHY: the salvaged docstring narrated the incident and referenced another PR;
replace it with the invariant the predictor must hold. The test only covers
the yield branch, so drop "terminal_paths" from its name.
will_retry gated notice suppression on recurring/paused/attempt/config only.
plan_retry has a yield branch: when the schedule's own next occurrence is at or
before the pending ladder rung it schedules nothing and clears state without
consuming an attempt. For a job on a cadence at or under a rung (<=5m, and the
15m/30m rungs for faster cadences) every failure hit that branch, the attempt
counter never advanced, and will_retry kept answering True — so during a
sustained outage every failure notice was held forever. The documented escape
('once the ladder is exhausted, the next failure alerts normally') was
unreachable: the ladder could never exhaust.
will_retry now recomputes the natural next occurrence exactly as
_advance_after_run will and answers True only when the rung precedes it — i.e.
exactly when plan_retry will actually park a re-run. Notices now go out on the
first failure for cadences the ladder cannot help, and slow jobs keep their
silent bounded re-runs.
(cherry picked from commit 9d3d6006204269103188a5ee366d8eb7c9f482a2)
Follow-up to the salvaged restart-wait commits:
- A drain or cron timeout of .inf now means "wait indefinitely" instead of
an OverflowError from the integer stop envelope, which crashed
`hermes gateway restart` and made `hermes update` silently fall back to
its 45s floor. The fleet "draining (up to Ns)" lines and the drain
progress report format the budget instead of int()-ing it, so an
unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
envelope and that configured cron reaches the CLI wait, instead of
re-deriving the formula; the negative wording assertion on the pending
footer is dropped (change-detector).
Neither kept test reaches a pyvenv.cfg read: the committed-generation path
returns before the cron fall-through, and the gateway overlay never parses
it. The `version=` arg, base/home dirs and the legacy cfg content were setup
nothing consumed. The cron test's selected_venv patch looks inert on head but
is what makes it red on base, so it now says so. Both tests still pass on head
and fail with base prod files (fake-win32 harness).
The 19-line block narrated review history and still carried the base claim
"rather than interpreting relocated pyvenv.cfg", which is now wrong: with no
generation committed the code falls through to the handed venv's pyvenv.cfg.
Four lines state the rule. Comment-only.
The `Path(sys.prefix).resolve() != resolved_venv` guard skipped exactly the
case _ensure_windows_gateway_venv_imports exists for (264ac72b67): a gateway
restarted under uv's base pythonw.exe that still needs venv/Lib/site-packages
(MCP SDK). When sys.prefix IS the venv its site-packages is already on
sys.path, so the guard turned the function into a no-op on non-PM installs.
It is not needed for #122183: the committed_venv early return alone keeps a
PM install off the leftover pre-PM venv (and a corrupt facts.json raising
there is fail-closed, matching hermes_bootstrap). Mutation-checked with a
fake-win32 harness: removing that early return turns
test_committed_generation_blocks_the_legacy_venv_overlay red.
_ensure_windows_gateway_venv_imports prepended VIRTUAL_ENV or <root>/venv
unconditionally. On a PM install hermes_bootstrap has already activated
the store Python onto the committed generation, so that late overlay put
a foreign-ABI tree (cp311 pydantic_core under 3.14) first on sys.path and
the gateway died on import.
- committed generation present -> return: bootstrap already made the
one activation decision; no second, late injection.
- committed_venv errors propagate (a corrupt facts.json is an error, not
permission to load the legacy venv).
- nothing committed -> only overlay a candidate that is the running
interpreter's own prefix, instead of parsing pyvenv.cfg versions (uv
writes version_info, which a version-only parser misses).
Co-authored-by: Robby Slamet <robbyslmt@users.noreply.github.com>
The salvage drops the try/except wrapper around committed_venv (a corrupt
facts.json raises as it did on main), so the "unreadable record falls
through" arm no longer describes the code. Keep the one invariant that
pins the bug: a committed generation, never the stale in-tree venv, is
what the cron child gets. Stack budget is two invariant tests.
Co-authored-by: Halldrix <halldrix@users.noreply.github.com>
platforms("windows") instead of a bare skipif: list_os_marked_tests.py
gates the Windows lane on that marker, so the skipif version ran on no
CI lane at all -- the antipattern root AGENTS.md calls out.
Four tests collapse to the two invariants that actually matter: the
committed generation wins, and no-generation/unreadable-record both
fall through to the handed venv. The legacy non-managed test is gone --
it passes on main by design, and the fall-through test's positive
control now covers that path inside the managed arm.
(cherry picked from commit 30c649f16bce08eeaabd5a407ae0429cf422556a)
The no-generation arm returned the bare store Python with a repo-only
overlay, which is worse than the main behavior it replaced: a store
Python with nothing committed is refused by _require_own_dependencies
with "hermes pm repair", and a cron script has no hermes_bootstrap to
refuse cleanly, so it runs and dies on its first third-party import. It
also stripped the overlay's site-packages entry, so
_windows_cron_bootstrap_argv warned on every spawn.
Fall through to the uv overlay instead: the handed venv keeps its own
base interpreter and its own packages, which is what worked on main.
The committed-generation arm is unchanged and still returns the store
Python with that generation's site-packages.
Also corrects the comment: site_packages is a pure path join on Windows
(pm/environments.py:250-255) and never raises there; the raise is the
pyvenv.cfg probe inside _recorded_venv:192, already covered here, and
dependency_site() is outside this try. (#123668 review)
(cherry picked from commit b3760eff7a1ea8207b0f90c4c9cbfdbc6ab753a2)
_windows_cron_python_invocation selected the dependency tree with
selected_venv, which falls back to base_venv and answers the leftover
pre-PM <root>/venv when no generation is recorded. That tree belongs to
whichever interpreter created it, so on a PM-managed install the managed
store Python 3.14 got a cp311 site-packages on PYTHONPATH and every cron
script died with "No module named 'pydantic_core._pydantic_core'" -- the
cron sibling of the gateway crash in #122183, which janviernine flagged
there and left out of scope.
committed_venv never answers with the in-tree venv. With nothing committed
this install provisioned no tree, so overlay none and let the child keep
the interpreter it was handed rather than borrowing a foreign ABI.
Every other selected_venv caller (tools/environments/local_pythonpath,
hermes_cli/_early_recovery, pm/extras, pm/environments_adopt) already gates
on runtime_facts_path().is_file(); this was the only unguarded one.
(cherry picked from commit dd2b7181ab0aec671ff1f8d0946e0d963cb50f52)
Follow-up to the light-mode contrast fix; no visual change on the catalog.
- Edit the existing light .heroEyebrow rule in place instead of shadowing
it from the appended block (it was left dead at the old #a8860a).
- One rule per component: .starPill, .capChip, .envChip, .copyBtn,
.highlight, .versionPill/.platformPill and .categoryChip each had colour
in one block and border/background in another; .categoryBtnActive set its
colour twice; .docsLink:hover/.emptyCtaSecondary:hover got a colour that
the later hover rule always overrode.
- The amber was hardcoded ~20 times although it is the theme's own
--ifm-color-primary-darker (the search icon's is --ifm-color-primary);
use the variables, including for --plugin-catalog-official.
- Share the light link blue between the list and detail pages via
--plugin-catalog-link so the two can't drift.
- Plugin detail/author pages used #8a6d00 for light star counts while the
list now uses the darker amber; align them on the same variable.
Verified: computed colour/background/border/shadow of every element on the
catalog, a plugin page, an official plugin page, an author page and the
empty state, in both themes plus hover on 19 controls, is identical to the
contributor's head except the intended star colour on detail/author pages.
Version pills and the eyebrow on a plugin page used the same pale colors that disappear on a white background.
(cherry picked from commit a1051e2565a1f2daad934362fc31319020e1c24e)
Gold and pale-blue chips, dates, and links sit around 1.2:1 on white, so the catalog grid washes out. Light mode now uses the dark amber and blue already chosen for the docs theme.
(cherry picked from commit b84d6d8fac4ff3f2ef5b9fe32558b8d2a79fae47)
The marker-preservation test is subsumed by the _db_flush_collect
regression test and the updated exact-dict assertion in
test_cached_agent_history_guard.py; keep one invariant test. Also note
why _build_replay_entry carries the stamp: replay rewrites are view-only
and marker-only flushes must skip rows already persisted.
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: Gaurav Saxena <gauravsaxena.jaipur@gmail.com>
The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.
Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
A profile-wide agent.service_tier: fast reaches every new session,
including local llama.cpp ones. The request builders already drop the fast
parameter on routes that do not bill for it, but session.info reported
fast: true whenever the tier was priority. The desktop then kept the Fast
switch visible so it could be turned off, tagged the model row Fast and
appended "· Fast" to the pill for a local model.
session.info now asks resolve_fast_mode_overrides, the gate the request
builders use, before it reports fast. While a model switch is pending, the
new pick decides, because the agent's base URL still belongs to the old
route. service_tier is still reported as stored. The TUI status bar reads
info.fast alone.
The package-manager switch (#102765) made engine detection read only PM's
store and pinned every llama.cpp backend to b10362. A machine with an
engine under runtimes/llamacpp/b<tag>/<backend>/ reported no engine, so
the local models pane showed one-click setup with models already on disk.
installed_engine() now moves the newest pre-PM install into the store
once. The old manifest's archive digests become the PM identity, the
package verifier runs llama-server --version, and os.rename puts the
directory in place before facts.json records it. A store that already
holds an engine is left alone, a store lock held by another PM operation
defers the move, and an install that fails verification stays where it
is for the rest of the process.
The lock returns to b10964 on all five backends, the default before the
PM switch. With b10362 pinned, the pane offered a downgrade from the
moved engine as an update. Upstream renamed the ROCm archives to 10.0 at
b10767, and the hip assets follow. Only the CUDA build is measured at
b10964.
Models used to live at <profile home>/models (adb2fdbf5c) before the managed
runtime moved them to the machine-scoped <root>/models (43e67d872f). GGUFs
staged under a named profile's old dir silently stopped being served.
adopt_legacy_models() renames them (and their assets/) into the current dirs,
so everything downstream keeps reading one directory: listing, presets,
delete, the --models-dir fallback. It runs at boot, before the "anything
staged?" check, and in the Local Models status route so the pane lists them
while the runtime is off.
os.rename only: instant within a filesystem, while shutil.move would silently
copy tens of GB across devices at session start; a cross-device dir stays put
with a warning. An existing destination name is never replaced.
Busy steer (the busy-mode and priority paths, and /steer) pushes a new message
into the running turn, which then answers it without a turn of its own.
Redirect already folded the message's reply_expected into the turn; steer
did not, so an unaddressed opener steered by an @mention could still end
on a bare silence marker. All of them now go through _fold_into_running_turn.
Also: stale comments on MessageEvent and the queued-terminal silence
verdict, and a crash-recovery case for an unaddressed row.
With reply_in_thread: false the whole channel is one session and the bot
answers top-level, so an unmentioned top-level message there is a
follow-up in a conversation the bot is part of, like a thread reply.
reply_expected is now False for a free-channel message only when it starts
its own session (a new top-level thread), else None. The bot-id set is
built inside _slack_reply_expected, as _channel_gate_allows does.
reply_expected was read from the event that opened the turn, so an addressed
message the same turn ended up answering could still end on a bare silence
marker and vanish:
- a queued chain paired the terminal turn's display kind with the opener's
flag; the recursive run now receives the pending event's flag, persists
it, and returns it as queued_terminal_reply_expected beside
queued_terminal_display_kind for the outer shaping to read;
- pending-message merges (merge_pending_message_event, text batching, busy
debounce) and an active-turn redirect folded the new message in but kept
the old flag; MessageEvent.absorb_reply_expected now folds it: an
addressed message wins, then an unknown one.
Also: reply_expected is persisted only when the adapter set it, so rows on
other platforms carry no null key; the turn_params pop and getattr reads
go (turn_params already flows into TurnContext); silence_allowed is a
module import; crash recovery keeps the diagnostic-mute check on machinery
turns and applies silence_allowed only to the silence verdict; the DEBUG
line no longer fires for machinery turns.
The Slack rule marked every admitted message that was not a DM, a mention
or a command as not addressed, so a plain "done?" in a thread the bot is part
of, or a reaction trigger, could end on a bare silence marker and vanish,
the case #111624 fixed (#110952).
reply_expected is now False only for a message that opens by @mentioning
someone else, or a top-level message a free-response channel admitted
without a mention. Reaction triggers and pipe-form self mentions count as
addressed; other thread replies are None (visible fallback). The
free-channel predicate moves into _slack_is_free_channel so the gate and
the rule read the same one. The test drives the real _handle_slack_message.
Docs describe the rule in its own note instead of the
ignore_other_user_mentions tip, and the messaging index documents the
human-turn fallback.
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.
`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.
Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.
(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
get_default_model_for_provider('bedrock') returns the static list's first
entry, so putting Opus 5.5 at [0] silently moved unconfigured Bedrock
users, gateway/API-server runs with no model, and '/model bedrock' from
Sonnet 5 to the priciest flagship. Opus 5.5 now sits second.
Replaces the Opus-5-only parametrized snapshot with a relationship: every
Claude id in the Bedrock static fallback resolves offline to the same
window as its bare id in DEFAULT_CONTEXT_LENGTHS. Covers Opus 5.5 (added to
the fallback in the previous commit) and any future picker addition,
without freezing a value.
The static Bedrock list is what the picker shows when live discovery
(ListFoundationModels + ListInferenceProfiles) is unavailable; it had no
Opus 5.5. Its 1M window comes from the anthropic.claude-opus-5 key in
BEDROCK_CONTEXT_LENGTHS (longest-substring match).
Salvaged from #119444 (catalog hunk only).
(cherry picked from commit 22b7e61f6b)
The native Anthropic curated list had no Opus 5.5 even though the
OpenRouter table in the same file already lists anthropic/claude-opus-5.5,
so the native picker showed it only as a live-discovered extra at the
bottom. Uses the hyphenated id that /v1/models returns.
Salvaged from #119444 (catalog hunk only; the thinking/OAuth/tool_choice
halves are covered by #121016 and #106414).
(cherry picked from commit 757ebf625ae0fee45aae8df29b9d140c16c27a8c)
DEFAULT_CONTEXT_LENGTHS declares claude-opus-5 as 1M, but the Bedrock
static table never got the entry, so the offline path resolved the 128K
default and the agent compressed context ~8x early. Add the key and pin
the table pairing with a test so the next 1M generation cannot drift.
Co-authored-by: JiaDe-Wu <JiaDe-Wu@users.noreply.github.com>
(cherry picked from commit c1be16ecbdd267b14124dd35c0b92ccfcd201818)
RegardV opened the first Bedrock Opus 5 context fix (#75432) that ryrenz
salvaged in #75824; JiaDe-Wu is co-author on that commit (#98494).
Co-authored-by: RegardV <regardv4life@gmail.com>
The adapter auth check is synchronous and runs on the adapter's event loop:
once per inline-button tap, and once per keystroke for Telegram inline
queries. Entering `_profile_runtime_scope` with its default
`hydrate_secrets=True` there calls `hydrate_profile_secret_sources`, which
takes the process-global secret-source lock and may resolve external secret
backends. That lock contention is the heartbeat-starvation class #99519
moved off the loop.
Startup and the per-profile message path already hydrate this profile's
sources off-loop, so the callback now binds the scope with
`hydrate_secrets=False` (the reconnect-retry precedent in the same module):
`build_profile_secret_scope` still re-reads the profile's `.env` per call,
so allowlist edits keep reaching the next tap.
The regression test also seeds the default profile's allowlist (file and
live os.environ) with a different user and asserts the secondary's bot
refuses them; on main that user was admitted on the secondary's buttons.
Review feedback (#120650): the prebuilt secret scope froze the profile's
`.env` at configure time, so an operator's runtime TELEGRAM_ALLOWED_USERS
edit reached the next text message (per-message scope re-read) but not the
next button tap, and it skipped hydrate_profile_secret_sources. Enter the
owning profile's runtime scope per call via _scope_or_null — the exact
pattern every sibling secondary handler already uses — so taps and messages
share one freshness and hydration story, and an unresolvable profile home
still binds nothing (fail-closed).
Under multiplex_profiles, a secondary profile that owns its own Telegram bot
gets its auth callback from _make_adapter_auth_check(profile_name=...), but
that branch ran _is_user_authorized with NO profile scope installed: the
callback is built at configure time and invoked from the adapter's event
loop, outside _profile_runtime_scope. The gate's scoped env read therefore
fell back to os.environ — the default profile's env — so an allowlisted
caller tapping an inline button (exec approval ea:) was refused with the
'bot is private' toast, while plain text messages from the same user in the
same DM were authorized (their handler runs inside the scope). This is the
per-credential secondary-bot row of the transport matrix; the shared
primary row was fixed in 74775df53f (#86296).
The secondary branch now re-enters the owning profile's runtime scope per
call, mirroring _make_profile_message_handler. The secret scope is prebuilt
at configure time so entering it in the callback never does file IO on the
event loop; an unresolvable profile home keeps the fail-closed behavior of
every other secondary handler. (#120639)