A hand-edited "completed": "2" still crashed every recorded run ("2" + 1),
and 1.0 was stored as 2.0 ("2.0/3"). load_jobs now coerces any non-int
counter to a non-negative int (0 when unparseable). Document the load-time
repair next to the direct-edit tip.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
A hand-edited "completed": null was patched with `.get("completed") or 0`
at five arithmetic readers. The display readers were still missed: the
cronjob tool listed a never-run one-shot as "1/1", and `hermes cron`
printed "None/3". Any new reader would bring the bug back.
Every reader gets its jobs through load_jobs: mark_job_run, update_job,
claim_dispatch, the due scan, merge_job_definition, list_jobs/get_job for
cronjob_job_args._repeat_display and hermes_cli/cron._job_rows. So reset a
null to 0 once in the load_jobs repair pass, which also persists the fix,
and change the per-site `or 0` patches back to their base form.
Co-authored-by: mochamgx <1114149@qq.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Since load_jobs re-runs itself under the job-store lock whenever it finds
a repair, every detection-time warning ran twice: once on the unlocked
pass and again on the locked re-read. The id-keyed "Skipping N non-dict
entries" warning had logged once on base, so this was a new regression.
The scalar 'jobs' warning doubled too. Only the list-junk warning had a
lock-depth guard.
Collect the repair details as notes and emit them next to the existing
"Auto-repaired jobs.json" line, which only the locked pass reaches. With
that, the per-warning depth guard is no longer needed. Also build the
junk list once instead of scanning jobs three times; the
isinstance(jobs, list) check was always true at that point.
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
#122114 guarded mark_job_run/claim paths against a hand-edited
"completed": null, but update_job and the job_definition merge still used
.get("completed", 0), which returns None when the key is present, so the
null was carried forward into the store. Use `or 0` there too, and pin the
behaviour with one test covering mark_job_run and update_job.
Fixes#123281
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: mochamgx <1114149@qq.com>
`repeat.get("completed", 0)` only falls back to 0 when the KEY is
absent. When the key exists with value None (a job created but never
successfully recorded a run), it returns None, and the following
`completed += 1` raises:
unsupported operand type(s) for +=: 'NoneType' and 'int'
Effect: the run is never recorded (last_run_at/last_status stay null)
and every fire logs a scheduling error. The job itself executes fine.
Same latent hazard at the `completed < times` comparison, which a None
value would also fail on.
Fix all three sibling sites with `repeat.get("completed") or 0`.
(cherry picked from commit 0c5024de901900a747849ca44750327b56477f3b)
An unlocked load_jobs that finds junk entries hands off to a locked re-read,
which runs the same analysis again, so every repair logged the "Skipping N
non-object entries" warning twice. Log it only on the locked pass (the one
that actually saves the repair).
An unlocked reader (list_jobs) that found a repairable store saved its parse-time snapshot. The shrink-merge restores only missing ids, so a locked writer's update to a job already in that snapshot (enabled, next_run_at, run claims) was reverted. Outside the lock, the repair now re-reads and saves under _jobs_lock(); this covers every repair kind, including the pre-existing bare-list, id-map and control-character repairs.
(cherry picked from commit 0f90a10212dbc360007a59bb09acc1e0ce676f61)
{"jobs": null} (or a string, number or bool) escaped load_jobs unchanged, so every reader crashed the same way the non-object list entries did. Replace it with an empty list and persist the repair. Ported from #123405.
(cherry picked from commit 1f380ec62c2ce65fd1d31656850b611234e3385e)
An all-invalid list filtered to [] skipped save_jobs (the 'if jobs and repair' guard), so every tick repeated the warning. Save whenever a repair ran (the shrink-merge still keeps concurrent valid jobs), and log the dropped entries' types instead of their raw values.
(cherry picked from commit 3b5b1aa61332cf79554fc2a24ae925908c37fe5d)
A null, string or number in the canonical {"jobs": [...]} list reached
every reader: the due scan raised AttributeError on each tick, so no job
fired, and list/resolve crashed too. Drop the junk with a warning and
self-heal the file, as the id-keyed map flatten already does.
(cherry picked from commit 2e2438d160967392184b2193cfc2e2ca124c2b7a)
run_one_job's dispatch-failure handler already routes an
_ExternalWorkerPostHandoffError through bookkeeping only (no incident, no
ping), but the recorded error still read "Restart-safe cron worker dispatch
failed", which is misleading in last_error / executions.db for a worker that
was spawned and may have run. Compute post_handoff first and label it
"Restart-safe cron worker failed after handoff: ..."; the pre-handoff prefix
is unchanged.
_launch_external_cron_worker returns through _wait_for_external_cron_worker
once the worker is spawned/acknowledged, and that waiter can raise (e.g. the
adopted worker died and recovery could not terminalize the row). Such an
error landed in the same run_one_job except branch as a genuine dispatch
failure, so a job that actually ran got a false "Restart-safe cron worker
dispatch failed" incident and failure ping on top of whatever the worker
itself delivered.
Wrap waiter failures in a dedicated _ExternalWorkerPostHandoffError and skip
_deliver_crash_failure for it, keeping the base bookkeeping-only behaviour
(mark_job_run + finish_execution, the latter a no-op on a worker-owned row).
Pre-handoff dispatch failures still open an incident and notify.
The dispatch-failure branch called _deliver_crash_failure before the
try/finally that closes the execution row. If delivery ever raised,
finish_execution was skipped and the row stayed claimed until dead-owner
recovery. Run delivery inside the try so the pre-existing finally still
terminalizes the attempt; delivery_error/outcome default to None when it
raises.
_deliver_crash_failure is already best-effort internally (incident upsert,
_deliver_result and _mark_incident_alerted each swallow their own errors),
and the sibling crash path in run_one_job calls it without a wrapper. The
extra try/except was a defensive layer that could only mask a genuine bug
in the notice helper; drop it so both call sites match. The existing
try/finally still guarantees finish_execution runs after mark_job_run.
webtecnica submitted an equivalent independent fix in #123433.
Fixes#123401
Co-authored-by: webtecnica <webtecnica@gmail.com>
A failed restart-safe handoff in run_one_job() recorded the failure on the
job and in the executions ledger, then returned before any incident or
delivery path ran: no cron_incidents row, no failure-lane notice. Route the
dispatch-failure branch through _deliver_crash_failure() so it opens the
same job+signature incident and delivers the same failure notice as any
other job failure, with the existing alerted-cooldown withholding repeats.
A notice-path exception no longer loses the bookkeeping: mark_job_run and
finish_execution still run with a "failed" delivery outcome (#123401).
(cherry picked from commit e5b5969df1b7ca212c5a6e27d30f4778fb835c52)
plan_retry's exhausted-ladder warning re-listed _ladder_instant's gates
(recurring, not paused, retry enabled) by hand, so the two could drift and
the warning fire for a job the ladder never applied to. Both now call
_ladder_applies(job). _ladder_instant checks the attempt count first, so the
exhausted path loads config exactly once (in the warning gate).
Behaviour unchanged: old-vs-new equivalence over 1,728 job states
(will_retry, plan_retry result, mutated job, log levels) shows 0 diffs.
WHY: will_retry hand-copied plan_retry's decision (recurring/paused, enabled,
ladder exhausted, next-rung instant, yield to the natural run) - the exact
drift that let fast jobs hold failure notices forever. Both now call a pure
_ladder_instant(job, natural_next, now); will_retry keeps only its own checks
(final finite repeat, uncomputable natural run) and reads the clock once.
The final-repeat check is spelled `times is not None and times > 0` as in
_advance_after_run. Behaviour is unchanged: an old-vs-new differential over
1728 job states matches on will_retry, plan_retry result, job state and logs.
Co-authored-by: Yuan Li <dskwelmcy@163.com>
WHY: on the last run of a finite repeat, _advance_after_run completes the
job and mark_job_run skips plan_retry (is_terminal_job), but will_retry still
predicted a re-run, so that final failure notice was held forever. Mirror
the terminal guard. Folds in the edge reported by #109991 (Liuzikaii).
WHY: the salvaged docstring narrated the incident and referenced another PR;
replace it with the invariant the predictor must hold. The test only covers
the yield branch, so drop "terminal_paths" from its name.
will_retry gated notice suppression on recurring/paused/attempt/config only.
plan_retry has a yield branch: when the schedule's own next occurrence is at or
before the pending ladder rung it schedules nothing and clears state without
consuming an attempt. For a job on a cadence at or under a rung (<=5m, and the
15m/30m rungs for faster cadences) every failure hit that branch, the attempt
counter never advanced, and will_retry kept answering True — so during a
sustained outage every failure notice was held forever. The documented escape
('once the ladder is exhausted, the next failure alerts normally') was
unreachable: the ladder could never exhaust.
will_retry now recomputes the natural next occurrence exactly as
_advance_after_run will and answers True only when the rung precedes it — i.e.
exactly when plan_retry will actually park a re-run. Notices now go out on the
first failure for cadences the ladder cannot help, and slow jobs keep their
silent bounded re-runs.
(cherry picked from commit 9d3d6006204269103188a5ee366d8eb7c9f482a2)
The 19-line block narrated review history and still carried the base claim
"rather than interpreting relocated pyvenv.cfg", which is now wrong: with no
generation committed the code falls through to the handed venv's pyvenv.cfg.
Four lines state the rule. Comment-only.
The no-generation arm returned the bare store Python with a repo-only
overlay, which is worse than the main behavior it replaced: a store
Python with nothing committed is refused by _require_own_dependencies
with "hermes pm repair", and a cron script has no hermes_bootstrap to
refuse cleanly, so it runs and dies on its first third-party import. It
also stripped the overlay's site-packages entry, so
_windows_cron_bootstrap_argv warned on every spawn.
Fall through to the uv overlay instead: the handed venv keeps its own
base interpreter and its own packages, which is what worked on main.
The committed-generation arm is unchanged and still returns the store
Python with that generation's site-packages.
Also corrects the comment: site_packages is a pure path join on Windows
(pm/environments.py:250-255) and never raises there; the raise is the
pyvenv.cfg probe inside _recorded_venv:192, already covered here, and
dependency_site() is outside this try. (#123668 review)
(cherry picked from commit b3760eff7a1ea8207b0f90c4c9cbfdbc6ab753a2)
_windows_cron_python_invocation selected the dependency tree with
selected_venv, which falls back to base_venv and answers the leftover
pre-PM <root>/venv when no generation is recorded. That tree belongs to
whichever interpreter created it, so on a PM-managed install the managed
store Python 3.14 got a cp311 site-packages on PYTHONPATH and every cron
script died with "No module named 'pydantic_core._pydantic_core'" -- the
cron sibling of the gateway crash in #122183, which janviernine flagged
there and left out of scope.
committed_venv never answers with the in-tree venv. With nothing committed
this install provisioned no tree, so overlay none and let the child keep
the interpreter it was handed rather than borrowing a foreign ABI.
Every other selected_venv caller (tools/environments/local_pythonpath,
hermes_cli/_early_recovery, pm/extras, pm/environments_adopt) already gates
on runtime_facts_path().is_file(); this was the only unguarded one.
(cherry picked from commit dd2b7181ab0aec671ff1f8d0946e0d963cb50f52)
A direct jobs.json expression edit while a retry was parked still fired once
at the old ladder instant; record the expression the retry was planned under,
as the quota-hold recovery fire does. Also drop an unreachable None guard in
_recovery_worthwhile (its only call is behind 'not blocked').
unreachable_retry.plan_retry parks a cron job at now + ladder delay, an
instant that is off the expression's lattice. The stale-cron guard on the
due scan classified it as a direct schedule edit (STALE_CRON_EXPR_EDIT) and
re-anchored it to the natural occurrence without firing, so the 5/15/30-minute
ladder never ran for cron jobs — the same failure mode the quota-hold recovery
fire had to exempt itself from. Record the ladder instant in the retry state
and let _reanchor_stale_cron accept either planner's parked instant as an
authorized one-shot off-lattice fire; interval jobs and legacy state without
the instant keep the previous behaviour.
_recovery_worthwhile subtracted same-tz aware datetimes directly, which is
wall-clock arithmetic and off by an hour when the hold-to-occurrence span
crosses a DST transition (weekly jobs make that reachable); use the module's
canonical _elapsed_seconds. Derive the schedule from the job and gate the
"scheduled fire" boolean at the call site under one name
(recover_consumed_fire) from scheduler -> mark_job_run -> plan_hold instead
of three. Collapse plan_hold's interval/cron arms, which duplicated the
not-blocked early return and the boundary park. Persist the expr fingerprint
only for the off-lattice recovery fire: the coalesce branch parks on the
lattice (STALE_CRON_MATCH, also under the croniter-missing fallback), so
is_recovery_fire is never consulted there. Drop a comment restating the
docstring.
The recovery fire from #117437 re-ran EVERY scheduled cron whose natural
occurrence fell past the hold boundary, so dense schedules fired twice
(hourly job, hold ending :58 -> fires :58 AND :00), and because the
recovery fire is itself a scheduled dispatch a second quota failure
re-parked it at the next boundary — a probe per hold window for the whole
outage, the exact cost the hold exists to prevent.
Gate the recovery on (a) the job not already carrying quota_hold_until
(the recovery fire failing again waits for the natural schedule) and (b)
the natural occurrence being at least half a cadence past the boundary,
reusing the half-period rule of cron.jobs._compute_grace_seconds. A full
cadence gap is impossible for the motivating case (weekly job, 20h hold:
gap is ~6.2 of 7 days), so half a period is the sparse threshold.
Co-authored-by: Marko Niskala <manis@aivelho.local>
Use the scheduler's own runnable predicate instead of a hand-rolled
enabled/state check, so a contradictory half-paused record is treated the same
way the scheduler treats it.
hermes_cli still reached into cron's private _jobs_lock and hand-built the
"created paused" record one function away from the module created so callers
never duplicate cron's schema. import_job_definitions() now holds the lock,
loads, merges and saves, and labels a merge ValueError with the job name;
_merge_cron_store keeps only the temp-store parse and the DistributionError
wrap, and takes the profile home instead of deriving it from dest.parent.parent.
Also drops the dead `and key != "repeat"` (repeat is reassigned right after)
and the comment that restated the module docstring.
A shipped job whose authored schedule was a past one-shot (or an unparseable
string) raised ValueError from _apply_schedule_update in the middle of
_copy_dist_payload: SOUL.md and skills were already replaced, the cron store
was not, and the message named no job. `hermes profile update` ended half-applied.
- _copy_dist_payload merges the cron store first, so the one step that can
reject shipped content runs while the profile is still whole; the merge
itself only writes after every record merged.
- merge_job_definition normalises string schedules via parse_schedule (a
hand-authored store previously persisted the raw string, and `cron resume`
then crashed on `.get`), and passes schedule_display only when the authored
record has one so the helper's display fallback applies.
- _merge_cron_store wraps the ValueError into a DistributionError naming the job.
- Paused/created stamps use hermes_time.now() like every other cron record.
The kept update test now covers the past-one-shot rejection (SOUL.md untouched,
local schedule kept) and a corrupt target store surfacing as DistributionError,
so dropping the error wrapper goes red.
cron/jobs.py is already ~3400 lines; JOB_DEFINITION_FIELDS and
merge_job_definition are only used by importers of a foreign cron store
(profile distributions), so they live in a small dedicated module instead
of growing the scheduler file. profile_distribution imports the new module
directly; jobs.py is left byte-identical to main (no re-export shim).
Refs #120823
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
On Windows (fr-FR, es-AR, de-DE reports) the zone name arrives in the ANSI
code page but is decoded under a UTF-8 LC_CTYPE (UTF-8 mode, or Piper/espeak
flipping the process locale mid-run) with surrogateescape. datetime.strftime
splices tzname() in as UTF-8, so "%Z" raised UnicodeEncodeError while the
system prompt was being built and every new/compressed conversation died.
Rework safe_strftime into a small repair: output is untouched for valid text
(the system prompt stays byte-identical), surrogateescape'd bytes decode back
through the ANSI code page ("heure d'été"), anything else degrades to U+FFFD,
and a raising "%Z" is rendered from the repaired tzname(). hermes_time no
longer imports agent.* at module load.
Route the remaining locale-name sites through it: cron quota-hold notice,
auxiliary cooldown notice, cron session titles, session_search dates,
insights, learning graph and billing renew dates. Tests use real datetimes
with a surrogate zone name instead of stubbed strftime.
Fixes#102910
Co-authored-by: Aniruddha Adak <127435065+aniruddhaadak80@users.noreply.github.com>
OpenRouter relays OpenAI's "this user has been blocked for a previous
policy violation" as HTTP 200 + an SSE error event. The SDK raises a
status-less APIError that no classifier rule matched, so it landed in
the retryable `unknown` bucket: every retry was re-sent (3 streamed
requests by default), a configured fallback only engaged after the
ladder, and the user was told the provider "looks temporarily
unavailable. Wait a minute and send /retry".
- error_classifier: match the ban phrase status-agnostically in
_provider_special_cases -> provider_policy_blocked (non-retryable,
fallback, no credential rotation: every key on a banned account is
banned; a 403 variant is not a bad key).
- turn_failure_copy: provider_policy_blocked chat copy now covers an
account block as well as data/privacy settings; add its cause gloss
so cron and subagent notices explain it instead of printing raw text.
- cron: provider_policy_blocked action (retrying won't help; pin another
model).
- docs: FAQ entry for the reply; auto-recovery ladder exclusion list.
Same class as the due-scan fix: the configured zone is one cached ZoneInfo, so two aware
datetimes in it compare and subtract by wall clock, and aware + timedelta is wall-clock
arithmetic that drops fold. Inside the repeated 01:xx hour:
- unreachable_retry.plan_retry: now(01:30 EST) + 300s landed on 01:35 EDT, 55 min in the
past, so the "re-run in 5 min" fired on the next tick.
- quota_hold: the provider-window end was an hour off (hold too long, or already expired)
and hold_active compared wall clocks.
- resume_job: the "slot elapsed while paused" check compared wall clocks and could drop
the kept slot (#113603 path).
- _repeat_alert_withheld: the repeat-alert window was measured in wall time.
Add jobs._seconds_after (UTC add, back to the caller's zone) next to the instant helpers
and use the instant comparisons there. Found by the C13 virtual-clock cron soak.
Use UTC-normalized comparisons and elapsed durations across due scans, claim ages, recovery, and the external misfire backstop. Add deterministic fall-back regressions for both folds and ordinary intervals.
Co-authored-by: Jaimin <95100522+Jaiminp007@users.noreply.github.com>
One resolver, hermes_cli.profiles.current_profile_name(): the HERMES_HOME override
(a multiplexed cron tick or routed gateway turn) names the profile first; the
dispatcher's HERMES_PROFILE pin is consulted only when no override is bound; the
process home last. #112888 added the HERMES_HOME-derived fallback but kept the env
pin FIRST, so under a multiplexer whose launch process carries a HERMES_PROFILE the
served profile's writes were still re-labelled with the host's name. The same
resolver replaces the per-module copies in hermes_cli/kanban.py, kanban_specify.py,
cron/lifecycle_guard.py and the kanban notify-target default.
Tests trimmed to two invariants (A->B->A under the override; control pin + generic
absence).
Closes#119859
Supersedes #112888
* refactor(fallback): share the pinned-owner chain rule
delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.
* fix(cron): a pinned job never falls back to the global chain
A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:
- _resolve_job_runtime walked the chain on an AuthError or transient
network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
fallback_model, so the conversation loop's provider ladder could swap a
pinned job mid-run.
Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".
Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.
No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* docs(cron): pinned jobs do not use fallback_providers
cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.
---------
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
Several profile-scope binders called set_hermes_home_override and
set_secret_scope before the try/finally that releases them. A raise in
scope setup (a corrupt or removed profile home) propagated with the
foreign override still bound to the caller's context, silently re-homing
every later read — and in _reregister_orphaned_adopters it also skipped
every remaining adopter. The set calls now run inside the try with
None-guarded resets; the same shape is fixed in the routed-turn scope,
the cron external worker (which also leaked the multiplex flag), the
kanban worker scope, the MCP OAuth paths, the launch-profile policy,
and model_switch, which releases partially-bound scopes on raise.
get_secret returned os.environ on a scoped miss whenever multiplex was
off, but non-multiplex hosts serve foreign homes too (dashboard/desktop
backend, per-profile cron, MCP owner scopes, kanban spawn-env builds),
where os.environ is the launch profile's. Bound scopes now carry the
home they were built for; serves_routed_profile detects a foreign scope
even when the binder deliberately skips the HERMES_HOME override, and
the miss returns the caller's default. Every production binder stamps
its home; own-home scopes keep the deliberate env overlay.
cron/scheduler.py and tui_gateway/prompt_turn.py each parsed
turn_exit_reason for max_iterations_reached( inline. Both now read
agent.turn_failure_copy.is_max_iteration_handoff so cron delivery and the
/goal judge gate cannot drift on what counts as a non-failed handoff.