A spilled hybrid got `-ot blk\.\d+\.ffn_.*\.weight=CPU`, which moves every
block's FFN — full-attention blocks included — regardless of how many bytes
actually had to leave the GPU. The GGUF reader now records FFN weight bytes
per block, and spill_overrides picks the fewest recurrent (n_head_kv==0)
blocks whose FFNs cover the decision's spill_bytes, emitting them as an
explicit `blk\.(i|j|...)\.ffn_.*\.weight=CPU` selector. Unknown sizes fall
back to every recurrent block, never a full-attention one.
Based on #120778 by @kvnloo (recurrent-index selector).
Covers the placement half of #113329; the overhead/logits calibration
(RUNTIME_OVERHEAD_BYTES, ub_logits_bytes) is untouched pending measured data.
Refs #113329
llama.cpp permits attention.sliding_window_pattern as a scalar period as
well as a per-layer array; current writers for Gemma-family and other
architectures use the scalar form, which read_gguf_header/profile_from_gguf
silently dropped (isinstance(v, list) is False), sending unlisted
architectures back to the all-full-attention fallback this fix was meant to
avoid.
Add GGUFHeader.sliding_window_pattern_period and expand it per llama.cpp's
set_swa_pattern() semantics, including the small set of architectures
(cohere2moe, modern-bert, smallthinker, laguna) whose loader uses
dense_first=True instead of the default False.
Also add binary GGUF round-trip tests (real read_gguf_header() parsing, not
GGUFHeader built by hand): array pattern with non-bool element types, both
scalar dense_first variants, and a length-mismatched array falling back
safely.
profile_from_gguf() only consulted a hardcoded architecture -> SWA-fraction
table (_SWA_LAYER_FRACTION). Any architecture absent from that table (e.g.
gemma4) fell through to pricing every layer as full attention, wildly
inflating the estimated KV cache and causing the physics check to refuse
models that fit and run fine.
GGUF files that declare a per-layer SWA layout already carry
attention.sliding_window_pattern (plus attention.key_length_swa /
value_length_swa for the SWA layers' dimensions) — gguf.py had no accessor
for any of the three. Add them, and make profile_from_gguf() prefer the
file's own per-layer pattern over the name table, falling back to the
table and then to all-global exactly as before when the file doesn't
declare one.
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
Boot presets and in-session growth priced the context window from card
capacity: total VRAM minus a fixed max(2 GiB, 9%) reserve. That reserve
covers a light desktop only. On an RTX 5090 with 4.5-6.3 GiB held by a
browser and other apps, Qwen3.8 27B booted at 216K, overflowed the card,
and Windows paged part of it to host memory without an error: decode
fell from ~90 to ~24 tok/s.
hardware.launch_budget subtracts what other programs hold now (device
total minus nvidia-smi free, minus our own server's footprint) plus
1 GiB of headroom. fit_to_free_memory narrows a resident plan's window
to that budget, down to the floor and never further: weights are never
moved to the CPU on a live reading, which once pinned a fitting model
to the CPU while the reading still counted the outgoing server.
- Boot: preset generation narrows every window through it.
- Idle: every 30 s with no model holding memory, the presets are
re-planned and the router re-reads them (GET /models?reload=1), so a
later on-demand load gets a window for the current desktop. The file
is never rewritten while a model is loaded, because reload unloads a
loaded model whose flags changed.
- Growth: the next rung must fit beside other programs, counting the
growing model's own memory as free.
Recommendations, quant selection and catalog pricing keep the capacity
budget. Unified-memory machines and failed probes keep today's plan.
The new tests pin the catalog 27B's window from 0 to 8.5 GiB of other
programs. Every boot, idle, growth and residency-cap test records its
probe_budget calls and fails on any made without planning=True, including
calls inside paths that swallow exceptions.
The package-manager switch (#102765) made engine detection read only PM's
store and pinned every llama.cpp backend to b10362. A machine with an
engine under runtimes/llamacpp/b<tag>/<backend>/ reported no engine, so
the local models pane showed one-click setup with models already on disk.
installed_engine() now moves the newest pre-PM install into the store
once. The old manifest's archive digests become the PM identity, the
package verifier runs llama-server --version, and os.rename puts the
directory in place before facts.json records it. A store that already
holds an engine is left alone, a store lock held by another PM operation
defers the move, and an install that fails verification stays where it
is for the rest of the process.
The lock returns to b10964 on all five backends, the default before the
PM switch. With b10362 pinned, the pane offered a downgrade from the
moved engine as an update. Upstream renamed the ROCm archives to 10.0 at
b10767, and the hip assets follow. Only the CUDA build is measured at
b10964.
Models used to live at <profile home>/models (adb2fdbf5c) before the managed
runtime moved them to the machine-scoped <root>/models (43e67d872f). GGUFs
staged under a named profile's old dir silently stopped being served.
adopt_legacy_models() renames them (and their assets/) into the current dirs,
so everything downstream keeps reading one directory: listing, presets,
delete, the --models-dir fallback. It runs at boot, before the "anything
staged?" check, and in the Local Models status route so the pane lists them
while the runtime is off.
os.rename only: instant within a filesystem, while shutil.move would silently
copy tens of GB across devices at session start; a cross-device dir stays put
with a warning. An existing destination name is never replaced.
On Windows, subprocess text=True without an explicit encoding decodes
child output with the ANSI code page (e.g. 'gbk'); non-ASCII bytes then
raise UnicodeDecodeError inside subprocess._readerthread, killing the
Hermes backend before it becomes ready and surfacing as the desktop boot
timeout.
Sweep every hermes_cli text=True subprocess call to encoding='utf-8',
errors='replace', and add an AST-based regression test that fails when a
future text-mode call omits the encoding.
Fixes#55658
utf-8-sig exists to tolerate BOMs that Windows tooling adds to files
users edit. /proc and /sys files are generated by the Linux kernel, never
BOM'd and absent on Windows, so -sig there only muddies the read/write
policy. Switch every literal /proc/ and /sys/ read to utf-8 and teach
the footgun read rule that string literals starting with /proc/ or
/sys/ are exempt (user-edited files keep utf-8-sig).
getconf _AVPHYS_PAGES counts page cache as used, so the Desktop statusbar
overreports Linux RAM. Read MemAvailable from /proc/meminfo, fall back to
MemFree, then the existing getconf probe. A reported 0 is a real available
value, not a missing field.
Fixes#102252
Symptom: a 22 GB catalog download ended with "[WinError 5] Access is denied:
'...\models\<model>.part'" even though the complete .gguf was on disk, and a
retry then found it and finished in seconds.
Root cause: download_file published the finished .part with shutil.move. On
Windows the just-closed file is commonly still open to an antivirus or
indexing scan, so os.rename fails with a permission error; shutil.move then
silently falls back to copy2 + unlink, which duplicates the whole file
(many seconds of disk churn and double the space) and finally reports the
leftover's failed delete as the download's failure. The except-path cleanup
could also replace the real error with its own unlink failure.
Fix: binaries.replace_when_released renames with os.replace and retries a
PermissionError with backoff through a bounded window (60 s), raising a
plain-language error if the hold outlasts it; there is no copy fallback.
Both the model download and the runtime-archive download publish through it,
and both cleanups suppress a failed leftover removal so the original error is
what the user sees. The job narrates the wait as "Finishing" so the bar does
not sit dead at 100%.
Tests: helper retries a transient refusal and lands the file; helper gives up
with the plain-language message chained to the OS error; route-level download
survives two refused renames with no .part left behind; a stuck leftover does
not mask the rename failure.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Residency was bounded by `local_runtime.models_max` alone, so a second model was admitted
against an already-full card. On Windows/WDDM that over-commit is not refused: the
allocation is paged to host memory and the child decodes at about a third of its speed for
the rest of its life — no error, no UI hint, and ejecting the incumbent afterwards does not
repair it (only a clean reload does).
The router now gets a cap priced from the same capacity budget the presets are priced
against: the count rises above one only while the largest staged model still fits twice, so
any pair of staged models is inside the card by construction and llama.cpp evicts its LRU
before an incoming child allocates. `models_max` stays a ceiling (a user's smaller number is
honoured) and an unpriceable budget or model keeps today's count.
Refs #116078
(cherry picked from commit d7abad7984d16d23ab2442b92d623b2270776586)
Review follow-up: the boot lock's mkdir/os.open ran in the context manager's
__enter__, outside the body's try/except, so a read-only runtimes dir or a
foreign-owned boot.lock raised OSError out of ensure_local_runtime — the
function main documented as never raising into session start. Catch OSError
around acquisition, warn, and yield unlocked (the contention-timeout shape).
Two Hermes backends starting within the same second (e.g. default +
non-default profile) each see no server.json yet and each spawn a
llama-server router on the stable port — the existing per-process
_SUPERVISOR singleton can't rule this out since every process starts
with its own copy. Wrap the state-check/spawn sequence in an OS-held
file lock (fcntl/msvcrt, mirroring managed_uv.py's install lock) so the
second caller waits for the first to publish its state and adopts it
instead of racing to spawn its own.
Fixes#116682
spawn_server also backs _subprocess_compat.bounded_probe_run (git probes,
PowerShell/tasklist scans, the update venv probe) and the scrub overrode an
explicit env= too, so every probe lost *_TOKEN/*_API_KEY/PASSWORD* (GH_TOKEN for
a gh credential helper, HF_TOKEN for a gated pull). Apply server_child_env in
supervisor._spawn and leave spawn_server's caller environment untouched.
The managed router child was spawned with the full process environment, so
every provider/tool credential (`*_API_KEY`, `*_TOKEN`, `*_SECRET`,
`*PASSWORD*`, `*_CREDENTIALS`) reached a native binary that talks to nobody
but us. On Windows that was not just a leak: the bundled libomp.dll died with
STATUS_HEAP_CORRUPTION (0xC0000374) during OpenMP initialisation with one
`*_API_KEY` present in the inherited Desktop environment and loaded fine
with only that variable removed — confirmed twice on a model-free
`ctypes.CDLL(libomp.dll).omp_get_max_threads()` probe (#116109).
Scrub those names at the single spawn boundary (`spawn_server`), leaving
PATH, CUDA_*/HSA_*/OMP_* and everything else untouched. The upstream
OpenMP defect itself is not ours to fix; keeping secrets out of the native
child is correct regardless and removes the trigger.
Keeps three invariants across #116156 (@fangliquanflq) and #116292 (@Finn763):
a configured providers.llamacpp entry wins over managed-server detection on the
runtime path; an external llama-server on a local_runtime.detect_ports port is
what /model > Local (switch_model with the picker's provider id) resolves to;
a staged GGUF alone keeps the picker's id resolvable so the runtime seam, not
the provider gate, reports a missing server.
Drops the endpoint.py config auto-load from #116292: both production callers of
resolve_llamacpp_endpoint() now pass the loaded config (#116156), so loading it
again inside the resolver was defense-in-depth. Drops the change-detector test
on the forwarded config object and the redundant negative control.
Docs: local-models page now shows detect_ports and the providers.llamacpp
override for a llama-server the user runs on a fixed port.
The /model -> local/ flow offers the 'llamacpp' row from staged GGUFs alone,
but resolve_provider_full only admitted that id behind a live endpoint, so
selecting a staged model died with "Unknown provider 'llamacpp'" before the
runtime seam could start or attach a server.
- providers.py: LLAMACPP_PROVIDER_ID/ALIASES are the single definition; the
llamacpp rung resolves when a server is reachable OR a model is staged.
- inventory.py: the picker row's slug comes from that definition.
- endpoint.py: honour local_runtime.detect_ports (every provider caller
invoked the resolver bare, leaving the documented knob dead).
- model_switch.py: a local-runtime alias failure surfaces the seam's own
message instead of an API-key hint.
Merges from origin/main resolved several files by taking main's side, which
re-inlined code paths this branch had already moved to PM. At HEAD that left:
- plugins/memory/hindsight: ImportError at module load (`_export_port_health_
grace_timeout` and `_MIN_CLIENT_VERSION` no longer exist) — the whole plugin
failed to import. Restored the side-env daemon client, folded main's
fail-closed profile-env rewrite guard into it, and dropped the version
auto-upgrade dance (the extra is pinned).
- mem0 / honcho / google_chat / memory_provider_migration / plugins_cmd_catalog:
imported the deleted tools.lazy_deps and hermes_cli.plugin_python_deps
modules; routed through pm.ensure_import / pm.sync_venv / pm.plugins_state.
scripts/ci/check_lazy_deps_imports.py exits 0 again.
- tools/skills_hub: the PLUGIN-COMPAT __getattr__ had been moved above
`_plugin_compat_prev_getattr = __getattr__`, so `SKILLS_DIR` recursed forever.
- tools/tirith_security: `import time` dropped; the circuit breaker NameError'd.
- hermes_cli/local_runtime/binaries: json/platform/os used without imports;
the manifest.json scan was for a layout PM no longer writes — the boot gate
now asks PM for an installed engine.
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:
- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
scalar snapshot reads as empty / returns the existing 5000 error instead of
violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
_combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
unchanged.
Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.
(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
* test(local-runtime): pin resume to the live managed llama.cpp port
A session that stored last boot's loopback URL must not keep the client
on a dead ephemeral port after the supervisor moves.
* fix(local-runtime): follow the live managed llama.cpp port on resume
Sessions persist last boot's loopback URL, so a supervisor port change
left the client on a dead endpoint. Drop that snapshot for llamacpp
and keep the live supervisor URL.
* fix(tui_gateway): drop the llamacpp snapshot URL at one seam
_resolve_agent_model_runtime already discards a persisted base_url when the
resolution came from the local runtime; the second blank in
_stored_session_runtime_overrides (wrapped in a try/except around an import
and a string compare) duplicated it.
* fix(cli): keep a launch-time --base-url on a same-provider llamacpp resume
Re-resolving the managed endpoint is for the snapshot URL a session persisted;
an explicit --base-url for the provider the session already ran on is user
intent and stays in charge. Also hand target_model to the resolver like the
provider-changed branch does.
* fix(gateway): rehydrated llamacpp overrides follow the live managed port
Same bug class as the CLI and TUI resume paths: after a gateway restart the
persisted /model override kept last boot's loopback URL over the freshly
resolved managed endpoint, so a supervisor that came back on an ephemeral port
(18434 busy) left the session on connection errors.
* docs(local-runtime): port-fallback warning no longer asks for a model re-pick
---------
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
Keep the two tests that were red on origin/main (probe failure keeps the clock and
unloads once telemetry recovers; confirmed busy still resets). Drop the is_idle() bool
contract test — it pins behaviour that did not change — and fold the probe-failure log
call onto two lines.
is_idle() treated any /slots or /metrics probe error as "busy", so sweep_idle()
popped the idle clock on every failed probe — one transient failure per sweep
and a resident model never reached IDLE_UNLOAD_S (21 GB pinned for hours with
zero requests).
Split the probe into a tri-state: _probe_idle() returns True (confirmed idle),
False (confirmed busy), None (probe failed). The sweep now resets the clock
only on a confirmed busy sighting; a failed probe keeps the existing clock and
logs at INFO, and unload still requires confirmed idleness past the threshold.
The public is_idle() bool contract is unchanged (probe failure reads as not
idle), so no caller outside the sweep ever unloads on a telemetry hiccup.
Closes#111154
llama.cpp b10964 (the pinned build) renamed --no-webui to --no-ui, so the
router refused to start with an unknown-argument error. The direct-I/O flag
is already selected per-build by _direct_io_args on main, so this reduces
to the one remaining rename and asserts it in the existing spawn test.
Refs #111323
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
Stop recommending a system-RAM spill when no curated model fits resident.
Preserve explicit model selection and the existing resident quality/speed
ranking, including the separate unified-memory policy.
Require a recommendation for automatic quickstart, expose Browse when
none exists, and rename Configure to Let me choose. Keep policy copy and
reason keys consistent across the four translated local-model sections.
Cover automatic refusal and explicit spilled setup against one budget,
plus the Browse, Download and Use interactions in the desktop pane.