Files
hermes-agent/hermes_cli/local_runtime
finn763 eaa5523f11 fix(local-runtime): cap resident models by the hardware budget, not a count
Residency was bounded by `local_runtime.models_max` alone, so a second model was admitted
against an already-full card. On Windows/WDDM that over-commit is not refused: the
allocation is paged to host memory and the child decodes at about a third of its speed for
the rest of its life — no error, no UI hint, and ejecting the incumbent afterwards does not
repair it (only a clean reload does).

The router now gets a cap priced from the same capacity budget the presets are priced
against: the count rises above one only while the largest staged model still fits twice, so
any pair of staged models is inside the card by construction and llama.cpp evicts its LRU
before an incoming child allocates. `models_max` stays a ceiling (a user's smaller number is
honoured) and an unpriceable budget or model keeps today's count.

Refs #116078

(cherry picked from commit d7abad7984d16d23ab2442b92d623b2270776586)
2026-09-20 20:42:05 -07:00
..
…
…