Files
hermes-agent/hermes_cli/local_runtime
emozilla 20700f298a fix(local-runtime): size launch windows beside other programs' GPU memory
Boot presets and in-session growth priced the context window from card
capacity: total VRAM minus a fixed max(2 GiB, 9%) reserve. That reserve
covers a light desktop only. On an RTX 5090 with 4.5-6.3 GiB held by a
browser and other apps, Qwen3.8 27B booted at 216K, overflowed the card,
and Windows paged part of it to host memory without an error: decode
fell from ~90 to ~24 tok/s.

hardware.launch_budget subtracts what other programs hold now (device
total minus nvidia-smi free, minus our own server's footprint) plus
1 GiB of headroom. fit_to_free_memory narrows a resident plan's window
to that budget, down to the floor and never further: weights are never
moved to the CPU on a live reading, which once pinned a fitting model
to the CPU while the reading still counted the outgoing server.

- Boot: preset generation narrows every window through it.
- Idle: every 30 s with no model holding memory, the presets are
  re-planned and the router re-reads them (GET /models?reload=1), so a
  later on-demand load gets a window for the current desktop. The file
  is never rewritten while a model is loaded, because reload unloads a
  loaded model whose flags changed.
- Growth: the next rung must fit beside other programs, counting the
  growing model's own memory as free.

Recommendations, quant selection and catalog pricing keep the capacity
budget. Unified-memory machines and failed probes keep today's plan.
The new tests pin the catalog 27B's window from 0 to 8.5 GiB of other
programs. Every boot, idle, growth and residency-cap test records its
probe_budget calls and fails on any made without planning=True, including
calls inside paths that swallow exceptions.
2026-09-27 22:41:39 -04:00
..
…