Files
hermes-agent/tests/hermes_cli/test_budget_source.py
emozilla 860ee24991 fix(local-runtime): launch and growth decisions price against capacity, not live-free VRAM
A model the pane promised '144K fully on GPU' could launch with its
weights pinned to CPU — single-digit tokens/s, high CPU, the card 60%
empty. Cause: preset generation runs during boot and refresh, while the
OUTGOING server instance still holds the card. The live-free probe read
the predecessor's residency as memory that doesn't exist, the fit
concluded 'weights don't fit', and the spill placement did exactly what
it was told: hold the 64K floor and pin FFN weights to host.

Both launch (bootstrap preset generation) and growth re-fit now price
against the capacity budget — total device memory minus margin — because
both execute through a server bounce: the old instance's memory is freed
before the new one loads a byte. Growth had the same bug in a nastier
costume: the model being grown vetoed its own next rung by reading its
own residency as unavailable.

Live-free remains the right probe for telemetry (hardware route,
statusbar), which reports the present, not a post-bounce future.

New contract test asserts every probe_budget call in bootstrap and
growth passes planning=True, with the symptom documented in the
assertion message.
2026-08-14 12:02:28 -04:00

52 lines
2.0 KiB
Python

"""Launch and growth decisions must price against CAPACITY, not live-free
VRAM. Both execute through a server bounce — the outgoing instance's memory
is freed before the new one loads — so a probe that reads the predecessor's
(or the grown model's own) residency as 'gone' vetoes configurations that
genuinely fit. Symptom when this regresses: a model the pane promised
'144K on GPU' launches with its weights pinned to CPU and single-digit
tokens/s while the card sits 60% empty."""
from __future__ import annotations
import ast
import inspect
def _planning_probe_calls(source: str) -> list[bool]:
"""Every probe_budget(...) call's planning= value in the source."""
tree = ast.parse(source)
out = []
for node in ast.walk(tree):
if (isinstance(node, ast.Call)
and getattr(node.func, "id", getattr(node.func, "attr", ""))
== "probe_budget"):
planning = any(
kw.arg == "planning"
and isinstance(kw.value, ast.Constant)
and kw.value.value is True
for kw in node.keywords)
out.append(planning)
return out
def test_bootstrap_presets_price_against_capacity():
import hermes_cli.local_runtime.bootstrap as bootstrap
calls = _planning_probe_calls(inspect.getsource(bootstrap))
assert calls, "bootstrap no longer probes a budget? update this test"
assert all(calls), (
"bootstrap prices launch decisions against live-free VRAM; a "
"restart/refresh probes while the outgoing server still holds the "
"card, pinning fitting models to CPU")
def test_growth_refit_prices_against_capacity():
import hermes_cli.local_runtime.growth as growth
calls = _planning_probe_calls(inspect.getsource(growth))
assert calls, "growth no longer probes a budget? update this test"
assert all(calls), (
"growth re-fits against live-free VRAM; the grown model's own "
"residency reads as unavailable and vetoes rungs that fit the "
"post-bounce card")