Files
hermes-agent/tests/hermes_cli/test_budget_source.py
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00

52 lines
2.0 KiB
Python

"""Launch and growth decisions must price against CAPACITY, not live-free
VRAM. Both execute through a server bounce — the outgoing instance's memory
is freed before the new one loads — so a probe that reads the predecessor's
(or the grown model's own) residency as 'gone' vetoes configurations that
genuinely fit. Symptom when this regresses: a model the pane promised
'144K on GPU' launches with its weights pinned to CPU and single-digit
tokens/s while the card sits 60% empty."""
from __future__ import annotations
import ast
import inspect
def _planning_probe_calls(source: str) -> list[bool]:
"""Every probe_budget(...) call's planning= value in the source."""
tree = ast.parse(source)
out = []
for node in ast.walk(tree):
if (isinstance(node, ast.Call)
and getattr(node.func, "id", getattr(node.func, "attr", ""))
== "probe_budget"):
planning = any(
kw.arg == "planning"
and isinstance(kw.value, ast.Constant)
and kw.value.value is True
for kw in node.keywords)
out.append(planning)
return out
def test_bootstrap_presets_price_against_capacity():
import hermes_cli.local_runtime.bootstrap as bootstrap
calls = _planning_probe_calls(inspect.getsource(bootstrap))
assert calls, "bootstrap no longer probes a budget? update this test"
assert all(calls), (
"bootstrap prices launch decisions against live-free VRAM; a "
"restart/refresh probes while the outgoing server still holds the "
"card, pinning fitting models to CPU")
def test_growth_refit_prices_against_capacity():
import hermes_cli.local_runtime.growth as growth
calls = _planning_probe_calls(inspect.getsource(growth))
assert calls, "growth no longer probes a budget? update this test"
assert all(calls), (
"growth re-fits against live-free VRAM; the grown model's own "
"residency reads as unavailable and vetoes rungs that fit the "
"post-bounce card")