Run models locally as a first-class provider. The CLI grows a managed llama.cpp runtime (engine install, model download, server supervision); the desktop app grows the full setup and management story on top of it. GUI surfaces ship behind the desktop --local launch flag (hermes desktop --local, or the flag on the packaged app); backend routes and the CLI are always live. Runtime (hermes_cli/local_runtime/): - curated GGUF catalog with per-machine variant selection: hardware probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice by context window - derived recommendation: quality-ranked picks gated by a predicted decode-speed floor, bandwidth-aware on unified memory; the decision table is pinned as a test (pick AND reason per memory class), and the Recommended badge explains its pick in a tooltip fed by the resolver's actual branch - engine install + model download with resumable split parts, cumulative plan-level progress, and staged-model integrity (a split GGUF counts only when every part is present) - server supervision: spawn/adopt/stop, router mode with per-model load progress relayed over SSE, abandoned-request cleanup Desktop: - Settings -> Providers -> Local models: one-click quickstart (install engine, download the recommended model, boot) plus per-model download/ activate/eject, fit-ranked catalog with context pills - model pickers (composer dropdown + Cmd+K) show staged local models, in-flight downloads as live progress rows, and load-into-memory bars - local-setup campaign tip for eligible hardware; System resources statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends - friendly dead-server errors, and failed agent builds retry on the next send instead of wedging the session Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
52 lines
2.0 KiB
Python
52 lines
2.0 KiB
Python
"""Launch and growth decisions must price against CAPACITY, not live-free
|
|
VRAM. Both execute through a server bounce — the outgoing instance's memory
|
|
is freed before the new one loads — so a probe that reads the predecessor's
|
|
(or the grown model's own) residency as 'gone' vetoes configurations that
|
|
genuinely fit. Symptom when this regresses: a model the pane promised
|
|
'144K on GPU' launches with its weights pinned to CPU and single-digit
|
|
tokens/s while the card sits 60% empty."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import ast
|
|
import inspect
|
|
|
|
|
|
def _planning_probe_calls(source: str) -> list[bool]:
|
|
"""Every probe_budget(...) call's planning= value in the source."""
|
|
tree = ast.parse(source)
|
|
out = []
|
|
for node in ast.walk(tree):
|
|
if (isinstance(node, ast.Call)
|
|
and getattr(node.func, "id", getattr(node.func, "attr", ""))
|
|
== "probe_budget"):
|
|
planning = any(
|
|
kw.arg == "planning"
|
|
and isinstance(kw.value, ast.Constant)
|
|
and kw.value.value is True
|
|
for kw in node.keywords)
|
|
out.append(planning)
|
|
return out
|
|
|
|
|
|
def test_bootstrap_presets_price_against_capacity():
|
|
import hermes_cli.local_runtime.bootstrap as bootstrap
|
|
|
|
calls = _planning_probe_calls(inspect.getsource(bootstrap))
|
|
assert calls, "bootstrap no longer probes a budget? update this test"
|
|
assert all(calls), (
|
|
"bootstrap prices launch decisions against live-free VRAM; a "
|
|
"restart/refresh probes while the outgoing server still holds the "
|
|
"card, pinning fitting models to CPU")
|
|
|
|
|
|
def test_growth_refit_prices_against_capacity():
|
|
import hermes_cli.local_runtime.growth as growth
|
|
|
|
calls = _planning_probe_calls(inspect.getsource(growth))
|
|
assert calls, "growth no longer probes a budget? update this test"
|
|
assert all(calls), (
|
|
"growth re-fits against live-free VRAM; the grown model's own "
|
|
"residency reads as unavailable and vetoes rungs that fit the "
|
|
"post-bounce card")
|