Commit Graph

38 Commits

Author SHA1 Message Date
teknium1
56b7a0fb37 fix(local-runtime): an unwritable boot lock degrades to an unlocked boot, never an OSError
Review follow-up: the boot lock's mkdir/os.open ran in the context manager's
__enter__, outside the body's try/except, so a read-only runtimes dir or a
foreign-owned boot.lock raised OSError out of ensure_local_runtime — the
function main documented as never raising into session start. Catch OSError
around acquisition, warn, and yield unlocked (the contention-timeout shape).
2026-09-20 15:53:25 -07:00
chelsealong
4352a7d053 fix(local-runtime): serialize managed runtime boot across processes
Two Hermes backends starting within the same second (e.g. default +
non-default profile) each see no server.json yet and each spawn a
llama-server router on the stable port — the existing per-process
_SUPERVISOR singleton can't rule this out since every process starts
with its own copy. Wrap the state-check/spawn sequence in an OS-held
file lock (fcntl/msvcrt, mirroring managed_uv.py's install lock) so the
second caller waits for the first to publish its state and adopts it
instead of racing to spawn its own.

Fixes #116682
2026-09-20 15:53:25 -07:00
teknium1
e56c229fed fix(local-runtime): name the MXFP4 row in the GGML size table 2026-09-20 15:29:34 -07:00
fangliquan
ae908a93ba fix(local-runtime): support MXFP4 GGUF tensors 2026-09-20 15:29:34 -07:00
teknium1
c6c595508d fix: scrub credentials at the llama-server spawn only, not the generic probe spawner (review follow-up)
spawn_server also backs _subprocess_compat.bounded_probe_run (git probes,
PowerShell/tasklist scans, the update venv probe) and the scrub overrode an
explicit env= too, so every probe lost *_TOKEN/*_API_KEY/PASSWORD* (GH_TOKEN for
a gh credential helper, HF_TOKEN for a gated pull). Apply server_child_env in
supervisor._spawn and leave spawn_server's caller environment untouched.
2026-09-20 11:50:48 -07:00
teknium1
63c16221fa fix(local-models): llama-server never inherits credential-shaped environment variables
The managed router child was spawned with the full process environment, so
every provider/tool credential (`*_API_KEY`, `*_TOKEN`, `*_SECRET`,
`*PASSWORD*`, `*_CREDENTIALS`) reached a native binary that talks to nobody
but us. On Windows that was not just a leak: the bundled libomp.dll died with
STATUS_HEAP_CORRUPTION (0xC0000374) during OpenMP initialisation with one
`*_API_KEY` present in the inherited Desktop environment and loaded fine
with only that variable removed — confirmed twice on a model-free
`ctypes.CDLL(libomp.dll).omp_get_max_threads()` probe (#116109).

Scrub those names at the single spawn boundary (`spawn_server`), leaving
PATH, CUDA_*/HSA_*/OMP_* and everything else untouched. The upstream
OpenMP defect itself is not ours to fix; keeping secrets out of the native
child is correct regardless and removes the trigger.
2026-09-20 11:50:48 -07:00
teknium1
34fb4ab482 fix: trim llamacpp salvage to invariant tests, drop redundant config load, document fixed-port servers
Keeps three invariants across #116156 (@fangliquanflq) and #116292 (@Finn763):
a configured providers.llamacpp entry wins over managed-server detection on the
runtime path; an external llama-server on a local_runtime.detect_ports port is
what /model > Local (switch_model with the picker's provider id) resolves to;
a staged GGUF alone keeps the picker's id resolvable so the runtime seam, not
the provider gate, reports a missing server.

Drops the endpoint.py config auto-load from #116292: both production callers of
resolve_llamacpp_endpoint() now pass the loaded config (#116156), so loading it
again inside the resolver was defense-in-depth. Drops the change-detector test
on the forwarded config object and the redundant negative control.

Docs: local-models page now shows detect_ports and the providers.llamacpp
override for a llama-server the user runs on a fixed port.
2026-09-20 10:03:32 -07:00
Hermes
04f37f74b0 fix(local): the picker's Local row id and the provider resolver share one definition (#116249)
The /model -> local/ flow offers the 'llamacpp' row from staged GGUFs alone,
but resolve_provider_full only admitted that id behind a live endpoint, so
selecting a staged model died with "Unknown provider 'llamacpp'" before the
runtime seam could start or attach a server.

- providers.py: LLAMACPP_PROVIDER_ID/ALIASES are the single definition; the
  llamacpp rung resolves when a server is reachable OR a model is staged.
- inventory.py: the picker row's slug comes from that definition.
- endpoint.py: honour local_runtime.detect_ports (every provider caller
  invoked the resolver bare, leaving the documented knob dead).
- model_switch.py: a local-runtime alias failure surfaces the seam's own
  message instead of an API-key hint.
2026-09-20 10:03:32 -07:00
beardthelion
9ac4e76e0f fix(tools,tui_gateway,cli,plugins): parseable non-dict JSON no longer crashes the remaining file scans
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:

- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
  envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
  scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
  the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
  scalar snapshot reads as empty / returns the existing 5000 error instead of
  violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
  dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
  _combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
  unchanged.

Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.

(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
2026-09-18 09:19:04 -07:00
Austin Pickett
1792e8bf5f fix(local-runtime): resumed llamacpp sessions follow the live managed port on every surface (#114336)
* test(local-runtime): pin resume to the live managed llama.cpp port

A session that stored last boot's loopback URL must not keep the client
on a dead ephemeral port after the supervisor moves.

* fix(local-runtime): follow the live managed llama.cpp port on resume

Sessions persist last boot's loopback URL, so a supervisor port change
left the client on a dead endpoint. Drop that snapshot for llamacpp
and keep the live supervisor URL.

* fix(tui_gateway): drop the llamacpp snapshot URL at one seam

_resolve_agent_model_runtime already discards a persisted base_url when the
resolution came from the local runtime; the second blank in
_stored_session_runtime_overrides (wrapped in a try/except around an import
and a string compare) duplicated it.

* fix(cli): keep a launch-time --base-url on a same-provider llamacpp resume

Re-resolving the managed endpoint is for the snapshot URL a session persisted;
an explicit --base-url for the provider the session already ran on is user
intent and stays in charge. Also hand target_model to the resolver like the
provider-changed branch does.

* fix(gateway): rehydrated llamacpp overrides follow the live managed port

Same bug class as the CLI and TUI resume paths: after a gateway restart the
persisted /model override kept last boot's loopback URL over the freshly
resolved managed endpoint, so a supervisor that came back on an ephemeral port
(18434 busy) left the session on connection errors.

* docs(local-runtime): port-fallback warning no longer asks for a model re-pick

---------

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-09-17 14:24:53 -04:00
Jeffrey Quesnelle
6005aa1fd9 Merge pull request #112443 from RyanUnderhill/codex/fix-runtime-archive-cache
fix(local-runtime): reject incomplete downloads before caching
2026-09-16 21:32:45 -04:00
emozilla
3c56591788 fix(local-runtime): isolate download staging files 2026-09-16 14:32:54 -04:00
Ryan Hill
ea09e2e401 fix(local-runtime): limit archive fix to incomplete downloads 2026-09-16 09:22:32 -07:00
Sora-bluesky
9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
Ryan Hill
745e3264ff fix(local-runtime): reject incomplete downloads and repair corrupt cache 2026-09-15 16:56:22 -07:00
teknium1
c1b7f28693 test(local-runtime): trim idle-probe salvage to two invariants
Keep the two tests that were red on origin/main (probe failure keeps the clock and
unloads once telemetry recovers; confirmed busy still resets). Drop the is_idle() bool
contract test — it pins behaviour that did not change — and fold the probe-failure log
call onto two lines.
2026-09-15 05:32:58 -07:00
Konstantin Khlopkov
b83ee209d9 fix(local_runtime): keep the idle clock across failed telemetry probes
is_idle() treated any /slots or /metrics probe error as "busy", so sweep_idle()
popped the idle clock on every failed probe — one transient failure per sweep
and a resident model never reached IDLE_UNLOAD_S (21 GB pinned for hours with
zero requests).

Split the probe into a tri-state: _probe_idle() returns True (confirmed idle),
False (confirmed busy), None (probe failed). The sweep now resets the clock
only on a confirmed busy sighting; a failed probe keeps the existing clock and
logs at INFO, and unload still requires confirmed idleness past the threshold.
The public is_idle() bool contract is unchanged (probe failure reads as not
idle), so no caller outside the sweep ever unloads on a telemetry hiccup.

Closes #111154
2026-09-15 05:32:58 -07:00
Hukla
b039e679ea fix(local-runtime): spawn llama-server with --no-ui
llama.cpp b10964 (the pinned build) renamed --no-webui to --no-ui, so the
router refused to start with an unknown-argument error. The direct-I/O flag
is already selected per-build by _direct_io_args on main, so this reduces
to the one remaining rename and asserts it in the existing spawn test.

Refs #111323
2026-09-15 05:31:16 -07:00
emozilla
e9e363c856 feat(local-runtime): prefer llama.cpp b10964 2026-09-14 21:49:13 -04:00
emozilla
ef1cfbbbf9 fix(local-runtime): contain Windows runtimes and safely stop orphans 2026-09-14 12:56:59 -04:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
emozilla
5764ae3168 fix(local-models): keep automatic recommendations GPU-resident
Stop recommending a system-RAM spill when no curated model fits resident.
Preserve explicit model selection and the existing resident quality/speed
ranking, including the separate unified-memory policy.

Require a recommendation for automatic quickstart, expose Browse when
none exists, and rename Configure to Let me choose. Keep policy copy and
reason keys consistent across the four translated local-model sections.

Cover automatic refusal and explicit spilled setup against one budget,
plus the Browse, Download and Use interactions in the desktop pane.
2026-09-11 12:43:21 -04:00
emozilla
0316d3d404 fix(local-runtime): unify memory accounting and effective-window MTP
Price weights, context, runtime, projector and batch overhead consistently
across catalog admission, initial launch, growth and restored windows.
Keep MTP and the larger window when lean batches avoid unnecessary spill.

Admit optional external drafts only when their complete footprint fits.
Use preset-only model discovery so refused files cannot autoload, and
preserve refusal/spill decisions atomically for desktop status read-back.

Add regression coverage for complete-footprint boundaries, MTP restarts,
growth admission, draft budgets and placement status transitions.

Builds on the overhead-accounting contribution in #102993 and the
restored-window MTP contribution in #106897. Does not adopt the 40%
host-RAM reserve or resolve the remaining requests in #102865/#106895.

Co-authored-by: infinitycrew39 <infinitycrew39@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-11 01:40:43 -04:00
Gille
485aaf69b7 fix(local-runtime): find nvidia-smi in WSL driver path 2026-09-06 23:34:20 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
c8d401c79a refactor(local_runtime): compact module docstrings (rules kept, prose trimmed) 2026-09-02 21:47:26 -07:00
Teknium
723d169e68 refactor(local_runtime): join one-name-per-line imports and wrapped calls; guard-clause reflows in sweep_idle/_reap_orphaned_children/physics_check 2026-09-02 21:34:34 -07:00
Teknium
7c7d967b4c refactor(local_runtime): catalog.entry_for_model replaces 2 hit-unpack sites; tighten _host_os_arch 2026-09-02 21:10:03 -07:00
Teknium
090fd9c6e9 refactor(local_runtime): collapse swallow-and-default try/except ladders to contextlib.suppress; single struct reader in read_gguf_header 2026-09-02 21:05:10 -07:00
Teknium
9aa1f66484 refactor(local_runtime): GGUFHeader int accessors via one property factory; supervisor best-effort child ops via _quiet; compact __init__ re-exports 2026-09-02 20:58:40 -07:00
Teknium
a4a7ab6286 refactor(local_runtime): unify managed-router GET (endpoint.managed_root/managed_get_json) across load_progress+capabilities; shared kv() closure in initial_window; GGUF scalar-type table inlined; compact rationale prose across the package 2026-09-02 20:53:51 -07:00
Teknium
0244f7fa7e refactor(local_runtime): unify split-GGUF id/scan helpers (gguf.model_id_from_stem, bootstrap.staged_in, binaries.manifest_verified); split generate_presets into per-model phases; drop dead /health probe in _state_endpoint (pid was already the sole verdict) 2026-09-02 20:35:53 -07:00
Teknium
cb1b44aacd refactor(local_runtime): compact supervisor/catalog/hardware — shared _RESIDENT status tuple, collapsed integrated verdict, tightened rationale comments 2026-09-02 20:22:10 -07:00
Teknium
4d1880e0bf fix(integration): restore subprocess encoding/stdin guards dropped in simplification
Simplification workers collapsed subprocess call sites into shared kwargs
helpers and dropped the Windows/TUI safety kwargs on the way:

- encoding='utf-8', errors='replace' restored on text=True runs in
  copilot_acp_client, hermes_cli/setup (vercel install), managed_uv
  (codesign steps), local_runtime/hardware._stdout, a2a adapter.
- stdin=subprocess.DEVNULL restored on copilot probe, verify/runner
  _SUBPROCESS_KW, iron_proxy._run, google_meet playwright/system_profiler,
  simplex convert, whatsapp _RUN_TEXT, mem0 ollama serve Popen.
  google_meet sudo/brew install keeps inherited stdin (user-confirmed,
  may prompt) — marked noqa: subprocess-stdin.
- Windows-safe SIGKILL: getattr(signal, 'SIGKILL', SIGTERM) in
  verify/runner; photon _kill call re-marked windows-footgun: ok
  (unreachable on win32).
- scripts/check_subprocess_stdin.py now recognizes **kwargs splats
  (**_KW / **_kw(...)) ONLY when the same-file definition provably sets
  stdin= — covers tui_gateway _capture_run_kwargs/run_kw. Parity test added.
2026-09-02 16:36:10 -07:00
Teknium
3ffd44acd3 refactor(hclib): models/runtime — models, inventory, runtime_provider, provider catalog, local_runtime, banner 2026-09-02 14:44:48 -07:00
emozilla
43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00