Commit Graph

18 Commits

Author SHA1 Message Date
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
emozilla
5764ae3168 fix(local-models): keep automatic recommendations GPU-resident
Stop recommending a system-RAM spill when no curated model fits resident.
Preserve explicit model selection and the existing resident quality/speed
ranking, including the separate unified-memory policy.

Require a recommendation for automatic quickstart, expose Browse when
none exists, and rename Configure to Let me choose. Keep policy copy and
reason keys consistent across the four translated local-model sections.

Cover automatic refusal and explicit spilled setup against one budget,
plus the Browse, Download and Use interactions in the desktop pane.
2026-09-11 12:43:21 -04:00
emozilla
0316d3d404 fix(local-runtime): unify memory accounting and effective-window MTP
Price weights, context, runtime, projector and batch overhead consistently
across catalog admission, initial launch, growth and restored windows.
Keep MTP and the larger window when lean batches avoid unnecessary spill.

Admit optional external drafts only when their complete footprint fits.
Use preset-only model discovery so refused files cannot autoload, and
preserve refusal/spill decisions atomically for desktop status read-back.

Add regression coverage for complete-footprint boundaries, MTP restarts,
growth admission, draft budgets and placement status transitions.

Builds on the overhead-accounting contribution in #102993 and the
restored-window MTP contribution in #106897. Does not adopt the 40%
host-RAM reserve or resolve the remaining requests in #102865/#106895.

Co-authored-by: infinitycrew39 <infinitycrew39@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-11 01:40:43 -04:00
Gille
485aaf69b7 fix(local-runtime): find nvidia-smi in WSL driver path 2026-09-06 23:34:20 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
c8d401c79a refactor(local_runtime): compact module docstrings (rules kept, prose trimmed) 2026-09-02 21:47:26 -07:00
Teknium
723d169e68 refactor(local_runtime): join one-name-per-line imports and wrapped calls; guard-clause reflows in sweep_idle/_reap_orphaned_children/physics_check 2026-09-02 21:34:34 -07:00
Teknium
7c7d967b4c refactor(local_runtime): catalog.entry_for_model replaces 2 hit-unpack sites; tighten _host_os_arch 2026-09-02 21:10:03 -07:00
Teknium
090fd9c6e9 refactor(local_runtime): collapse swallow-and-default try/except ladders to contextlib.suppress; single struct reader in read_gguf_header 2026-09-02 21:05:10 -07:00
Teknium
9aa1f66484 refactor(local_runtime): GGUFHeader int accessors via one property factory; supervisor best-effort child ops via _quiet; compact __init__ re-exports 2026-09-02 20:58:40 -07:00
Teknium
a4a7ab6286 refactor(local_runtime): unify managed-router GET (endpoint.managed_root/managed_get_json) across load_progress+capabilities; shared kv() closure in initial_window; GGUF scalar-type table inlined; compact rationale prose across the package 2026-09-02 20:53:51 -07:00
Teknium
0244f7fa7e refactor(local_runtime): unify split-GGUF id/scan helpers (gguf.model_id_from_stem, bootstrap.staged_in, binaries.manifest_verified); split generate_presets into per-model phases; drop dead /health probe in _state_endpoint (pid was already the sole verdict) 2026-09-02 20:35:53 -07:00
Teknium
cb1b44aacd refactor(local_runtime): compact supervisor/catalog/hardware — shared _RESIDENT status tuple, collapsed integrated verdict, tightened rationale comments 2026-09-02 20:22:10 -07:00
Teknium
4d1880e0bf fix(integration): restore subprocess encoding/stdin guards dropped in simplification
Simplification workers collapsed subprocess call sites into shared kwargs
helpers and dropped the Windows/TUI safety kwargs on the way:

- encoding='utf-8', errors='replace' restored on text=True runs in
  copilot_acp_client, hermes_cli/setup (vercel install), managed_uv
  (codesign steps), local_runtime/hardware._stdout, a2a adapter.
- stdin=subprocess.DEVNULL restored on copilot probe, verify/runner
  _SUBPROCESS_KW, iron_proxy._run, google_meet playwright/system_profiler,
  simplex convert, whatsapp _RUN_TEXT, mem0 ollama serve Popen.
  google_meet sudo/brew install keeps inherited stdin (user-confirmed,
  may prompt) — marked noqa: subprocess-stdin.
- Windows-safe SIGKILL: getattr(signal, 'SIGKILL', SIGTERM) in
  verify/runner; photon _kill call re-marked windows-footgun: ok
  (unreachable on win32).
- scripts/check_subprocess_stdin.py now recognizes **kwargs splats
  (**_KW / **_kw(...)) ONLY when the same-file definition provably sets
  stdin= — covers tui_gateway _capture_run_kwargs/run_kw. Parity test added.
2026-09-02 16:36:10 -07:00
Teknium
3ffd44acd3 refactor(hclib): models/runtime — models, inventory, runtime_provider, provider catalog, local_runtime, banner 2026-09-02 14:44:48 -07:00
emozilla
43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00