Commit Graph

9 Commits

Author SHA1 Message Date
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
84b8682203 refactor(gateway): collapse defensive layers, drop dead loggers, compact call sites across lifecycle/monitoring/media modules 2026-09-02 19:21:46 -07:00
Teknium
3238af61c1 refactor(gateway): tier tables for memory/disk pressure, shared ISO parser for drain markers, hand-compacted docstrings 2026-09-02 18:54:58 -07:00
Teknium
5fa7066054 refactor(gateway): compact restart/drain/cache-pressure/monitor modules; extract hook-dir loader and capture-path indexer 2026-09-02 18:33:07 -07:00
Teknium
98dccb4ca8 refactor(gateway): unify /proc + coercer helpers, merge timestamp regexes, collapse defensive layers in lifecycle/monitoring/media modules 2026-09-02 18:23:35 -07:00
Teknium
3c7069bdcb refactor(gateway): dead-code removal, helper unification, defensive-layer collapse and rationale-preserving comment compaction across 40 modules
authz_mixin, browser_control_broker, delivery, delivery_ledger, display_config, drain_control,
hosted_room_links/peer/policy_checkpoint, hosted_rooms, platform_registry, relay/__init__,
relay/ws_transport, run.py and slash_commands.py (comments), session_context, session_state,
streaming_tts_consumer, turn_lease and small modules.

- HostedRoomPolicyCheckpoint._apply_event -> per-kind handler table
- WebSocketRelayTransport._handle_frame -> frame-handler table
- GatewayAuthorizationMixin: unified adapter setting/flag/extra readers
- dead symbols removed (verified zero references): RoomLinkProbe/select_room_link,
  relay_bot_username, is_restart_loop_tripped, debug_rows, DeadTargetRegistry.all_dead,
  BrowserControlBroker.detach_owner/_prune_tickets, StreamingTTSConsumer.started/_enqueue_done/
  _iter_stream_chunks/_next_stream_chunk, RecoverableHandleCache.status_for, _auth_env,
  _copy_default_catalog, _parse_timestamp_prefix, _present_* helpers, _send_result_error_kind,
  _truthy_env, SessionFieldView/TurnLeaseTokenView dunder shims, and their orphaned tests.
- lost WHY/invariant text from the earlier compaction restored compactly (541 hunks audited)
2026-09-02 13:30:50 -07:00
Shannon Sands
ba5dc00bb2 fix(memory-status): review follow-ups — incident-keyed dismissal, honest copy, single mobile offset (NS-656)
Addresses the human review findings on the memory-pressure feature:

* [P2] Dismissal hid later incidents of the same kind. The gateway now
  publishes `boot_id` (the lifecycle sentinel's started_at — changes on
  every gateway life) in the /api/status memory block, and the dashboard
  keys OOM-restart dismissal on it: acknowledging one restart no longer
  mutes the NEXT one (the OOM-loop case this banner exists for). Live
  pressure dismissals now also reset once pressure is demonstrably back
  to "ok" — "unknown" (stale heartbeat) is absence of evidence and
  clears nothing. Dismissal storage moved to a JSON list; old bare-string
  entries fail JSON.parse and degrade to a clean reset.

* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
  heartbeat), not proof the OOM killer acted — banner copy now says
  "restarted unexpectedly, most likely because it ran out of memory"
  instead of stating OOM as fact.

* [P3] Mobile header clearance was applied per-banner (mt-14 on both
  MemoryPressureBanner and ProfileScopeBanner) AND on the content
  (pt-14), double/triple-stacking 56px gaps when banners were visible.
  Replaced with a single h-14 spacer above the banner stack.
2026-08-13 20:30:12 -07:00
Shannon Sands
e11d1ddc7f feat(status): surface memory pressure and suspected-OOM restarts to users (NS-656)
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.

This is the read-side fix:

* New gateway/memory_status.py distills the existing 30s loop heartbeat
  (gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
  sentinel into a compact `memory` block: pressure ok/elevated/critical/
  unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
  Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
  future-dated heartbeats degrade pressure to "unknown" so a dead
  gateway's final gasp can't render a live "critical" banner forever.
  Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
  level would make a later unclean death "suspected OOM", warn at that
  level while the process is still alive.

* lifecycle_ledger.record_startup now carries prior_unclean_exit /
  prior_suspected_oom onto the reclaimed sentinel — previously the
  verdict survived only in append-only diag prose. Flags age out on the
  next sentinel rewrite (scoped to the life after the crash).

* /api/status serves the block (profile-aware, executor-offloaded,
  fail-safe to pressure=unknown). Deliberately NOT folded into
  components/overall: memory pressure is advisory, and flipping overall
  to "degraded" on it would page NAS's availability sweep for a
  condition the eviction valve is already handling. Public-safety:
  coarse numbers/enums/booleans only — same disclosure class as the
  existing nous_session_valid field, added for the same NAS-sweep
  audience.

* Dashboard: new MemoryPressureBanner (app-shell, next to
  ProfileScopeBanner) with worst-first trigger precedence
  (critical > suspected-OOM restart > elevated), per-trigger
  session-scoped dismissal, and escalation re-opening past a dismissal.
  i18n keys optional with English fallbacks, matching the
  managingProfileBanner convention.

Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.

NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.

Refs NS-656; context: NS-608, NS-657, OOF-77.
2026-08-13 20:30:12 -07:00