Commit Graph

33817 Commits

Author SHA1 Message Date
Teknium
9ca7db8232 test(multiplex): shared-listener ingress invariants
- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
  its runtime scope; another profile's secret is 401 at that URL; the bare path is
  untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
  in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
  of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
2026-09-12 01:53:15 -07:00
Teknium
65fd3a2b9c feat(status): show each served profile's shared-listener callback URLs
hermes -p <name> gateway status, hermes gateway status and hermes status (under
Serves:) list the /p/<profile>/<path> URL per inbound-port platform the live
multiplexer serves, read from the <profile>:<platform> ingress_url in the default
home's gateway_state.json (hermes_cli/gateway_multiplex_served.py). The dashboard's
messaging payload carries the same ingress_url and the Channels page renders it.
The dashboard's 409 guard now covers only api_server/webhook (the mirrored pair):
enabling Twilio/LINE/Teams/... on a secondary is allowed because the gateway serves
it.
2026-09-12 01:53:15 -07:00
Teknium
6fc76ce751 feat(platforms): bind through shared_ingress.bind_listener in every inbound-port adapter
sms, line, teams, bluebubbles, msgraph_webhook, whatsapp_cloud, wecom_callback and
feishu (webhook mode) build their aiohttp app exactly as before and hand it to
bind_listener(): standalone and default-profile behaviour is unchanged (same host,
port, reuse_address, access_log), while a multiplex secondary publishes the app for
/p/<profile>/ forwarding instead of binding. BlueBubbles registers the /p/<profile>/
URL with its server; LINE builds media URLs off the shared listener's prefix when no
LINE_PUBLIC_URL is set; WeCom skips its own port-in-use probe in shared mode.
2026-09-12 01:53:15 -07:00
Teknium
c0d8d6b5d5 feat(multiplex): serve secondary profiles' inbound-port platforms on the default's shared listener
Under gateway.multiplex_profiles a secondary profile with Twilio / LINE / Teams /
BlueBubbles / Microsoft Graph / WhatsApp Cloud / WeCom-callback / Feishu-webhook
credentials was refused WHOLE at config load (SecondaryPortBindingConfigError): every
one of its platforms was skipped because these adapters bind their own port and the
default profile owns the one listener.

The refusal is gone. gateway/platforms/shared_ingress.py gives port-binding adapters
a shared-listener mode: the runner stamps `_shared_listener_profile` on a secondary's
port-binder, `bind_listener()` publishes the adapter's fully wired aiohttp app instead
of starting a TCPSite, and the default listener (api_server, or the webhook adapter
when there is no api_server) forwards `/p/<profile>/<tail>` to the served profile's
adapter whose router matches `/<tail>`, under that profile's runtime scope. The request
is therefore verified by the NAMED profile's adapter with its own secret and replies
leave through that adapter; the un-prefixed path keeps serving the default byte for
byte; an unknown profile or a profile without an adapter for the path is a 404, never
another profile's bot. api_server and webhook stay MIRRORS (the default's own adapter
answers /p/<profile>/ for them) and are the only port-binders a secondary must not
enable; `SHARED_LISTENER_MIRROR_PLATFORMS` is that set in gateway/config.py.

The adapter records its public callback URL (`ingress_url`) in gateway_state.json
under `<profile>:<platform>` and logs it once at connect, so the operator knows what
to paste into the vendor console.
2026-09-12 01:53:15 -07:00
Teknium
36b1a6e62c test(cron): crash-safety tests claim the fire before asserting not-due
advance_next_run alone no longer commits an occurrence — a restart before the
fire claim restores it (#107485). The at-most-once assertions model the real
tick sequence (advance → claim) so they keep guarding the mid-run crash case.
2026-09-12 01:51:59 -07:00
Teknium
6cd4fbd640 docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
2026-09-12 01:51:59 -07:00
Teknium
eaa5ac94ea fix(cron): fire one catch-up for a slot missed during a restart gap (#107485)
tick() advances a recurring job's next_run_at BEFORE dispatch so a crash
mid-run cannot re-fire it on every restart (at-most-once). That leaves a
window — advance persisted, fire claim not yet taken — in which the process
dying (interpreter already finalizing, executor refusing new futures, SIGKILL,
the Desktop idle-exit from #107485) loses the occurrence silently: the
restarted scan sees only the future next_run_at, writes no execution row and
no log line, and a daily job skips a day.

Contract (documented in cron-internals.md): every recurring occurrence is
accounted for — it runs once, or its skip is logged with a reason.

- The due scan stamps `pending_slot = {scheduled_at, at, by}` in the same save
  that records `last_dispatch`; claim_job_for_fire and mark_job_run clear it,
  and an explicit schedule / next_run_at / enabled / state rewrite (edit,
  pause, resume, run-now) drops it.
- A later scan that finds the stamp with a provably gone owner (this process
  and the job not in its running set, or another owner past the fire-claim
  lease / dead pid) restores scheduled_at as next_run_at ONCE and logs a
  WARNING (cron/occurrences.py::unclaimed_pending_slot). The restored instant
  then meets the ordinary policy — completed_occurrence() blocks a second fire
  of a slot that already ran, the grace window classifies it late, past-grace
  collapses the backlog into one run, cron.catch_up_missed: false skips it
  with a logged reason. Never a replay of N slots.

Same store fields on both topologies: a standalone `hermes -p X gateway run`
and a profile served by the default multiplexer evaluate the identical record.
2026-09-12 01:51:59 -07:00
Teknium
2bc471cb35 fix(cron): skip catch-up only with a future anchor 2026-09-12 01:51:59 -07:00
Teknium
df51797e2e feat(cron): let planned downtime skip missed recurring runs 2026-09-12 01:51:59 -07:00
Teknium
be82d52cb7 docs: migrating from per-profile gateways to the multiplexer
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.
2026-09-12 01:49:28 -07:00
Teknium
df95f378a6 feat(dashboard): "Migrate to a single multiplexed gateway" on the System page
GET /api/gateway/migrate/plan returns the CLI plan JSON; POST
/api/gateway/migrate spawns `hermes gateway migrate --multiplex --yes`
detached (action log gateway-migrate.log). The Gateway card shows the
button only for a multi-profile install that is not yet multiplexed, and
disables it while listing the blockers.
2026-09-12 01:49:28 -07:00
Teknium
07a4ae016a feat(update): auto-migrate to one multiplexed gateway when unblocked
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
2026-09-12 01:49:28 -07:00
Teknium
e2fc493427 feat(gateway): hermes gateway migrate --multiplex / --standalone
Moves a per-profile-gateway install onto one multiplexed default gateway:
table-driven preflight (duplicate credential via the gateway's own
fingerprint; secondary port-binders without a /p/<profile>/ ingress),
--dry-run, apply (stop + uninstall each secondary's service, record it in
<default>/gateway_migration.json, flip gateway.multiplex_profiles through
the config API, restart/install the default on the same service manager,
verify served_profiles), and --standalone rollback from the manifest.
Idempotent; refuses cleanly when already multiplexed or blocked.

`hermes profile create` points at `hermes gateway restart` when a live
multiplexer is detected (the served set is snapshotted at startup).
2026-09-12 01:49:28 -07:00
Teknium
bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00
Teknium
5aa17c0590 docs(multiplex): list the newly per-profile caches in the isolation table 2026-09-12 01:35:05 -07:00
Teknium
208bd0b65a fix(multiplex): per-profile plugin caches and the Yuanbao active adapter
Ramp Router efforts cache + warm/disk flags, xAI and OpenRouter image catalogs,
Hindsight append-capability verdict, memory-provider skill registry, OpenViking
atexit provider, Honcho loopback flow status, Langfuse client (os.environ-only
credentials) and disk-cleanup's protected cron paths held one profile's
credential- or home-derived value process-wide; YuanbaoAdapter._active_instance
was last-connected-wins across profiles.

Keyed by home key / credential fingerprint under an override, credentials read
through the secret scope, warm threads run under copy_context(); unscoped module
slots stay for the single-profile path and the existing monkeypatch tests.
2026-09-12 01:35:05 -07:00
Teknium
5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
Teknium
4b8c01f691 fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).

Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
2026-09-12 01:35:05 -07:00
Teknium
2cc6d88c1d fix(multiplex): build Bedrock and Entra credential clients from the routed profile
boto3 and azure-identity freeze the credential chain into the client at
construction, and the process env under a multiplexed turn belongs to the launch
profile. A region-only (bedrock) / config-only (lru_cache) slot therefore signed a
served profile's calls with the launch profile's keys and served its account's
model list to everyone.

Under a HERMES_HOME override the clients are built from the profile's secret
scope (AWS_* / AZURE_* from its .env) and cached per (home, service, region) /
(home, config); the discovery cache key carries the home. The unscoped path keeps
the region slot and the maxsize=1 lru byte-for-byte.
2026-09-12 01:35:05 -07:00
kshitij
044a77b3b6 fix(gateway): failed-turn boundary keyed on the durable conversation tail; exception path classifies overflow before writing
Review findings on the salvage (all reproduced with a real SessionStore):

1. Primary persisted-agent path skipped the boundary. The agent's turn-start
   flush already persists the user row stamped with the inbound platform id, so
   `has_platform_message_id` saw THIS turn's own row, took the "duplicate" branch
   and skipped the whole block — including the new assistant boundary. The
   transcript stayed `[..., 'user']`, exactly the open tail #107070 is about.
   Fresh sessions hid it a second way: `session_meta` is appended after the
   agent-flushed user row, so a naive "newest row" tail read sees `session_meta`.
2. The exception fallback appended the boundary unconditionally; a redelivery of
   an already-closed turn produced `['user', 'assistant', 'assistant']`.
3. The exception fallback wrote the user row + boundary before classifying a
   400/500-on-long-session as overflow, growing a session that is already too
   large (the #1630 no-grow rule the persist path honours).

Fix: `SessionDB.latest_conversation_role()` (newest active row excluding the
`session_meta`/`system` bookkeeping rows the model never sees) behind
`SessionStore.transcript_tail_role()`, which resolves the same route
`load_transcript` reads via the existing `_compression_tip_for_session_id`.
One `_hmwa_close_failed_turn()` appends the boundary iff that tail is an open
user row; both the persist path and the exception fallback call it, so the
user-row dedupe no longer gates the boundary and a redelivery never stacks.
The overflow verdict in `_hmwa_agent_error_reply` is an early return ahead of
every transcript write. The `failed_turn_notice` kwarg and its dead
`or _hmwa_failed_turn_notice(...)` fallback are gone; the notice is derived
where each consumer needs it.

Tests (each red with the production change reverted, green here): boundary
keyed on the durable tail with the user write deduped (agent-flushed row →
closed; redelivery → nothing); fresh-session agent-flushed failed first turn
closed despite `session_meta` (real store); exception-path redelivery adds no
second boundary (real store, every lineage location, contract asserted from
store state); exception-path overflow persists nothing. Live E2E:
`evals/gateway_failure_ownership/probe.py` (real AIAgent + fixture provider)
20/20; the two `failed provider input` turns that previously left an open user
tail now close with the "not processed" row.
2026-09-12 13:00:29 +05:30
kshitijk4poor
b9f543951e refactor(gateway): compute the failed-turn notice once; the reply and the boundary row share it
The notice was derived twice per failed turn (reply + persisted row) from
the same input with different gates, so the shown text and the stored row
were not guaranteed identical. Classify once, pass failed_turn_notice into
_hmwa_persist_turn_transcript. Inline the one-line boundary-row staticmethod
at its two sites, drop the no-producer "notice already present" guard, and
assert the notice constants in tests instead of substrings.
2026-09-12 13:00:29 +05:30
kshitijk4poor
79680b04bd fix(gateway): a deduped platform retry of a failed turn adds no second boundary row
The boundary row closes the user row it follows. When Telegram redelivers
the same failed message the dedupe branch skips the user row, so writing the
boundary unconditionally stacked two consecutive assistant rows and the
notice would be concatenated twice on replay. Append it only alongside the
user row; one test replays the retry.
2026-09-12 13:00:29 +05:30
kshitijk4poor
07b0c7e1bc refactor(gateway): failed-turn boundary row built once; notice derived where it is persisted
The tool-evidence scan sliced messages[history_offset:] unguarded — after
mid-turn compression that slice is empty and the "not processed / resend"
notice would have been emitted over a turn whose tools DID run. Route it
through media_repair._current_turn_messages, which already falls back to the
last user row. Drop the notice=None default nobody used, compute the notice
inside _hmwa_persist_turn_transcript instead of threading it as a 12th
kwarg, and share one boundary-row builder between the persist path and the
exception fallback.
2026-09-12 13:00:29 +05:30
kshitijk4poor
c4c7fe13c3 fix(gateway): restore main's generic error-reply wording
PR #107088 changed "Try again or use /reset to start a fresh session."
to "Use /reset to start a fresh session if needed." That rewording is
unrelated to the boundary-row fix; the PARTIAL notice appended right
after it already carries the "verify before resending" caveat. Keep
main's text so the diff stays scoped to the transcript boundary.

Follow-up to #107088 (fangliquanflq), refs #107070.
2026-09-12 13:00:29 +05:30
kshitijk4poor
241752ac97 fix(gateway): keep the context-overflow error reply free of the partial-effect notice
PR #107088 routed the "Session too large" reply through
_hmwa_add_failed_turn_notice with the PARTIAL notice ("some actions may
already have run; verify"). Context overflow is a deterministic request
rejection, not an indeterminate-effect failure — #107567 just tightened
that verdict — so the overflow branch returns main's exact text again.

One invariant test: a 400 on a >50-row history yields the /compact
guidance alone, without the partial notice.

Follow-up to #107088 (fangliquanflq), refs #107070.
2026-09-12 13:00:29 +05:30
fangliquanflq
371f0a057b fix(gateway): make failed-turn retry guidance side-effect safe 2026-09-12 13:00:29 +05:30
fangliquanflq
25c69ee63c fix(gateway): close failed turns before future replay 2026-09-12 13:00:29 +05:30
hermes-seaeye[bot]
28326a6ccd fmt(js): npm run fix on merge (#108902)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-12 07:09:15 +00:00
Teknium
bf867d3c74 chore(skills): ship dynamic-workflow as an optional skill
Install with hermes skills install official/autonomous-ai-agents/dynamic-workflow.
Orchestration-campaign guidance is niche enough not to sit in every default prompt index.
2026-09-12 00:02:10 -07:00
Teknium
fb4ed7b284 docs(skills): dynamic-workflow v2 — background-first delegation, campaign lessons from #102117
Main moved under take-1: top-level delegate_task is background-forced (results
re-enter as messages; synthesizing on the same turn reads files that do not
exist yet), per-task `toolsets` is gone (children inherit the parent's set),
DELEGATE_BLOCKED_TOOLS is {delegate_task, clarify, memory, send_message,
cronjob_manage} (execute_code is NOT stripped), and the 457-char description
was truncated by the 60-char routing budget. All four bbopen review items fixed.

Adds the campaign shape learned running the 1,863-session / 19-hour
whole-codebase simplification fan-out (#102117): shared brief + exclusive
file ownership, commit-per-step as the only handoff, fleet ceiling before the
OAuth refresh stampede, per-round integration with a frozen base and a full
suite on the combined tree, `rev-list --count` per branch before declaring a
round done (168 late-slice commits were once left behind), one serialized
test runner, forward-port last, live QA as its own wave, refuting the parent's
own heuristics with the same attempt/refuter mechanic, HANDOFF.md on restart.

Frontmatter now meets the hardline standard (platforms, ≤60-char description,
modern section order); `/tmp` replaced by the terminal temp dir so the skill is
correct on Termux and Windows. Catalog + sidebar + generated page added.
2026-09-12 00:02:10 -07:00
teknium1
8fb0350b61 docs(skills): tighten dynamic-workflow per donovan-yohan review
Address all 5 review points against actual delegate_task behavior:
- child toolsets are subject to delegate restrictions (leaf strips
  delegate_task/clarify/memory/send_message/execute_code), not 'full'
- durable work has lighter options than kanban (cron one-shot,
  managed background terminal) for simpler cases
- unique per-run /tmp/wf_<name>_<uuid> dir + freshness/count check so
  a stale interrupted run isn't read as success
- note that one delegate_task batch is capped by
  delegation.max_concurrent_children; large fan-out needs bounded waves
- delegate_task exposes no per-task model/profile field (per-task keys
  are goal/context/toolsets/role); model/profile-scoped runs go via
  delegation config, cron, kanban, or separate process
2026-09-12 00:02:10 -07:00
teknium1
351041463d feat(skills): add dynamic-workflow orchestration skill
Adapts Claude Code's research-preview dynamic workflows (plan-in-code
fan-out, hundreds of subagents per session) to Hermes invariants.

The ported mechanic is plan/loop/intermediate-state-out-of-context, not
more subagents. Documents the two real orchestration layers and the hard
capability boundary between them:
- Layer A (execute_code): deterministic fan-out, SANDBOX_ALLOWED_TOOLS
  only, cannot call delegate_task
- Layer B (delegate_task batch): LLM-judgment fan-out

Plus the synchronous trap (delegate_task is turn-scoped, cancelled on new
message; durable/resumable = kanban swarm) and the genuinely-new piece:
the adversarial-convergence verification recipe (N independent attempts
with varied framings + M refuters, keep only located claims that survive
refutation, iterate to convergence).

Self-contained: inlines the load-bearing fan-out hygiene rather than
hard-depending on local-only skills; references the shipped kanban swarm
subsystem for the durable path.
2026-09-12 00:02:10 -07:00
kshitijk4poor
436ec48985 fix(gateway): generation-guarded slot release on /stop; displaced lease tokens swept
- _drop_turn_slot takes the post-bump run_generation from _interrupt_running_turn
  and forwards it to _release_running_agent_state: _interrupt_and_clear_session
  awaits adapter.interrupt_session_activity between bump and release, so a
  successor claiming the slot in that window must not have its sentinel/lease
  wiped by the displaced /stop tail (sync eviction path forwards too, for the
  same guard).
- _drop_turn_slot then sweeps lease_tokens from generations older than current:
  a hung evicted turn's finalizer may never run, and each such generation pinned
  its token (and _SessionLease) forever. Identity-checked + idempotent release,
  so a live successor's token is untouched.
- _claim_one_turn_restore folds the two spellings of the one-shot "earliest
  snapshot wins" rule (/moa direct write, /model --once setdefault) into one
  helper; /model --once passes its pre-switch snapshot so the earliest-wins
  contract is expressed once.
2026-09-12 12:24:24 +05:30
kshitijk4poor
443c2785fa docs(gateway): state the sudo mid-command rationale once, at the decision point
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
2026-09-12 12:15:07 +05:30
kshitijk4poor
1ecdfd18db fix(gateway): consult the installed unit only when root
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.

Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
2026-09-12 12:15:07 +05:30
kshitijk4poor
b18c5752e3 test(gateway): keep one invariant test per bare-name owner
Trims the nine salvaged tests to four, one per contract: SUDO_USER's default
home keeps the bare name across the unit sync; a home the unit does not pin
keeps its suffix (the #105525 guard); a bare unit pinning a profiles/<name>
home beats the profile branch; the production sync itself preserves the name.
2026-09-12 12:15:07 +05:30
kshitijk4poor
a32b8cd4a3 refactor(gateway): share the sudo-invoker home lookup, drop unreachable guards
`_native_service_homes()` re-implemented the root+SUDO_USER -> `pw_dir/.hermes`
resolution that `hermes_cli.main._resolve_sudo_user_profile_env` already did.
Both now call `hermes_constants.sudo_invoker_default_home()`; main.py appends
`profiles/<name>` to it.

Also removes the local `from pathlib import Path as _Path` (Path is a module
import) and narrows the except to `KeyError`: `import pwd` cannot fail once
`os.geteuid` exists, and `getpwnam(str)` raises nothing else.
2026-09-12 12:15:07 +05:30
JoaoMarcos44
1ff56d9ed9 fix(gateway): keep the unit anchor Linux-only and pin the order it depends on
Review follow-ups to the unit-anchored service identity.

`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.

Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.

Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.

The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.

Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.

Refs #108674
2026-09-12 12:15:07 +05:30
JoaoMarcos44
a4a10eb520 fix(gateway): anchor system service identity on the installed unit
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).

The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.

Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.

The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.

Refs #108674
2026-09-12 12:15:07 +05:30
Hukla
fdb9b50a15 fix(gateway): preserve sudo system service name 2026-09-12 12:15:07 +05:30
kshitijk4poor
24699e07a6 chore: map huklaa's contributor email 2026-09-12 12:15:07 +05:30
brooklyn!
99d7f17a1c docs(desktop): explain automatic tutorial cutoff in settings 2026-09-12 01:35:33 -05:00
brooklyn!
995e79afe4 feat(desktop): retire tutorials after the first month 2026-09-12 01:35:33 -05:00
kshitijk4poor
526ce1b64c test(gateway): eviction-stub release-thread fakes accept session_key kwarg 2026-09-12 12:03:39 +05:30
kshitijk4poor
789eec791b refactor(gateway): drop the write-only one-shot generation stamp; /stop and eviction share _drop_turn_slot
- The `one_turn_restore["run_generation"]` stamp had no reader once settlement
  moved to the invalidate chokepoint (the finalizer guards on its own generation
  via _is_session_run_current); delete it and the test lines that set it.
- release-slot-then-evict-cached-agent (#44212 rationale) was duplicated in
  /stop and eviction; one _drop_turn_slot owns it.
- Best-effort interrupt uses the repo's _log_suppressed seam like run_shutdown.
- turn_lease.rebind resolves the lease via token.lease like release does
  (identity, not a session_id lookup); _held_turn_lease hands back the token
  map so release/rebind stop re-peeking session state.
- Test trims: vacuous isinstance, registry-internals asserts.
2026-09-12 12:03:39 +05:30
kshitijk4poor
90970d85f2 fix(gateway): a repeated /model --once keeps the earliest restore target
`/model X --once` then `/model Y --once` before any turn replaced the
pending snapshot with one taken while X was live, so slot cleanup restored
X and made the first temporary model permanent (ehz0ah, review on #106966).
setdefault keeps the snapshot from the first command — the user's standing
override — as the restore target. One producer-driven regression test.
2026-09-12 12:03:39 +05:30
kshitijk4poor
c71a411fe5 test(gateway): release-slot fakes accept the finalizer's run_generation kwarg
The turn finalizer now releases only its own generation
(`_release_running_agent_state(key, run_generation=...)`); two test doubles
were zero-kwarg lambdas and raised TypeError inside the finally.
2026-09-12 12:03:39 +05:30
kshitijk4poor
c133095cc5 refactor(gateway): eviction and /stop share one interrupt-and-invalidate core
_hm_evict_running_agent copied the head of _interrupt_and_clear_session
(peek → sentinel check → request_hard_interrupt → invalidate), minus the
turn-process reaper the stop path spawns — so tool subprocesses of an
evicted turn were never reaped. Extract the sync core
(_interrupt_running_turn) and call it from both; the raising-interrupt
guard now protects /stop as well.
2026-09-12 12:03:39 +05:30
kshitijk4poor
549a6fb359 fix(gateway): stop, reset and eviction settle a one-shot override before displacing the turn
#107013 made the finalizer's one-shot restore (/moa, /model --once)
generation-guarded so a displaced turn cannot clobber its successor — but
every displacing path (/stop, /new, idle or reaped eviction) bumps the
generation before that finalizer runs, so the guard skipped the restore and
the one-shot model stayed in force for every later message (main restored
unconditionally). Settle the snapshot inside
_invalidate_session_run_generation, the chokepoint all of those paths go
through, so the displaced finalizer then finds nothing to restore.

/moa now records its prior override in the same conversation.one_turn_restore
snapshot /model --once uses (via _snapshot_session_model_override) instead of
per-turn event attributes, which removes _restore_moa_one_shot,
_moa_run_generation and the "stamp only when None" plumbing; the finalizer
checks ownership with the existing _is_session_run_current. One test drives
the /stop-mid-turn settlement and the stale-finalizer no-op; the moa restore
test now exercises the shared path.
2026-09-12 12:03:39 +05:30
kshitijk4poor
18da3f673a refactor(gateway): release turn leases by token identity only
Every `TurnLeaseToken` is now constructed by `SessionTurnLeaseRegistry.acquire`
with its concrete `_SessionLease`, so `release()` no longer needs the
`getattr(token, "lease", None) or self._leases.get(...)` fallback that #107013
left in place. Make `lease` a required constructor argument and resolve the
lease from the token alone; the mapping lookup could only ever return the same
object (or a stale alias after rotation, which is exactly the case identity
release exists to avoid).
2026-09-12 12:03:39 +05:30