- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
its runtime scope; another profile's secret is 401 at that URL; the bare path is
untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
hermes -p <name> gateway status, hermes gateway status and hermes status (under
Serves:) list the /p/<profile>/<path> URL per inbound-port platform the live
multiplexer serves, read from the <profile>:<platform> ingress_url in the default
home's gateway_state.json (hermes_cli/gateway_multiplex_served.py). The dashboard's
messaging payload carries the same ingress_url and the Channels page renders it.
The dashboard's 409 guard now covers only api_server/webhook (the mirrored pair):
enabling Twilio/LINE/Teams/... on a secondary is allowed because the gateway serves
it.
sms, line, teams, bluebubbles, msgraph_webhook, whatsapp_cloud, wecom_callback and
feishu (webhook mode) build their aiohttp app exactly as before and hand it to
bind_listener(): standalone and default-profile behaviour is unchanged (same host,
port, reuse_address, access_log), while a multiplex secondary publishes the app for
/p/<profile>/ forwarding instead of binding. BlueBubbles registers the /p/<profile>/
URL with its server; LINE builds media URLs off the shared listener's prefix when no
LINE_PUBLIC_URL is set; WeCom skips its own port-in-use probe in shared mode.
Under gateway.multiplex_profiles a secondary profile with Twilio / LINE / Teams /
BlueBubbles / Microsoft Graph / WhatsApp Cloud / WeCom-callback / Feishu-webhook
credentials was refused WHOLE at config load (SecondaryPortBindingConfigError): every
one of its platforms was skipped because these adapters bind their own port and the
default profile owns the one listener.
The refusal is gone. gateway/platforms/shared_ingress.py gives port-binding adapters
a shared-listener mode: the runner stamps `_shared_listener_profile` on a secondary's
port-binder, `bind_listener()` publishes the adapter's fully wired aiohttp app instead
of starting a TCPSite, and the default listener (api_server, or the webhook adapter
when there is no api_server) forwards `/p/<profile>/<tail>` to the served profile's
adapter whose router matches `/<tail>`, under that profile's runtime scope. The request
is therefore verified by the NAMED profile's adapter with its own secret and replies
leave through that adapter; the un-prefixed path keeps serving the default byte for
byte; an unknown profile or a profile without an adapter for the path is a 404, never
another profile's bot. api_server and webhook stay MIRRORS (the default's own adapter
answers /p/<profile>/ for them) and are the only port-binders a secondary must not
enable; `SHARED_LISTENER_MIRROR_PLATFORMS` is that set in gateway/config.py.
The adapter records its public callback URL (`ingress_url`) in gateway_state.json
under `<profile>:<platform>` and logs it once at connect, so the operator knows what
to paste into the vendor console.
advance_next_run alone no longer commits an occurrence — a restart before the
fire claim restores it (#107485). The at-most-once assertions model the real
tick sequence (advance → claim) so they keep guarding the mid-run crash case.
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
tick() advances a recurring job's next_run_at BEFORE dispatch so a crash
mid-run cannot re-fire it on every restart (at-most-once). That leaves a
window — advance persisted, fire claim not yet taken — in which the process
dying (interpreter already finalizing, executor refusing new futures, SIGKILL,
the Desktop idle-exit from #107485) loses the occurrence silently: the
restarted scan sees only the future next_run_at, writes no execution row and
no log line, and a daily job skips a day.
Contract (documented in cron-internals.md): every recurring occurrence is
accounted for — it runs once, or its skip is logged with a reason.
- The due scan stamps `pending_slot = {scheduled_at, at, by}` in the same save
that records `last_dispatch`; claim_job_for_fire and mark_job_run clear it,
and an explicit schedule / next_run_at / enabled / state rewrite (edit,
pause, resume, run-now) drops it.
- A later scan that finds the stamp with a provably gone owner (this process
and the job not in its running set, or another owner past the fire-claim
lease / dead pid) restores scheduled_at as next_run_at ONCE and logs a
WARNING (cron/occurrences.py::unclaimed_pending_slot). The restored instant
then meets the ordinary policy — completed_occurrence() blocks a second fire
of a slot that already ran, the grace window classifies it late, past-grace
collapses the backlog into one run, cron.catch_up_missed: false skips it
with a logged reason. Never a replay of N slots.
Same store fields on both topologies: a standalone `hermes -p X gateway run`
and a profile served by the default multiplexer evaluate the identical record.
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.
GET /api/gateway/migrate/plan returns the CLI plan JSON; POST
/api/gateway/migrate spawns `hermes gateway migrate --multiplex --yes`
detached (action log gateway-migrate.log). The Gateway card shows the
button only for a multi-profile install that is not yet multiplexed, and
disables it while listing the blockers.
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
Moves a per-profile-gateway install onto one multiplexed default gateway:
table-driven preflight (duplicate credential via the gateway's own
fingerprint; secondary port-binders without a /p/<profile>/ ingress),
--dry-run, apply (stop + uninstall each secondary's service, record it in
<default>/gateway_migration.json, flip gateway.multiplex_profiles through
the config API, restart/install the default on the same service manager,
verify served_profiles), and --standalone rollback from the manifest.
Idempotent; refuses cleanly when already multiplexed or blocked.
`hermes profile create` points at `hermes gateway restart` when a live
multiplexer is detected (the served set is snapshotted at startup).
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
Ramp Router efforts cache + warm/disk flags, xAI and OpenRouter image catalogs,
Hindsight append-capability verdict, memory-provider skill registry, OpenViking
atexit provider, Honcho loopback flow status, Langfuse client (os.environ-only
credentials) and disk-cleanup's protected cron paths held one profile's
credential- or home-derived value process-wide; YuanbaoAdapter._active_instance
was last-connected-wins across profiles.
Keyed by home key / credential fingerprint under an override, credentials read
through the secret scope, warm threads run under copy_context(); unscoped module
slots stay for the single-profile path and the existing monkeypatch tests.
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.
Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).
Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
boto3 and azure-identity freeze the credential chain into the client at
construction, and the process env under a multiplexed turn belongs to the launch
profile. A region-only (bedrock) / config-only (lru_cache) slot therefore signed a
served profile's calls with the launch profile's keys and served its account's
model list to everyone.
Under a HERMES_HOME override the clients are built from the profile's secret
scope (AWS_* / AZURE_* from its .env) and cached per (home, service, region) /
(home, config); the discovery cache key carries the home. The unscoped path keeps
the region slot and the maxsize=1 lru byte-for-byte.
Review findings on the salvage (all reproduced with a real SessionStore):
1. Primary persisted-agent path skipped the boundary. The agent's turn-start
flush already persists the user row stamped with the inbound platform id, so
`has_platform_message_id` saw THIS turn's own row, took the "duplicate" branch
and skipped the whole block — including the new assistant boundary. The
transcript stayed `[..., 'user']`, exactly the open tail #107070 is about.
Fresh sessions hid it a second way: `session_meta` is appended after the
agent-flushed user row, so a naive "newest row" tail read sees `session_meta`.
2. The exception fallback appended the boundary unconditionally; a redelivery of
an already-closed turn produced `['user', 'assistant', 'assistant']`.
3. The exception fallback wrote the user row + boundary before classifying a
400/500-on-long-session as overflow, growing a session that is already too
large (the #1630 no-grow rule the persist path honours).
Fix: `SessionDB.latest_conversation_role()` (newest active row excluding the
`session_meta`/`system` bookkeeping rows the model never sees) behind
`SessionStore.transcript_tail_role()`, which resolves the same route
`load_transcript` reads via the existing `_compression_tip_for_session_id`.
One `_hmwa_close_failed_turn()` appends the boundary iff that tail is an open
user row; both the persist path and the exception fallback call it, so the
user-row dedupe no longer gates the boundary and a redelivery never stacks.
The overflow verdict in `_hmwa_agent_error_reply` is an early return ahead of
every transcript write. The `failed_turn_notice` kwarg and its dead
`or _hmwa_failed_turn_notice(...)` fallback are gone; the notice is derived
where each consumer needs it.
Tests (each red with the production change reverted, green here): boundary
keyed on the durable tail with the user write deduped (agent-flushed row →
closed; redelivery → nothing); fresh-session agent-flushed failed first turn
closed despite `session_meta` (real store); exception-path redelivery adds no
second boundary (real store, every lineage location, contract asserted from
store state); exception-path overflow persists nothing. Live E2E:
`evals/gateway_failure_ownership/probe.py` (real AIAgent + fixture provider)
20/20; the two `failed provider input` turns that previously left an open user
tail now close with the "not processed" row.
The notice was derived twice per failed turn (reply + persisted row) from
the same input with different gates, so the shown text and the stored row
were not guaranteed identical. Classify once, pass failed_turn_notice into
_hmwa_persist_turn_transcript. Inline the one-line boundary-row staticmethod
at its two sites, drop the no-producer "notice already present" guard, and
assert the notice constants in tests instead of substrings.
The boundary row closes the user row it follows. When Telegram redelivers
the same failed message the dedupe branch skips the user row, so writing the
boundary unconditionally stacked two consecutive assistant rows and the
notice would be concatenated twice on replay. Append it only alongside the
user row; one test replays the retry.
The tool-evidence scan sliced messages[history_offset:] unguarded — after
mid-turn compression that slice is empty and the "not processed / resend"
notice would have been emitted over a turn whose tools DID run. Route it
through media_repair._current_turn_messages, which already falls back to the
last user row. Drop the notice=None default nobody used, compute the notice
inside _hmwa_persist_turn_transcript instead of threading it as a 12th
kwarg, and share one boundary-row builder between the persist path and the
exception fallback.
PR #107088 changed "Try again or use /reset to start a fresh session."
to "Use /reset to start a fresh session if needed." That rewording is
unrelated to the boundary-row fix; the PARTIAL notice appended right
after it already carries the "verify before resending" caveat. Keep
main's text so the diff stays scoped to the transcript boundary.
Follow-up to #107088 (fangliquanflq), refs #107070.
PR #107088 routed the "Session too large" reply through
_hmwa_add_failed_turn_notice with the PARTIAL notice ("some actions may
already have run; verify"). Context overflow is a deterministic request
rejection, not an indeterminate-effect failure — #107567 just tightened
that verdict — so the overflow branch returns main's exact text again.
One invariant test: a 400 on a >50-row history yields the /compact
guidance alone, without the partial notice.
Follow-up to #107088 (fangliquanflq), refs #107070.
Install with hermes skills install official/autonomous-ai-agents/dynamic-workflow.
Orchestration-campaign guidance is niche enough not to sit in every default prompt index.
Main moved under take-1: top-level delegate_task is background-forced (results
re-enter as messages; synthesizing on the same turn reads files that do not
exist yet), per-task `toolsets` is gone (children inherit the parent's set),
DELEGATE_BLOCKED_TOOLS is {delegate_task, clarify, memory, send_message,
cronjob_manage} (execute_code is NOT stripped), and the 457-char description
was truncated by the 60-char routing budget. All four bbopen review items fixed.
Adds the campaign shape learned running the 1,863-session / 19-hour
whole-codebase simplification fan-out (#102117): shared brief + exclusive
file ownership, commit-per-step as the only handoff, fleet ceiling before the
OAuth refresh stampede, per-round integration with a frozen base and a full
suite on the combined tree, `rev-list --count` per branch before declaring a
round done (168 late-slice commits were once left behind), one serialized
test runner, forward-port last, live QA as its own wave, refuting the parent's
own heuristics with the same attempt/refuter mechanic, HANDOFF.md on restart.
Frontmatter now meets the hardline standard (platforms, ≤60-char description,
modern section order); `/tmp` replaced by the terminal temp dir so the skill is
correct on Termux and Windows. Catalog + sidebar + generated page added.
Address all 5 review points against actual delegate_task behavior:
- child toolsets are subject to delegate restrictions (leaf strips
delegate_task/clarify/memory/send_message/execute_code), not 'full'
- durable work has lighter options than kanban (cron one-shot,
managed background terminal) for simpler cases
- unique per-run /tmp/wf_<name>_<uuid> dir + freshness/count check so
a stale interrupted run isn't read as success
- note that one delegate_task batch is capped by
delegation.max_concurrent_children; large fan-out needs bounded waves
- delegate_task exposes no per-task model/profile field (per-task keys
are goal/context/toolsets/role); model/profile-scoped runs go via
delegation config, cron, kanban, or separate process
Adapts Claude Code's research-preview dynamic workflows (plan-in-code
fan-out, hundreds of subagents per session) to Hermes invariants.
The ported mechanic is plan/loop/intermediate-state-out-of-context, not
more subagents. Documents the two real orchestration layers and the hard
capability boundary between them:
- Layer A (execute_code): deterministic fan-out, SANDBOX_ALLOWED_TOOLS
only, cannot call delegate_task
- Layer B (delegate_task batch): LLM-judgment fan-out
Plus the synchronous trap (delegate_task is turn-scoped, cancelled on new
message; durable/resumable = kanban swarm) and the genuinely-new piece:
the adversarial-convergence verification recipe (N independent attempts
with varied framings + M refuters, keep only located claims that survive
refutation, iterate to convergence).
Self-contained: inlines the load-bearing fan-out hygiene rather than
hard-depending on local-only skills; references the shipped kanban swarm
subsystem for the durable path.
- _drop_turn_slot takes the post-bump run_generation from _interrupt_running_turn
and forwards it to _release_running_agent_state: _interrupt_and_clear_session
awaits adapter.interrupt_session_activity between bump and release, so a
successor claiming the slot in that window must not have its sentinel/lease
wiped by the displaced /stop tail (sync eviction path forwards too, for the
same guard).
- _drop_turn_slot then sweeps lease_tokens from generations older than current:
a hung evicted turn's finalizer may never run, and each such generation pinned
its token (and _SessionLease) forever. Identity-checked + idempotent release,
so a live successor's token is untouched.
- _claim_one_turn_restore folds the two spellings of the one-shot "earliest
snapshot wins" rule (/moa direct write, /model --once setdefault) into one
helper; /model --once passes its pre-switch snapshot so the earliest-wins
contract is expressed once.
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.
Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
Trims the nine salvaged tests to four, one per contract: SUDO_USER's default
home keeps the bare name across the unit sync; a home the unit does not pin
keeps its suffix (the #105525 guard); a bare unit pinning a profiles/<name>
home beats the profile branch; the production sync itself preserves the name.
`_native_service_homes()` re-implemented the root+SUDO_USER -> `pw_dir/.hermes`
resolution that `hermes_cli.main._resolve_sudo_user_profile_env` already did.
Both now call `hermes_constants.sudo_invoker_default_home()`; main.py appends
`profiles/<name>` to it.
Also removes the local `from pathlib import Path as _Path` (Path is a module
import) and narrows the except to `KeyError`: `import pwd` cannot fail once
`os.geteuid` exists, and `getpwnam(str)` raises nothing else.
Review follow-ups to the unit-anchored service identity.
`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.
Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.
Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.
The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.
Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.
Refs #108674
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).
The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.
Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.
The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.
Refs #108674
- The `one_turn_restore["run_generation"]` stamp had no reader once settlement
moved to the invalidate chokepoint (the finalizer guards on its own generation
via _is_session_run_current); delete it and the test lines that set it.
- release-slot-then-evict-cached-agent (#44212 rationale) was duplicated in
/stop and eviction; one _drop_turn_slot owns it.
- Best-effort interrupt uses the repo's _log_suppressed seam like run_shutdown.
- turn_lease.rebind resolves the lease via token.lease like release does
(identity, not a session_id lookup); _held_turn_lease hands back the token
map so release/rebind stop re-peeking session state.
- Test trims: vacuous isinstance, registry-internals asserts.
`/model X --once` then `/model Y --once` before any turn replaced the
pending snapshot with one taken while X was live, so slot cleanup restored
X and made the first temporary model permanent (ehz0ah, review on #106966).
setdefault keeps the snapshot from the first command — the user's standing
override — as the restore target. One producer-driven regression test.
The turn finalizer now releases only its own generation
(`_release_running_agent_state(key, run_generation=...)`); two test doubles
were zero-kwarg lambdas and raised TypeError inside the finally.
_hm_evict_running_agent copied the head of _interrupt_and_clear_session
(peek → sentinel check → request_hard_interrupt → invalidate), minus the
turn-process reaper the stop path spawns — so tool subprocesses of an
evicted turn were never reaped. Extract the sync core
(_interrupt_running_turn) and call it from both; the raising-interrupt
guard now protects /stop as well.
#107013 made the finalizer's one-shot restore (/moa, /model --once)
generation-guarded so a displaced turn cannot clobber its successor — but
every displacing path (/stop, /new, idle or reaped eviction) bumps the
generation before that finalizer runs, so the guard skipped the restore and
the one-shot model stayed in force for every later message (main restored
unconditionally). Settle the snapshot inside
_invalidate_session_run_generation, the chokepoint all of those paths go
through, so the displaced finalizer then finds nothing to restore.
/moa now records its prior override in the same conversation.one_turn_restore
snapshot /model --once uses (via _snapshot_session_model_override) instead of
per-turn event attributes, which removes _restore_moa_one_shot,
_moa_run_generation and the "stamp only when None" plumbing; the finalizer
checks ownership with the existing _is_session_run_current. One test drives
the /stop-mid-turn settlement and the stale-finalizer no-op; the moa restore
test now exercises the shared path.
Every `TurnLeaseToken` is now constructed by `SessionTurnLeaseRegistry.acquire`
with its concrete `_SessionLease`, so `release()` no longer needs the
`getattr(token, "lease", None) or self._leases.get(...)` fallback that #107013
left in place. Make `lease` a required constructor argument and resolve the
lease from the token alone; the mapping lookup could only ever return the same
object (or a stale alias after rotation, which is exactly the case identity
release exists to avoid).