Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The re-offer window (2 x 600 s = 1200 s) was shorter than a legitimate in-flight
bot_relay.deliver hold (lock wait + two attempts = 1320 s) and than the
Desktop's own deliver deadline (1500 s), so a slow-but-live delivery with no
reply on disk yet was handed out AGAIN by the next drain: a double turn on the
target (or target_busy), whose fast error then won first-settled-wins over the
real answer. And every re-offer bumped the mtime and opened a new window, so an
envelope nobody ever answered was re-offered every 1200 s forever - the 6 h sweep
is mtime-based too - long after the sender's waiter had given up.
- REOFFER_AFTER_SECONDS = DESKTOP_DELIVER_TIMEOUT_SECONDS + 60: past the point
where the Desktop has provably posted its own delivery_timeout reply (the
gateway-side hold ends before it by construction), silence means a dead
Desktop. The false "never in flight that long" comment is gone.
- REPLY_WAIT_SECONDS = REOFFER_AFTER_SECONDS + DESKTOP_DELIVER_TIMEOUT_SECONDS
+ 60, so the waiter is still listening when the one re-offered delivery hits
its own deadline.
- One re-offer per envelope (`reoffered_at` stamped on the claimed file); once
created_at + REPLY_WAIT_SECONDS passes unanswered the drain writes a
delivery_timeout reply instead, so the sender learns and no turn loop runs
against a target nobody is waiting for. Age is created_at, not bumped mtime.
- bot_mode.envelope_ttl_seconds applies to the re-offer leg exactly as to the
outbox: the message is back in the queue from claim + REOFFER_AFTER_SECONDS,
and a drain that comes a whole TTL later refuses it with queued_expired.
- write_reply's first-settled-wins is now documented as safe BECAUSE two
deliveries of one envelope can no longer overlap.
The existing constants test pins REOFFER > Desktop deadline > live hold and
REPLY_WAIT > REOFFER + Desktop deadline; the drain-side test drives the real
outbox.drain handler through re-offer, no second re-offer, and the timeout
reply. Docs bullet reworded to the new window.
After outbox.drain moved an envelope to claimed/, a Desktop that disconnected
before bot_relay.deliver left it there with no reply: the sender's waiter
learned nothing until its deadline, every later drain saw an empty outbox, and
the 6h sweep deleted the message. claim_pending_envelopes now re-offers claimed
envelopes unanswered for REOFFER_AFTER_SECONDS (two turn attempts — a live
delivery never runs that long without the Desktop posting its own timeout
reply), bumping the mtime so each re-offer opens a new window; the claim itself
now stamps the mtime so the window counts from the claim, not the enqueue.
write_reply is idempotent by envelope id: the first settled reply is kept, so a
re-offered delivery's second outcome never displaces the answer the waiter read.
Slim redo of #111207 (@JoaoMarcos44): the Desktop already drains on every
reconnect (b469be8cc3, 1eb771e2ff), so no Desktop change and no per-envelope
receipt store are needed for the loss the PR reproduced. Closes the residual of
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
`bot_relay.deliver` refused any call that named the sender when the calling transport carried
an authenticated identity, on the premise — stated in its comment and in the shipped test
docstring — that "the Desktop and server-internal callers carry no login identity". That is
false for one transport shape: a Desktop socket onto a gateway whose `/api/status` reports
`auth_required`. There the Desktop upgrades as `?ticket=`, and `consume_ticket` stamps the
signed-in `{user_id, provider}` onto the WS transport (hermes_cli/web_server_chat.py). Only
`?internal=` callers are identity-exempt and the Desktop cannot present one, while the relay
always sends `from_connection` — so every cross-connection `message_agent` DM whose TARGET is
such a gateway was refused with 4095 before any turn ran. The envelope has already been
atomically claimed and the refusal is written back as its reply, so the DM is lost and the
sender's waiter exits 1.
Scope: one direction (into a gated gateway), oauth-mode URL/cloud connections only. SSH-mode
connections authenticate with a dashboard token, stamp no identity, and are unaffected; so are
loopback/token installs, which is why this survived review. The legacy token is rejected in
gated mode, so token mode is not an escape for a gated gateway. Group chats do not use this
handler. The `<peer>/<agent>` route over the api_server peer link still reaches a gated
machine, but it is fire-and-forget with no reply to the sender — a degraded alternative.
The sender fields a logged-in client names are still not trusted — but the author is neither
refused nor dropped. Dropping it (this PR's first head) made the turn the HUMAN's to the
recipient's memory: Honcho's `sync_turn` routes an unattributed turn into the human session and
`_bot_turn_write_refusal` / `on_memory_write` guard only on `is_bot`, so conclusion, profile
and mirror writes went through — violating "a bot author's turn is written into that bot's own
a2a session, never the human's". The author is now derived from the caller's minted identity
(`_principal_digest`, stable and unspoofable): `bot:<principal>/relay`, `is_bot`, whether or
not the client named a sender — which also closes the pre-existing unattributed delivery a
logged-in client could make by naming none. The human-facing "Message from 🤖 …" signature
stays in the text the sender composed. `prompt.submit` continues to accept only an in-process
`DeliveryAuthor`.
Tests pin both sides: every sender shape from a logged-in client delivers with the principal
author and never the claimed one (one principal, one author; different principal, different
author) on the subprocess and live paths; a ws-ticket identity is a login identity so the
Desktop is one; and, reaching the downstream memory routing and write guards, the principal
author lands in its own Honcho a2a session, never the human's, with conclusion / profile /
mirror writes refused — alongside the failure mode it prevents. The test that encoded the old
premise is replaced rather than deleted.
(cherry picked from commit 0b669833f81867c3a3d7fde990d7e226226feef9)
Every gateway's `default` profile is `@hermes`, so a remote default titled
"CoS Bot" was unreachable by any bare form: `resolve_remote_target` matched
handle/profile only, the prompt roster offered it as `@hermes` (which local
resolution routes to THIS gateway's default), and the DM it sent arrived
stamped `(@hermes)` so a reply hit the recipient's own default.
- `tools/bot_relay.py`: `_target_aliases` adds the Bot Mode title's mention
slugs; `resolve_remote_target` accepts them (exact handle/profile keeps
precedence over a colliding title); `remote_target_forms` picks the shortest
form no LOCAL profile and no other remote row answers to — bare handle,
title slug, then `handle@connection` — independent of roster order;
`qualify_sender_stamp` rewrites a relayed `Message from 🤖 … (@handle):`
to that form.
- `tools/bot_mode_probe.py::local_taken_forms` feeds the local handles and
friendly-name slugs into the roster paragraph; `bot_mode_dm` lists ambiguous
candidates with the same forms.
- `tui_gateway/methods_bot_relay.py::bot_relay.deliver` re-stamps the sender
before the transport branch (the recipient knows `from_connection`; the
sender does not know its own Desktop id).
Slim redo of #103767's Python half (@fangliquanflq): that PR always
qualified every remote form with its connection id and reworked the Desktop
mention Map; this keeps bare forms where they are unambiguous and leaves the
Desktop picker untouched.
Part of #103731
Salvages #103767
Co-authored-by: fangliquanflq <fangliquanflq@users.noreply.github.com>
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
The cross-machine reply waiter was spawned as `python -c '<generated
source>'` through terminal_tool from the SENDER's turn. The approval gate
flags inline interpreter code ("script execution via -e/-c flag"), and
`approvals.single_query_mode` defaults to `deny` for the one-shot `-Q` turn
a bot replies from — so exactly when bots talked to each other (a teammate
answering a DM from its own delivery turn), the waiter was refused. The
message was delivered and, since 54d7f75590, honestly reported as queued with
a notification_error; the reply still never woke the sender, and every
cross-machine bot-to-bot exchange ended after one hop.
Spawn it as `tools/bot_mode_dm.py --wait-reply <reply file> <label>
<budget>`, the shape the local delivery runner already uses and the gate
lets through. The runner's --wait-reply mode is the old generated loop,
stdlib only, same three notifications and exit codes. Roster fields ride as
shlex-quoted argv instead of source text, so a hostile handle or connection
id is a label, not code, and the Windows path rewrite is the delivery
runner's (forward slashes for Git Bash), replacing the raw-string trick.
teknium1 suggested this follow-up when closing #97051.
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:
- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
scalar snapshot reads as empty / returns the existing 5000 error instead of
violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
_combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
unchanged.
Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.
(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Both bot-to-bot DM retry gates (tools/bot_mode_dm.py::_run_local_turn and
tui_gateway/methods_bot_relay.py bot_relay.deliver) classified a failed turn from
`(proc.stderr or proc.stdout)`. A failed `hermes … -Q` turn prints the provider
prose on stdout (it is the turn's final_response) and `session_id: …` on stderr
on every run, so the classifier only ever saw the banner: every 429/5xx/context
overflow classified `unknown`, the policy-gated re-run never fired, and the relay
sender always got `reason: unknown`. Classify from both streams in the order the
CLI writes them (tools.bot_failure_reasons.turn_failure_text) at both sites.
Classifying alone double-delivers: the failed attempt's turn-start persist already
wrote the DM as the Bot Chat's unanswered tail row, and the re-run (a fresh
process replaying the same session + payload) appended a second identical copy.
The re-run's child env now carries HERMES_RESUME_UNANSWERED_TURN=1
(tools.bot_relay.retry_turn_env); the quiet one-shot consumes it before the turn
(hermes_cli.quiet_single_query.adopt_unanswered_turn) and re-stages the
identical unanswered tail row as the pending CLI user dict, stamped durable, so
_stage_turn_user_message reuses it and the flush writes no second row. Opt-in
by design: an identical tail alone cannot tell a re-run from a person's
deliberate re-send.
Slimmer redo of #105529 by @jonpol01 (same diagnosis and shape; the marker is
consumed at the CLI seam that already reads the dispatcher's env instead of
inside agent/turn_context.py).
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
The Desktop delivers each target's claimed envelopes in the order the drain returns them, one
turn at a time, and says so at the call site; the user guide states the same guarantee
("messages to the same Bot are delivered in order, one turn at a time"). The claim sorted the
outbox directory by filename, which is uuid4().hex — so the order handed to the Desktop was
random, and a sender's second instruction to one agent could be delivered before its first.
Sort by the mtime _sweep_stale already treats as an envelope's age, with the name only breaking
ties. The whole-second created_at field cannot separate two DMs sent in the same second.
The test forces the filenames into the reverse of the send order, which random ids reproduce
half the time: it fails on the filename sort, on newest-first, and when the mtime is dropped.
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:
- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.
tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.
Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
The HERMES_SESSION_* prefix match also stripped non-identity knobs
(HERMES_SESSION_STALL_TIMEOUT tunes gateway stall watchers). The strip now
uses gateway.session_context's session env names (_VAR_MAP) — synced with
the session binding surface as vars are added — imported unguarded like
tools/approval_context already does on a hotter path; the except fallback
was an unreachable branch whose 3-name tuple silently narrowed the strip.
Test pins STALL_TIMEOUT as the kept negative case.
Quiet chat -Q inherited the dispatcher's HERMES_SESSION_KEY, so a nested
message_agent notify was addressed to the grandparent and B never woke.
Bind this session's key, strip inherited session identity from the delivery
child env, and continue owned notify completions in-process before stdout.
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
bot_relay.deliver accepted from_profile, from_handle and from_connection from any admitted JSON-RPC client. The handler now refuses them with error 4095 when the calling transport carries a browser login identity, since a logged-in browser never relays for another connection. The DeliveryAuthor docstring now says the author is trusted because an admitted client relays it, not because the sender is verified.
delivery_turn_author kept the bare bot:<profile> id when the sender's connection was the Desktop's own "local", so a DM relayed from that machine collided with the recipient's profile of the same name. A relayed DM always crosses gateways, so the connection id is now part of the id whenever the Desktop sends one, and only the direct message_agent path in tools/bot_mode_dm.py stays bare. The relay.ts and session_auto_continue.py comments added earlier are cut to one line each.
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
When the target Bot Chat is already open on this gateway, the relay handler delivers through
`prompt.submit` with `queued: true`, and that branch dropped the envelope's sender. The model
still saw the text prefix, but the turn reached the agent unattributed, the exact case the
subprocess branch fixes.
The relay handler now stamps the author on the submit as a `DeliveryAuthor`, an in-process object
a JSON client cannot build, so `prompt.submit` accepts it the way it accepts a hosted-room callback
and refuses a dict with error 4124. The busy queue keeps an authored envelope in its own slot, the
drain hands the author to the turn runner, and the runner passes it to an agent that declares the
keyword. A plain prompt after an authored dm carries no author.
Local deliveries to a desktop-owned Bot Chat take the live-owner mailbox instead. The admission
intent and the mailbox record now carry the author, a retry under the same id with a different
author is refused, and the owner gateway hands the author to the turn it runs. Isolated compute
turns still run unattributed, because the compute-host frame has no author field.
a bot dm arrived as an ordinary user message. the only trace of the sender
was the "Message from" text prefix, which the model reads and nothing else
does. the recipient's memory provider saw its own configured user.
message_agent now passes the sender as {"id": "bot:<profile>", "name":
<handle>, "is_bot": true} to the delivery runner (--author <json>), which sets
HERMES_TURN_AUTHOR on the recipient one-shot only. the -Q turn reads it and
passes turn_author into run_conversation. the desktop relay forwards the
envelope's from_profile/from_handle to bot_relay.deliver, which sets the same
variable on its delivery turn. the runner drops any inherited author first so
a delivery without one stays unattributed. the text prefix is unchanged.
The sender-side waiter gave up at 900s while the Desktop held bot_relay.deliver open for 1500s, so a turn finishing between minute 15 and minute 25 wrote a reply nobody read. REPLY_WAIT_SECONDS now rebuilds the Desktop budget from the same numbers and waits 60s past it. The two turn constants move into tools/bot_relay.py so the gateway handler and the waiter share one definition.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Read text files with the encoding utf-8-sig so a BOM at the start of a
file does not cause a Unicode decode error (Windows editors add BOMs).
Reconstructed from ethie/pm commits 48a32b135b + 013219e814 onto the
current upstream/main base: only the utf-8 -> utf-8-sig transforms were
carried (370 exact line pairs across 205 files); pm-rename hunks that
rode in the original commit were left to the pm-store commit, and
utf8sig hunks entangled with content changes ride their owning commit.
Rebuilt on ethie/pm-clean off ac6c8028e0 (upstream/main).
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:
- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
try shutil.which('hermes') before the bare-name fallback, so
environments with a PATH but no venv sibling resolve exactly what an
interactive shell would. Platform test switched os.name -> sys.platform
('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
errors='replace' on both subprocess.run sites — without them the
child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
with which=None, and encoding-pin assertions in the deliver transport
test.
Refs #93590, #93597, #93601
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):
1. waiter_command embeds the reply path in generated python -c source
with !r. repr escapes each backslash, but the Windows execution layer
folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
and SyntaxErrors the whole waiter script. Raw-string literals keep
the folded single backslash a literal; POSIX paths have no
backslashes so the prefix is a no-op there, and \' inside a raw
literal still cannot terminate the string, keeping the #93091
injection defense intact.
2. local_delivery_command hardcoded "hermes", relying on PATH — absent
in service contexts (systemd units, desktop launchers, non-login SSH
shells), so delivery died with ENOENT. It now resolves the CLI next
to this gateway's own interpreter (venv bin/Scripts sibling,
hermes.exe on Windows) with a bare-name fallback. The #93091
per-profile turn-lock recognition in bot_mode_dm now matches the CLI
element by basename (split on both separators) so resolved absolute
paths still take the lock instead of silently bypassing it.
Fixes#93590
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:
- Desktop relay drain forwards bot_relay.deliver's error.data.reason
into bot_relay.reply (and prefers it for the attention badge over
free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
text, so the completion notification the sending agent receives is
machine-branchable.
Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.
bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
wrapping it in --run-delivery would make the child contend with its
parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
loaded CI runners; the wait-duration message assert matches ~Ns generally.
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:
- tools/bot_relay.py: expose the 6h stale sweep as
cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
into gateway housekeeping — previously it ran only when the Desktop
drained the outbox, so plaintext envelopes/replies queued while the
Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
waiter deliveries, which have no plaintext DM tempfile to reclaim.
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.
Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.