The #70760 commit widened _redact_process_result(result, *, task_id=) and passed the
caller's task_id from _handle_process. tests/gateway/test_completion_delivery.py
monkeypatches the function with a one-arg stub (CI red on #118845), and the caller's id
is the wrong identity anyway: a poll/log/wait result belongs to the process OWNER, whose
task_id the session record already carries. Resolve it inside the function from
result["task_id"] or the session looked up by result["session_id"]; the test pins the
owner's id ("t1") instead of the caller's.
A plugin registered on transform_terminal_output only ever saw foreground
`terminal` output: tools/terminal_tool_result.py::_apply_output_transform_hook
runs from finalize_foreground_result and nowhere else. Background output
reached the model through a different seam — process_manage poll/wait/log/kill
results, `list` previews and the completion/heartbeat/watch notifications all
pass through tools/process_registry.py::_redact_process_result — which redacted
but never transformed, so a fleet redaction or summarising plugin silently did
nothing for backgrounded commands.
Apply the same hook helper at that shared seam (one new
transform_process_output wrapper) and at the two gateway agent-notify sites
that read session.output_buffer directly. The order matches the foreground
path and teknium1's review note on #71401: hook first, redaction after, so a
replacement the plugin returns is still masked. returncode is None while the
process runs; env_type is not recorded per process and is passed empty.
Not changed: the spawn acknowledgement ("Background process started") carries
no command output, and the persistent local shell already goes through
_run_foreground and was transformed — the issue's reading of that branch was
wrong; the real gap was the process_manage/notification seam.
Fixes#70760
Slim redo of #71401 (Christopher-Schulze): same seam and ordering, without
the ANSI-stripping relocation and render helper.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
(cherry picked from commit 84661de52f78ccb84054e1ded92b07166e40546d)
List reconciliation could race the live stdout reader, publish a completion before buffered descendant output was ingested, then close the pipe underneath that reader. For selectable POSIX pipes, mark the direct child exited but ask the reader to perform the final drain and remain the sole completion publisher. Keep the existing fallback for readers that cannot be coordinated.
`terminal(background=true, heartbeat=N)` emits a `heartbeat` event every N seconds
(floor 60) carrying only the output produced since the previous one, plus the usual
completion notice. The agent stays current on a long bounded job (merge train, full
test suite, deploy) and reacts to a failure within N seconds instead of at exit.
Why: `notify=true` fires once at exit, so multi-hour merge trains ran silently — the
parent's prompt cache went cold and every "what's up" cost 5–10 rediscovery calls;
30 days of batch sessions show 4,930 hand status polls (14% of all tool calls).
`notify=[pattern]` cannot serve this: its lifetime cap (8) exists precisely to stop
periodic output from flooding the agent.
Mechanism: one daemon timer thread for every heartbeat session (reader threads block
on the pipe and cannot keep time); delta output via a total-ingested counter over the
rolling buffer; heartbeats only where the completion notice can be delivered (same
async-support/subagent gates), never after exit. `heartbeat_seconds` is checkpointed.
Rendered on CLI, gateway and TUI through the existing process-notification paths.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The live-work docks retire a finished background process ~60 s after it
ends; the registry only knew started_at, so the age of an exit was not
observable. _move_to_finished is the single choke point every exit path
(reader loop, reconcile, kill) passes through, so the stamp lives there.
completion_reason rides along so a killed process can read "killed"
instead of "exit -15".
process(action='wait'/'poll'), the already_exited and kill snapshots still used a
hard-coded 2000-char tail with no output_cut, so the polling fallback bot-mode.md
documents for api_server/one-shot senders delivered a silently truncated reply.
One _completion_output() helper now sizes every delivery route.
SIGKILL/taskkill are asynchronous; poll()/isalive() right after kill() still
say alive, so the exact #115490 scenario (a SIGTERM-ignoring child that needed
escalation) was reported as 'Kill incomplete' on every run (10/10 in a live
probe) and the session stayed in _running until the reader thread finalised it
as a plain exit. Re-probe survivors for up to 1s before deciding.
Post-kill, check the session tree (Popen poll, PTY aliveness,
PID-scope identity plus descendant scan). Survivors keep the session
running instead of persisting a false killed receipt and pruning it.
Fixes#115490.
A bot DM's reply comes back as the delivery process's completion notification,
which carried the last 2000 characters of that process's output — the right
tail for a build log, but a bot may send 16,000 characters (MESSAGE_MAX_CHARS)
and its teammate's answer routinely runs longer than 2000. The sender got the
tail of the reply, header included in the cut, with nothing saying so, and
relayed it as the whole answer.
The registry now sizes a completion per process (ProcessSession
.completion_output_chars, checkpointed) and declares a cut (output_cut) when
the output did not fit; the notice names the cut and the process log that has
the rest. tools/bot_mode_dm._spawn_delivery — the one spawn point for every
DM lane (local, live-owner, Desktop relay waiter) — asks for MESSAGE_MAX_CHARS
plus header room. Everything else keeps the 2000-char tail.
Fixes#115334
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.
Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.
The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).
Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.
Fixes#113667
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).
#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.
- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
(RuntimeError subclass) so the dispatcher can tell a host refusal from a
card failure.
- _record_task_failure(infrastructure=True): run + event carry
`infrastructure: true`, consecutive_failures is left alone, the breaker
never trips; the card stays `ready` with the linger remedy in
last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
is such a refusal inside the rate-limit cooldown window, so a dead bus is
retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
managed gateway.
#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.
- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
is scope-wrapped when the user bus is reachable. Without a bus it degrades
with a loud once-per-process warning naming the consequence and both
remedies (enable-linger, KillMode=process) rather than refusing — the
unit's KillMode/lifetime is unknowable and a long-lived sequencer without
linger must keep working.
Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.
Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
scoped_spawn_lost_user_bus() re-derived the user bus from an empty base env,
so it only ever looked under /run/user/<uid>. The worker itself is launched
with systemd_user_bus_env(worker_env), which honours a configured
XDG_RUNTIME_DIR. On a host whose bus lives outside the default runtime dir,
any unrelated systemd-run exit was therefore misread as "bus gone": the job
error named a missing bus that was still there, the cached scope verdict
flipped to False, and the next 60s of cron fires dispatched without cgroup
isolation.
The check now takes the spawn env, drops the bus address the spawn already
carried, re-derives from that, and decides on the DBUS_SESSION_BUS_ADDRESS
key rather than on dict truthiness (a non-empty env with no bus was never
"bus present").
Review finding: scoped_spawn_lost_user_bus used systemd_user_bus_env({}) instead of the worker's spawn env, so a configured XDG_RUNTIME_DIR made unrelated wrapper exits flip the scope cache to unscoped dispatch.
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.
Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.
Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
`ProcessRegistry._move_to_finished` pops the session out of `_running`, then
saves the receipt, releases handles and writes the checkpoint, and only THEN
enqueues the completion and sets `_completion_event`. A quiet one-shot parent
whose turn ends inside that window called `wait_for_pending_completions`,
found nothing in `_running`, drained an empty queue and exited without the
follow-up turn. That is the CI flake in
tests/tools/test_completed_process_results.py::test_headless_terminal_result_survives_cli_exit
(`follow_ups == []`), which also hit unrelated branches.
Consider `_finished` sessions whose event is not yet set as pending too.
Repro: a 6s sleep before the enqueue plus a 2s delay before the parent's first
wait fails the E2E 2/2 on main and passes 2/2 with this change.
_release_finished_handles read the dataclass fields through getattr
fallbacks and swallowed every exception; use the real attributes, suppress
only OSError/ValueError on the pipe close (an stdin flush can hit EPIPE),
and rely on ptyprocess/pywinpty close() idempotence for the master fd.
The call-site comment claimed the reader had always drained the pipe;
on kill_process/_reconcile_local_exit the reader may still be reading,
so state what actually happens (its next read raises, the loop exits).
Widen #75162: _prune_if_needed() drops finished sessions (TTL expiry and
oldest-finished eviction at MAX_PROCESSES) — release their Popen/PTY
handles there too, covering sessions inserted into _finished without
passing through _move_to_finished(). The release helper is idempotent,
so double-close on the normal path is a no-op. Adds two tests: prune
releases handles of dropped sessions, and a still-running session's
pipe stays open.
Finished sessions retained their subprocess.Popen pipe objects (and PTY
masters) until the finished-process TTL (FINISHED_TTL_SECONDS, default 30
minutes) elapsed. Under heavy background churn — deployments, archivers,
watchers — finished-but-unpruned sessions accumulated one open pipe FD
each, exhausting the gateway process's file descriptor budget and
surfacing as a 'file descriptor limit' error on new background spawns.
The registry never rejects spawns (it prunes oldest-finished at
MAX_PROCESSES), so the real defect was the retained-handle leak, not a
registry-cap rejection. The fix closes each finished session's Popen
stdout/stderr/stdin streams and PTY master in _move_to_finished(), right
after the reader loop drains EOF. poll()/wait()/read_log() serve output
from the buffered output_buffer — never from the pipe — so the release is
lossless.
Tests: 4 new cases in TestFinishedHandleRelease — Popen pipes closed,
PTY closed, no-handle sessions safe, and poll() still serves buffered
output after the release. All 4 fail on main (reproduction) and pass
with the fix.
Programs run via terminal(background=true, pty=true) can block forever when
they probe their terminal — device-status (ESC[5n), window-size (ESC[18t),
cursor-position (ESC[6n), or DEC private-mode (ESC[?N$p) queries — because
nothing on the PTY master side answers, and the raw query bytes leak into
captured output.
- tools/pty_query_responder.py: incremental byte scanner that strips the
handled queries from PTY output (chunk splits included) and produces
bounded replies; everything else passes through untouched.
- tools/process_registry.py: wire the responder into _pty_reader_loop
(POSIX only — ConPTY answers its own queries); flush partial escape
tails at end-of-stream.
- tests mirror the codex fixtures plus a live-PTY E2E where a subprocess
blocks on ESC[6n until answered.
Carve the direct-child list reconciliation from Indigo Karasu's earliest
PR #60506 (2a96ae2cbf806ccdc3e9b584911774f32622f421), corroborated by
fangliquanflq's narrow #81385 (50ffd243d92627e4a03a3ee8427ad0f5e090c3ab).
Run the existing helper after task/session filtering and reuse the
idempotent owner-stamped completion path. Do not import cross-session
disclosure, bare-PID healing, forget RPCs, or reader rewrites.
A real-child regression fails before this change and passes after it:
the direct child exits while its descendant keeps writing to stdout;
listing reports exit without consuming the result or waiting for EOF.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).
wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.
Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.
Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
The warning fires once per process but was prefixed with the first
job's unit suffix, reading as a per-job notice for a host-level
condition. Drop the suffix; the message now says what applies to every
later dispatch.
Follow-up to the cherry-picked #102431 fix, addressing the review findings:
- The two real-helper scheduler tests ran the Linux-only helper unmarked
and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
scripts/check-windows-footguns.py --all (lint lane red). The remedy text
is now built once in the helper and passed into the warning, so the
"scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
once (helper level); default config still Popens externally and
`require_restart_safe_scope: true` raises (scheduler level, real helper).
Dropped the stubbed duplicate, the standalone config-raise test and the
in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
with the same `except Exception` guard as the sibling
`failure_nudge_threshold` read, so a config error no longer escapes
the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
`cron.require_restart_safe_scope` documented in the cron user guide.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).
Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.
The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.
Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.
Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:
- tools/process_registry.py, tools/environments/{modal,singularity}.py:
`_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
call time (same seam as `tools/skills_tool._skills_dir`, so the existing
`monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
the `_X_resolved` flags and the lifecycle reset keep their shape.
tools/file_tools.py drops its private `file_read_max_chars` memo and reads
the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
startup) is ignored in favour of the routed profile's config.yaml. Both
sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
the static schema text is profile-neutral and `dynamic_schema_overrides=`
rebuilds the `display_hermes_home()` / create-dir hint per
`get_definitions()`, so a routed profile's model sees its own paths.
Refs #95685.
Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.
Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.
Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
Execute the selected no-op rather than freeze its spelling, while rejecting
/bin/true to model the NixOS failure. Mark the regression Linux-only and
retain the current user-bus environment handling.
Consolidates the earlier NixOS scope-probe report and fix in #102587 with
the PATH-independent payload from #105436. The fallback resolver is not
needed when /bin/sh is used directly.
Co-authored-by: Scott Garrand <sgarrand@gmail.com>
Merge upstream b1f003e186 while preserving PM runtime ownership and
Python 3.14 worker startup, Windows signing, and macOS wait recovery.
Keep retired runtime modules deleted. Port upstream updater preflight
checks into the checkout strategy and preserve live build logging.
Carry checkpoint filename handling and process recovery into the current
module layout. Regenerate locks and adapt incoming platform test markers.
Focused Python and JavaScript tests, desktop and root-test typechecks,
conflict-path lint checks, lock validation, and retired-import checks pass.
The full test suite and packaged release builds were not run.
_user_systemd_socket_ready() accepts systemd/private alone, which is enough for
systemctl --user but not for the systemd-run --user that restart-safe workers
need; systemd_user_bus_env() requires the bus socket. Replace the uid threading
through five helpers with one _wait_for_target_user_bus(uid) that polls
/run/user/<uid>/bus, and move the post-enable wait + restart hint out of
_ensure_linger_enabled into _ensure_system_service_linger so the activity probe
runs only when linger was actually just enabled. Kanban applies the bus env
unconditionally like the cron sibling. Refs #104893.
A system-level gateway unit has no ordering against user@<uid>.service and
linger may be enabled after boot, so the bus can appear after the one-shot
adoption in run_gateway() ran. Derive XDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESS
fresh for the availability probe and every scoped spawn (cron worker, Kanban
worker, PTY/pipe terminal spawns, scope cleanup) so the 60s failure TTL can
actually recover. Refs #104893.
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.
_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).
The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.
Fixes#104893
A process that finishes while the child is alive needs no handoff, but if the child never
polls/waits/logs it, the result vanished: the completion notice is suppressed in the parent
and the child's summary never mentions it. Finalization now attaches exit code + output
tail as unread_completions, rendered in the parent's delegation notice.
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).
- process_manage(action="handoff", session_id, data="<purpose>"), children
only: process_registry.transfer_ownership flips owner_task_id/task_id/
session_key to the parent under the registry lock, so the completion is
stamped with the parent's owner at exit, passes the parent's sa- filter,
and is reaped by the parent, not the child. Cap 3 per child; an exited,
foreign, or non-child request is a tool error. The purpose rides the
event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
orphaned_processes on the result before cleanup kills the leftovers; the
parent's delegation block renders both.