Fence owner discovery by profile, lease and loopback endpoint, and hand the authenticated URL to the existing Ink transport without acquiring a competing lease. Distinguish lease age from turn activity in unsupported-owner recovery.
Client slice only: requires the integration runtime to advertise shared_runtime_url and provide the session-attach handshake. Classic CLI attachment remains an integration gap.
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted
by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The
request path has honoured it since it landed, but the picker resolved probe
credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed
`/v1/models` with an EMPTY key.
Against an authenticated endpoint the probe 401s, discovery returns nothing,
and the provider falls back to its single configured default model. The
picker shows ONE model, indistinguishable from an endpoint that genuinely
serves one — while inference keeps working, because that path mints
correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models
discovered with an empty key, 26 with the minted token.
Both picker probe sites already funnel through `_entry_credentials()`, so
the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and
each site falls back to `resolve_probe_token()` after api_key/key_env. An
explicit static key still wins, so existing configs are unaffected.
The identity is keyed on the COMMAND, never the minted token: the token
rotates on every refresh, so keying on its value would change the group
fingerprint constantly and force a re-probe on every open. Two entries on
one URL with different helpers still get distinct rows.
`resolve_probe_token()` lives in agent.command_token_source, which already
owns key_cmd minting, and shares the CommandTokenSource cache with the
request path — a cache read, not a fresh sign-in. Fail-closed: a helper
needing an interactive sign-in degrades to today's empty-key behaviour
rather than taking down every other provider's row.
`_model_flow_named_custom` (the `hermes model` setup flow) is the sibling
path — it builds its own `Authorization: Bearer` from the same incomplete
resolution — and is fixed the same way, with one ordering constraint: the
value persisted to config.yaml is computed BEFORE the mint, so a short-lived
bearer can never be written back to shadow the key_cmd meant to re-mint it.
Tests drive the real code paths and assert on the credential each probe
receives rather than on function source, so a semantics-preserving refactor
does not fail them. Verified they fail with the fix reverted.
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Carry the owning task's notification subscriptions independently of dependency
edges, within the creation transaction. Prefer its durable session over worker
and request-local sessions while preserving explicit overrides. Cover worker
CLI create and built-in decomposition, and retain conversation route anchors.
Auto-subscribe no longer upgrades an inherited passive subscription.
Slim adaptation of Christopher-Schulze's session-precedence fix in #85687,
expanded to durable subscription provenance and sibling creation paths.
Related: #85575, #85687
Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then
58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess
probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in
children and a plain CLI negative control, with recording transport only.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.
_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).
The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.
Fixes#104893
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.
Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
``gated_cache`` bypassed both the fresh-hit and the stale-while-revalidate branches of
``cached_provider_model_ids`` whenever the disk entry contained Astra. For an entitled
account that is every entry: live discovery rewrites the entry with Astra → next call is gated
again → a blocking /models round-trip on every picker open, for exactly the providers the user
is most likely on. The parallel prefetch's staleness check didn't know about the gate either, so
it skipped the slug and the blocking call landed in the serial picker loop the prefetch exists to
avoid.
Gate on provenance, not contents: only live discovery ever writes Astra into a same-credential
entry (the static/offline paths filter it out), so a fresh entry IS the entitlement record. The
one filter that matters stays — a stale entry served because the refresh failed drops Astra, and
the entry itself is left intact so the next successful fetch restores it.
The regression test now pins both halves: fresh entry served with Astra and zero live calls;
failed refresh past the fresh window serves the entry minus Astra.
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.
``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.
``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.
Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.
Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.
The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.
Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
Based on the earliest batching proposal by BrunoBza (#104686) and JoaoMarcos44 structured correction (#104703). Slim redo into topical siblings and shared TUI poller/post-turn routing. Reported by Xipong (#104671). Local shell-to-loopback probes demonstrate 12 to 1 turn dispatches; suites await the campaign test lock.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
`hermes_cli/banner.py::_git_run` is the shared spawn path for every banner and
passive update-check git probe (rev-parse, rev-list, remote get-url, ls-remote).
It carried the UTF-8 text contract but not `creationflags=windows_hide_flags()`,
so on Windows each probe run from a GUI-hosted backend (desktop-spawned
`hermes serve`, `tui_gateway` import kicking off `prefetch_update_check`) flashes
a console window. Every other short-lived helper in `hermes_cli/` already passes
the flag; this brings the banner path in line.
Surfaced by CI on this PR: the stray `git ... origin` spawn from the prefetch
daemon landed in `test_env_probe_run_hides_console_window`'s process-wide
`subprocess.run` capture and tripped its call-count assertion. The test now
scopes its assertion through the module's `_spawns` helper like its siblings,
so an unrelated daemon spawn cannot fail it (the production fix is what makes
that stray spawn carry the flag in the first place).
A/B: base `_git_run` -> no `creationflags` kwarg; fixed -> 0x08000000.
Add explicit RFC 8628 device login and oauth.flow selection while keeping
browser PKCE and the SDK runtime refresh path. Reuse issuer/resource
validation, configured client authentication and profile-scoped storage.
Only persist an approved, validated grant; never echo endpoint error bodies.
Slim redo of #104752 by @wjorgensen, replacing duplicate HTTP/storage
wrappers with the existing SDK and two real-wire invariant tests.
Refs #104742
Co-authored-by: Wes Hermes <weshermes@Wess-Mac-mini.localdomain>
ssh bypasses stdin=DEVNULL and GIT_TERMINAL_PROMPT: when a git child
dials an SSH remote whose host key is unknown, ssh opens /dev/tty
directly and its yes/no prompt steals the caller's terminal — exactly
what noninteractive_git_env exists to prevent. Pin core.sshCommand to
"ssh -o BatchMode=yes" at the config-injection layer so the ssh child
fails instead of prompting; an agent-authenticated ssh still succeeds,
and an explicit user GIT_SSH_COMMAND env var still takes precedence
(#104591).
The startup update check resolved `git remote get-url origin` with the
user's global config in scope while the subsequent fetch runs under
noninteractive_git_env (GIT_CONFIG_GLOBAL=/dev/null). A global
url.<https>.insteadOf rewrite therefore made an SSH origin masquerade as
HTTPS, the SSH-avoiding fast path was skipped, and the fetch dialed the
raw SSH origin — whose host-key prompt opens /dev/tty directly and
steals the CLI's keystrokes (#104591).
Probe the origin URL under the same isolated env so both sides observe
the URL the fetch will actually dial.
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.
Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Emit the one-shot reason marker outside the CLI facade; parse whole codes before falling back to legacy prose. Explicit coordination and unknown codes cannot be labeled target_busy.
Fixes#104784
Co-authored-by: William Echo <2054936695@qq.com>
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Salvaged from #104272. Preserve restart and fatal-exit policy while classifying the planned restart code as success. Earlier analysis in #13604 by Justin Kausel.