Widen PR #87757 to cover the ZIP path: _update_via_zip() also calls
_finish_dashboard_update_cleanup() but never runs _reload_config_modules,
so the Windows git-broken fallback would still crash with the same
ImportError (cannot import name 'bounded_probe_run' from the stale cached
hermes_cli._subprocess_compat).
- new _reload_process_scan_modules() called inside
_finish_dashboard_update_cleanup itself, so every current and future
call site is covered; reloads dependency-first
(_subprocess_compat, then dashboard_procs)
- reload failures log at warning (a miss surfaces seconds later as an
ImportError in the same process)
- regression tests: reload-before-kill ordering, node-failure skip,
stale-module symbol restoration (the exact #87134 boundary state),
nonfatal reload failure, and the #87757 reload-list contract
Attribution mapping for the PR #84982 salvage. The commit email is not
linked to a public GitHub account, so contributor_audit --strict fails
without it; login confirmed from the PR author field.
Regression test for the _ref_map merge (salvaged from #79515): the live
0.19.3 driver splits action refs into refs[] while content_refs re-lists
every node with empty actions; the empty entries must not clobber the
action-bearing ones. Caught live: every typed click refused with
browser_ref_stale until the merge fix.
Custom providers could only authenticate from a static credential (inline
api_key or a key_env env var). Enterprise gateways -- SSO/OIDC brokers, cloud
IAM, internal auth proxies -- issue short-lived bearers instead, so a value
copied into .env is stale within the hour: long sessions start returning 401s
and the user has to restart or run an external cron that rewrites .env.
The existing `secrets.command` source does not cover this: it runs once per
process at startup (subsequent calls are no-ops by design), so it cannot
re-mint a credential mid-session.
Add providers.<name>.key_cmd: a command that prints a token, wrapped at
resolution in a zero-argument callable. Both wire clients already accept a
callable api_key and invoke it per request (the Entra ID path established
this), so chat_completions, codex_responses and anthropic_messages all work
unchanged and always send a fresh credential. The callable also routes the
Anthropic client through its per-request Authorization hook, which is what
OAuth-gated gateway routes require -- so no per-vendor auth wiring is needed
anywhere in core.
- cached until shortly before the advertised expiry (60s leeway), so the
helper runs about once per token lifetime rather than once per request
- expiry is read from the OAuth 2.0 relative `expires_in` when present, and
otherwise from an absolute ISO 8601 deadline (`expiry`, `expiresOn`), which
is what CLI token helpers commonly print. Reading only `expires_in` treated
those helpers as advertising no TTL at all, cached their token for the life
of the process, and returned 401 on every request once the real deadline
passed. ISO parsing reuses hermes_cli.auth._parse_iso_timestamp rather than
adding another datetime parser.
- no synthetic expiry: when no TTL is advertised, or the advertised one is
unparseable or already past, the token is used and refreshed on 401 instead
of re-minted on an invented schedule
- stdout contract matches OAuth 2.0 token endpoints and existing agent
helpers (bare token or JSON access_token/expires_in); multi-line output is
rejected rather than guessed at, so a misconfigured helper surfaces as a
clear error instead of a corrupt-credential 401
- precedence: explicit --api-key still wins; otherwise key_cmd beats a
static api_key/key_env on the same entry
- failures never include the helper's output (may hold a partial token) or
the command string (may embed a client secret)
Resolution happens on two paths. agent/auxiliary_client.py resolves named
custom providers itself rather than calling _resolve_named_custom_runtime, so
key_cmd is honoured in both: wiring only the runtime resolver leaves the main
agent turn working while every auxiliary call (title generation, compression,
vision, embedding) falls back to the no-key-required placeholder and 401s.
Precedence is identical on both paths, so one config entry cannot yield two
different credentials depending on which resolver the caller reached.
Closes#84162
Signed-off-by: LordMelkor <kray@block.xyz>
Rework of salvaged PR #6372 (@ag9920) onto current main:
- /save promoted from CLI-only JSON snapshot to a cross-platform session
export: `/save [json|md|html] [filename] [redact]` on CLI and every
gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
(hermes_cli/session_export.py + session_export_html.py) instead of the
PR's new hermes_state formatter — new helpers normalize_save_format /
render_session_for_save / default_save_filename are shared by both
surfaces.
- `redact` arg runs the export through the force-mode secret redaction
pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
filenames with basename, and lands in gateway/slash_commands.py (the
handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
single-session HTML export.
Closes#4249. Closes#51200.
The inflight-turn journal can outlive the turn it recorded (reclaim,
reconnect or restart races skip the settle that clears it). On session
resume the fold then re-appends journaled assistant rows to a transcript
that already holds the committed replies, so the conversation ends with
duplicate answers in scrambled order. The fold also carried the stale
entry's streamId onto the resumed state on an idle resume, which kept the
journal entry alive (persistInFlightTurnState only clears when streamId is
null) and re-folded the same tail on every open.
Detect text-level staleness before the append path: when every recoverable
journaled assistant row already exists as committed text in the base
transcript, treat the entry as caught up and clear it. Only keep a stream
target when the resumed session is genuinely running (keepPending), so an
idle resume self-heals instead of re-folding.