One-time, explicit copy of provider credentials from one profile to other
profiles from a desktop page. Owner submission (rule 5): repo and tag v1.0.1
pinned at 1599d21d1bce11d94424f749531dadf8d619f42c.
The pin merged in #115599 (eb1ac9b) carries a security gap found during that
review: the run_command_evaluator capability was never consulted, because
AutonomyPolicy was never constructed anywhere in src/. A project at autonomy
level 0 still executed command-type objective evaluators, since the security
allowlist constrains which executable may run and never whether running is
permitted at all.
4f733a8 wires the policy in and fails closed below level 3.
Verified at the new sha against a clean NousResearch/hermes-agent main
checkout: doctor_plugin ok with no findings, read_declaration resolves both
deps from pyproject. Upstream repo gates all pass.
The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
Split the single connection row of the gateway's shaped provider-error reply
into three: an interrupted established connection (reset/EOF/RemoteProtocolError),
a refused/unroutable endpoint (the case #86570 wrote the "not running or is
unreachable" wording for), and a cause-free SDK ``APIConnectionError`` that
supports neither diagnosis. A reset says nothing about whether the endpoint is
up, so telling the user to restart a server that just answered sends them to
debug the wrong thing (#116323).
Selectively adapted from Lei-k/hermes-agent#16 (497cc47556b1) via PR #109701,
rebased onto the current reply contract (rate-limit > auth > policy > connection,
every reply names a slash command, no operator jargon).
Transient provider outages no longer end the turn: bounded auto-recovery ladder after retries and fallback, visible on CLI/TUI/gateway/API/cron (#85426, #107307)
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
Desktop lint: /<script[\s\S]*?<\/script>/gi in a feed sanitiser scored as
'script injection' and failed pinned-source-validate for rss-reader. Mask
JS regex literals for the markup-shaped rule only; <script in a string
literal (an innerHTML payload) and createElement('script') still fail.
Install scanner: "printenv" as a whole-string entry of a read-only
allowlist (frozenset({..., "printenv"})) fired dump_all_env high →
caution on hermes-jev. Extend the literal-token demotion: a token that is
the ENTIRE quoted literal on a line that executes nothing steps down like
an alternation member; "sudo" inside subprocess.run([...]) and
os.system("printenv") keep high.
A/B vs origin/main: attack probes identical (23 rows), in-tree sweep 319
entries 0 worse/0 changed; both new tests red on base. Bumps
PLUGIN_SCANNER_VERSION to v6 so cached caution verdicts refresh.
Signed-off-by: teknium1 <teknium1@users.noreply.github.com>
Chained /v1/responses turns no longer replay earlier tool calls or double the stored history after history repair/compaction (#89891, rebuild of #70695)
Two corrections to 0b9a9a0f5c.
The `if allow_split_turn else -1` gate did not skip the scan it claimed to:
`_ensure_last_user_message_in_tail` runs the identical lookup as its first
statement, so the micro-compaction pass paid the same scan and the batch path
paid it twice. Reverted to the unconditional call, which also removes the `-1`
sentinel every reader had to reason about.
The dedupe stopped at two of five spellings of the same predicate. Three more
sites inline it: the in-flight replay's "a real request follows the summary"
check, the handoff-candidate admission test (as its negation), and the merge
pre-check. All now call `_is_real_user_turn`. The one remaining inline pair is
deliberately different — it also requires `_is_real_user_message`, which rejects
metadata-flagged scaffolding this predicate cannot see.
Equivalence: all three predicates are pure, so the negated and reordered forms
are the same test; 221 tests pass across the compressor/anchor/micro-compaction
files, including the source-shape anchor-order test that the call shape here
leaves untouched.
Follow-up to the oversized-turn exception merged in #116181.
- `_find_last_user_message_idx` and `_real_user_indices_desc` each spelled out
the same actionable-and-not-synthetic predicate; both now call one
`_is_real_user_turn`. No behaviour change — same two classmethods, same rows.
- The newest-user index is only read by the split exception, which rolling
micro-compaction disables, so the scan is skipped on that pass instead of
running and being discarded.
`_is_actionable_user_turn` / `_is_synthetic_compression_user_turn` are pure, so
the dedupe is equivalence by construction; the row set each scan returns is
unchanged.
The OpenAI-compatible SSE writers carried tool progress and reasoning but no
lifecycle status, so an API client waiting through a provider outage (now the
auto-recovery ladder) saw a silent socket with no way to tell "waiting on the
provider" from "hung". _spawn_stream_agent wires status_callback into the
agent and both writers (/v1/chat/completions, /v1/responses) emit
`event: hermes.status` with {kind, text}, redacted like every other API-bound
error text. The Responses writer's tag dispatch moves to a table so the new tag
does not grow an if/elif ladder.
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.
Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.
Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
hermes doctor validated model.provider but never the auxiliary blocks, so a
task whose provider could not be resolved (and therefore silently ran on the
main model) passed as healthy. The config check now runs each routed block
through resolve_runtime_provider — the entry point the tasks themselves use —
and turns a resolver error into a finding with the resolver's reason; the
green line shows the resolved provider@host so a block that fell to the public
default endpoint is visible too. Docs: the openai direct-API alias, its
endpoint precedence, and the new warning/doctor behaviour.
_resolve_review_runtime swallowed every resolver error at DEBUG and returned
the parent runtime, so a dead auxiliary.background_review block ran reviews on
the main model indefinitely with nothing in agent.log and nothing on screen
(#116055: the configured model appeared zero times in session_model_usage).
The fallback now logs a WARNING naming the provider, model and reason on every
review and pushes the same message once per agent through _emit_warning — the
rail every surface (CLI, TUI/Desktop, gateway) already renders for the
reasoning_effort notice. curator and the MoA slot resolver had the identical
debug-only swallow; both are WARNING now.
auxiliary.<task>.provider: openai was expanded to custom + the user's OpenAI
endpoint only by agent/auxiliary_client.py (compression, vision, title
generation). hermes_cli/runtime_provider.py::resolve_runtime_provider — the
path background_review, curator, MoA slots and delegation use — had no such
expansion, so "openai" hit auth.resolve_provider's registry lookup and raised
"Unknown provider 'openai'". The alias table now lives once in
runtime_provider_custom.py (the direct-alias/custom sibling) and both paths
call it; resolve_runtime_provider applies it before the ladder.
_host_gated_env_key_candidates also pairs OPENAI_API_KEY with a base_url that
is exactly OPENAI_BASE_URL: the alias lands on that proxy when no block
base_url is set, and the key was issued for it — the host gate otherwise sent
the "no-key-required" placeholder there while the aux-client path used the key.
Slim redo of #116083 (same direction: shared alias, applied in the runtime
resolver) without the extra key gate and effective_provider threading.
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.
Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.
Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
Add ``resume_thread_id`` to CodexAppServerSession: when set, the first
ensure_started() issues ``thread/resume`` (with the same cwd / personality /
developerInstructions / model params thread/start sends) instead of starting
an empty thread, and verifies codex handed back the requested id. A refused
or mismatched resume raises the typed CodexThreadResumeError once; the next
ensure_started() falls through to ``thread/start`` on the same handshaken
client (initialize now runs once per client, not once per attempt).
Why: the thread id lived in memory only, so every new AIAgent for the same
Hermes session — a later API-server request or the first turn after a
restart — started a fresh codex thread and the model lost its own memory of
the conversation (#100531). The runtime decides the fail-closed policy; this
adapter only speaks the wire contract (verified against codex-cli 0.147.0:
thread/resume{threadId,...} -> result.thread.id; unknown id -> -32600 "no
rollout found"; killed writer -> -32600 "already has an active writer").
Salvaged from #103352 (thread/resume + id cross-fill + mismatch guard).
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.
The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.
Refs #83080
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.
One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.
Part of #97548
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.
Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.
Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).
Fixes#116191
Custom providers already declare their wire protocol (api_mode / legacy
transport) on the entry, but nothing outside routing could read it: context
length resolution keyed Codex detection on the hostname, which a proxy on
127.0.0.1 never matches. get_custom_provider_api_mode() returns the canonical
transport of the first entry serving a base_url (None custom_providers loads
config, like the sibling context_length helper) so metadata decisions can
follow the route instead of the host.
Slim port of the helper from #116262 by @JoaoMarcos44.
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.
Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.
Part of #116213