Commit Graph

38482 Commits

Author SHA1 Message Date
teknium1
cdbbbfd882 chore(plugin-catalog): disclosure line (catalog review) 2026-09-19 19:13:07 -07:00
apoapostolov
a7f63109e6 feat(plugin-catalog): pin provider-status public edition 2026-09-19 19:13:07 -07:00
Apostol Apostolov
3ee2d10098 feat(plugin-catalog): pin provider-status catalog card 2026-09-19 19:13:07 -07:00
Apostol Apostolov
97f731020b feat(plugin-catalog): add provider-status 2026-09-19 19:13:07 -07:00
Brendan Ryan
4abd595f2a fix(plugins): update MPP catalog pin and payment disclosure 2026-09-19 19:12:30 -07:00
Brendan Ryan
f2a85c16f7 feat(plugins): add the Tempo MPP catalog entry
Co-authored-by: Derek Cofausper <256792747+decofe@users.noreply.github.com>
2026-09-19 19:12:30 -07:00
teknium1
f69161d437 chore(plugin-catalog): contributor map (catalog review) 2026-09-19 19:11:53 -07:00
Me
b22ebf4833 fix(catalog): pin reviewed Corpus kit and disclose marketing and PII 2026-09-19 19:11:53 -07:00
Me
63e8d3e32c chore(catalog): re-pin corpus to 5e1a5a7 (checkout/handoff copy fix) 2026-09-19 19:11:53 -07:00
Me
38bae549c6 feat(catalog): add Corpus plugin entry 2026-09-19 19:11:53 -07:00
teknium1
33bbd49fc6 chore(plugin-catalog): disclosure line, category -> tools (catalog review) 2026-09-19 19:11:16 -07:00
Abhilash Reddi
4666ce6d83 feat(plugin-catalog): add Adspirer 2026-09-19 19:11:16 -07:00
teknium1
2062388b64 chore(plugin-catalog): contributor map (catalog review) 2026-09-19 19:09:03 -07:00
Chloé
14802a586c fix(plugin-catalog): update pstack maintainer Zoeille → Cloeille
GitHub username changed. Updates repo URL, maintainer, docs_url,
contributor email, and pins sha to latest commit.
2026-09-19 19:09:03 -07:00
VGFreakXBL
58bc617c8b feat(plugin-catalog): add sticky-notes community plugin
Pin VGFreakXBL/hermes-sticky-notes at v0.2.0. Desktop-only SDK plugin
(no tools, hooks, env). Owner submission.
2026-09-19 19:08:27 -07:00
teknium1
95a183c6f5 chore(plugin-catalog): contributor map (catalog review) 2026-09-19 19:07:50 -07:00
tobenwarrior
2e58fb3b04 feat(plugin-catalog): add provider-copy community plugin
One-time, explicit copy of provider credentials from one profile to other
profiles from a desktop page. Owner submission (rule 5): repo and tag v1.0.1
pinned at 1599d21d1bce11d94424f749531dadf8d619f42c.
2026-09-19 19:07:50 -07:00
Sahil-SS9
c29e6b2f4b fix(plugin-catalog): re-pin hermes-project-stewardship to 4f733a8
The pin merged in #115599 (eb1ac9b) carries a security gap found during that
review: the run_command_evaluator capability was never consulted, because
AutonomyPolicy was never constructed anywhere in src/. A project at autonomy
level 0 still executed command-type objective evaluators, since the security
allowlist constrains which executable may run and never whether running is
permitted at all.

4f733a8 wires the policy in and fails closed below level 3.

Verified at the new sha against a clean NousResearch/hermes-agent main
checkout: doctor_plugin ok with no findings, read_declaration resolves both
deps from pyproject. Upstream repo gates all pass.
2026-09-19 19:07:20 -07:00
teknium1
30de041b01 docs: explain the three gateway connection-failure replies
The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
2026-09-19 18:24:11 -07:00
Halldrix
6d2765b755 fix(gateway): anchor errno/winerror connection markers with word boundaries 2026-09-19 18:24:11 -07:00
Lei-k
37c15650a7 fix(gateway): distinguish interrupted and unreachable model connections
Split the single connection row of the gateway's shaped provider-error reply
into three: an interrupted established connection (reset/EOF/RemoteProtocolError),
a refused/unroutable endpoint (the case #86570 wrote the "not running or is
unreachable" wording for), and a cause-free SDK ``APIConnectionError`` that
supports neither diagnosis. A reset says nothing about whether the endpoint is
up, so telling the user to restart a server that just answered sends them to
debug the wrong thing (#116323).

Selectively adapted from Lei-k/hermes-agent#16 (497cc47556b1) via PR #109701,
rebased onto the current reply contract (rate-limit > auth > policy > connection,
every reply names a slash command, no operator jargon).
2026-09-19 18:24:11 -07:00
Teknium
8a92051f20 Merge pull request #116328 from NousResearch/fix/boa-w3-new-reports-cron-codex
fix(cron): missing-credential preflight verdict names the profile and HERMES_HOME it read (#116213)
2026-09-19 14:33:58 -07:00
Teknium
059b3d416f Merge pull request #116353 from NousResearch/boa-w3-recovery
Transient provider outages no longer end the turn: bounded auto-recovery ladder after retries and fallback, visible on CLI/TUI/gateway/API/cron (#85426, #107307)
2026-09-19 14:33:20 -07:00
Teknium
8e0828a39c Merge pull request #116346 from NousResearch/fix/boa-w3-new-reports-aux-openai
auxiliary provider "openai" routes identically on every aux path; dead review routes are reported (#116055, salvage #116083)
2026-09-19 14:32:59 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
Teknium
abc21bb427 Merge pull request #116344 from NousResearch/boa-w3-session-guard
session.create refuses a model its provider cannot serve instead of a dead first turn (#96817, salvage #96845)
2026-09-19 14:32:15 -07:00
Teknium
b5fcf635dc Merge pull request #116343 from NousResearch/boa-w3-codex-thread
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
2026-09-19 14:31:55 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
Teknium
5322b4e637 Merge pull request #116335 from NousResearch/boa-w3-small
Exhausted connect retries now say which host, how many attempts and how large the request was (CLI, TUI/Desktop, gateway)
2026-09-19 14:31:14 -07:00
Teknium
1372de375a Merge pull request #116329 from NousResearch/fix/boa-w3-new-reports-codex-proxy-ctx
fix(context): proxied Codex routes (custom codex_responses provider, HERMES_CODEX_BASE_URL) resolve the 272K Codex window, not the 1.05M direct-API catalog (#116191, credit #116199 #116262)
2026-09-19 14:30:54 -07:00
teknium1
50e9cd7bd5 fix(plugin-guard): two intake false positives — regex literal <script, allowlist "printenv"
Desktop lint: /<script[\s\S]*?<\/script>/gi in a feed sanitiser scored as
'script injection' and failed pinned-source-validate for rss-reader. Mask
JS regex literals for the markup-shaped rule only; <script in a string
literal (an innerHTML payload) and createElement('script') still fail.

Install scanner: "printenv" as a whole-string entry of a read-only
allowlist (frozenset({..., "printenv"})) fired dump_all_env high →
caution on hermes-jev. Extend the literal-token demotion: a token that is
the ENTIRE quoted literal on a line that executes nothing steps down like
an alternation member; "sudo" inside subprocess.run([...]) and
os.system("printenv") keep high.

A/B vs origin/main: attack probes identical (23 rows), in-tree sweep 319
entries 0 worse/0 changed; both new tests red on base. Bumps
PLUGIN_SCANNER_VERSION to v6 so cached caution verdicts refresh.

Signed-off-by: teknium1 <teknium1@users.noreply.github.com>
2026-09-19 14:30:36 -07:00
Teknium
60e166cbf5 Merge pull request #116326 from NousResearch/fix/responses-turn-boundary-anchor
Chained /v1/responses turns no longer replay earlier tool calls or double the stored history after history repair/compaction (#89891, rebuild of #70695)
2026-09-19 14:30:05 -07:00
Teknium
a80ec24fb0 Merge pull request #116291 from NousResearch/fix/boa-partial-ultra-pill
fix(desktop,tui): reasoning pill says ultra sends max on this route instead of a distinct Ultra level (#61634)
2026-09-19 14:29:44 -07:00
Teknium
f22b7cedf9 Merge pull request #116288 from NousResearch/feat/boa-partial-429-retry-at-reset
feat(desktop): 429 usage-limit card can schedule one retry for when the limit resets (#98852)
2026-09-19 14:29:23 -07:00
kshitij
f88c6fc46e refactor(compression): finish the row-test dedupe and drop a gate that saved nothing
Two corrections to 0b9a9a0f5c.

The `if allow_split_turn else -1` gate did not skip the scan it claimed to:
`_ensure_last_user_message_in_tail` runs the identical lookup as its first
statement, so the micro-compaction pass paid the same scan and the batch path
paid it twice. Reverted to the unconditional call, which also removes the `-1`
sentinel every reader had to reason about.

The dedupe stopped at two of five spellings of the same predicate. Three more
sites inline it: the in-flight replay's "a real request follows the summary"
check, the handoff-candidate admission test (as its negation), and the merge
pre-check. All now call `_is_real_user_turn`. The one remaining inline pair is
deliberately different — it also requires `_is_real_user_message`, which rejects
metadata-flagged scaffolding this predicate cannot see.

Equivalence: all three predicates are pure, so the negated and reordered forms
are the same test; 221 tests pass across the compressor/anchor/micro-compaction
files, including the source-shape anchor-order test that the call shape here
leaves untouched.
2026-09-20 02:23:46 +05:30
kshitij
0b9a9a0f5c refactor(compression): dedupe the actionable-user row test, skip an unread scan
Follow-up to the oversized-turn exception merged in #116181.

- `_find_last_user_message_idx` and `_real_user_indices_desc` each spelled out
  the same actionable-and-not-synthetic predicate; both now call one
  `_is_real_user_turn`. No behaviour change — same two classmethods, same rows.
- The newest-user index is only read by the split exception, which rolling
  micro-compaction disables, so the scan is skipped on that pass instead of
  running and being discarded.

`_is_actionable_user_turn` / `_is_synthetic_compression_user_turn` are pure, so
the dedupe is equivalence by construction; the row set each scan returns is
unchanged.
2026-09-20 02:11:36 +05:30
teknium1
789d02811d docs(agent): list turn_recovery_autorecover in the turn-phase sibling map 2026-09-19 12:38:55 -07:00
teknium1
b9f29e2180 feat(api_server): stream agent status lines as hermes.status SSE events
The OpenAI-compatible SSE writers carried tool progress and reasoning but no
lifecycle status, so an API client waiting through a provider outage (now the
auto-recovery ladder) saw a silent socket with no way to tell "waiting on the
provider" from "hung". _spawn_stream_agent wires status_callback into the
agent and both writers (/v1/chat/completions, /v1/responses) emit
`event: hermes.status` with {kind, text}, redacted like every other API-bound
error text. The Responses writer's tag dispatch moves to a table so the new tag
does not grow an if/elif ladder.
2026-09-19 12:38:55 -07:00
teknium1
0752127c5e feat(agent): bounded auto-recovery ladder after retries and fallback are spent (#85426, #107307)
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.

Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.

Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
2026-09-19 12:38:55 -07:00
teknium1
15fec561db feat(doctor): resolve every routed auxiliary.<task> block and report the dead ones
hermes doctor validated model.provider but never the auxiliary blocks, so a
task whose provider could not be resolved (and therefore silently ran on the
main model) passed as healthy. The config check now runs each routed block
through resolve_runtime_provider — the entry point the tasks themselves use —
and turns a resolver error into a finding with the resolver's reason; the
green line shows the resolved provider@host so a block that fell to the public
default endpoint is visible too. Docs: the openai direct-API alias, its
endpoint precedence, and the new warning/doctor behaviour.
2026-09-19 12:30:32 -07:00
teknium1
9910041d0f fix(background_review): say so when the configured review route falls back to the main model
_resolve_review_runtime swallowed every resolver error at DEBUG and returned
the parent runtime, so a dead auxiliary.background_review block ran reviews on
the main model indefinitely with nothing in agent.log and nothing on screen
(#116055: the configured model appeared zero times in session_model_usage).
The fallback now logs a WARNING naming the provider, model and reason on every
review and pushes the same message once per agent through _emit_warning — the
rail every surface (CLI, TUI/Desktop, gateway) already renders for the
reasoning_effort notice. curator and the MoA slot resolver had the identical
debug-only swallow; both are WARNING now.
2026-09-19 12:30:32 -07:00
teknium1
956a8c843d fix(aux): provider "openai" resolves the same on the runtime and aux-client paths
auxiliary.<task>.provider: openai was expanded to custom + the user's OpenAI
endpoint only by agent/auxiliary_client.py (compression, vision, title
generation). hermes_cli/runtime_provider.py::resolve_runtime_provider — the
path background_review, curator, MoA slots and delegation use — had no such
expansion, so "openai" hit auth.resolve_provider's registry lookup and raised
"Unknown provider 'openai'". The alias table now lives once in
runtime_provider_custom.py (the direct-alias/custom sibling) and both paths
call it; resolve_runtime_provider applies it before the ladder.

_host_gated_env_key_candidates also pairs OPENAI_API_KEY with a base_url that
is exactly OPENAI_BASE_URL: the alias lands on that proxy when no block
base_url is set, and the key was issued for it — the host gate otherwise sent
the "no-key-required" placeholder there while the aux-client path used the key.

Slim redo of #116083 (same direction: shared alias, applied in the runtime
resolver) without the extra key gate and effective_provider threading.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-19 12:30:31 -07:00
teknium1
e77e24a6a6 fix: persist the codex thread id per session and thread/resume it across an API-server restart
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.

Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.

Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
2026-09-19 12:27:01 -07:00
Ben Awad
ba586ed4b5 fix: let CodexAppServerSession resume a stored codex thread
Add ``resume_thread_id`` to CodexAppServerSession: when set, the first
ensure_started() issues ``thread/resume`` (with the same cwd / personality /
developerInstructions / model params thread/start sends) instead of starting
an empty thread, and verifies codex handed back the requested id. A refused
or mismatched resume raises the typed CodexThreadResumeError once; the next
ensure_started() falls through to ``thread/start`` on the same handshaken
client (initialize now runs once per client, not once per attempt).

Why: the thread id lived in memory only, so every new AIAgent for the same
Hermes session — a later API-server request or the first turn after a
restart — started a fresh codex thread and the model lost its own memory of
the conversation (#100531). The runtime decides the fail-closed policy; this
adapter only speaks the wire contract (verified against codex-cli 0.147.0:
thread/resume{threadId,...} -> result.thread.id; unknown id -> -32600 "no
rollout found"; killed writer -> -32600 "already has an active writer").

Salvaged from #103352 (thread/resume + id cross-fill + mismatch guard).
2026-09-19 12:27:01 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
teknium1
21ef1b97f9 fix(context): proxied Codex routes resolve the Codex OAuth window, not the direct-API catalog
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.

Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.

Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).

Fixes #116191
2026-09-19 12:17:05 -07:00
JoaoMarcos44
193e568494 feat(config): expose a custom route's canonical api_mode for metadata lookups
Custom providers already declare their wire protocol (api_mode / legacy
transport) on the entry, but nothing outside routing could read it: context
length resolution keyed Codex detection on the hostname, which a proxy on
127.0.0.1 never matches. get_custom_provider_api_mode() returns the canonical
transport of the first entry serving a base_url (None custom_providers loads
config, like the sibling context_length helper) so metadata decisions can
follow the route instead of the host.

Slim port of the helper from #116262 by @JoaoMarcos44.
2026-09-19 12:17:05 -07:00
teknium1
fc49f7619d fix(cron): a missing-credential preflight verdict names the profile and HERMES_HOME it read
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.

Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.

Part of #116213
2026-09-19 12:14:32 -07:00
teknium1
03cfc36f51 docs: describe the session.create model×provider refusal
Hosts speaking the gateway protocol need to know the new -32602 shape
(error.data model/provider/suggestions) and which pairs stay permissive.
2026-09-19 12:12:05 -07:00