The lifecycle guard classifies a `.py` script as Python and skips its
shell reference walk, so `interpreter=/bin/bash` on a `.py` whose body is
`bash restart.sh` was created and then executed as shell. Refuse any
interpreter whose name (or symlink target's name) is not a Python image,
reusing the guard's own `_INTERPRETER_IMAGE_RE`.
One parametrized test covers every refusal (bare name, missing,
directory, non-executable, /bin/bash, python -> /bin/bash symlink); the
two run-path tests now use a `python3`-named wrapper.
Add an optional per-job `interpreter` field so a cron Python `script` /
`monitor_script` can run under a user-managed venv instead of Hermes' own
Python, letting scripts import packages the Hermes runtime does not carry
(#8714). Nothing is installed, frozen, or restored automatically.
- cron/jobs.py: persist + normalize the field (absent => record unchanged;
empty string clears it on update).
- cron/scheduler_script.py: _resolve_cron_interpreter() validates the path
at run time (absolute/~ required, regular file, executable on POSIX);
_script_argv runs [interpreter, script] and skips the managed-store
bootstrap/PYTHONPATH overlays, which exist for Hermes' own venv.
Threaded through _run_job_script, the claim-heartbeat wrapper, the
pre-run prompt path and monitor scripts.
- hermes_cli: --interpreter on `cron create` / `cron edit`; shown in
details and `cron list`.
- tools/cronjob_tools.py: programmatic/CLI lane only, like model and
reasoning_effort — absent from the model-facing schema.
Shell scripts (.sh/.bash) still always run under bash. Revives #8741.
Ported onto current main from #70500 (the scheduler moved to
cron/scheduler_script.py and the CLI/tool became table-driven since the
PR's base).
Co-authored-by: MestreY0d4-Uninter <241404605+MestreY0d4-Uninter@users.noreply.github.com>
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.
The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.
Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.
Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
Users who followed the previous documented setup (express key, GEMINI_BASE_URL
unset) now 403 on every surface with only an error-text hint. Call the
behaviour change out in the Gemini guide with the one-line config fix.
Google now issues AQ.-prefixed keys for both Google AI Studio and Vertex AI
express mode (the AIza Studio format is being phased out), so Hermes no longer
reroutes by key shape (#115306, #115323). Rewrite the 'Vertex AI Express Mode
Keys' section to teach the configured-base contract: default Studio host with
GEMINI_BASE_URL unset, explicit aiplatform base for express keys (completed to
the publishers form), and 403-surface guidance.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
`model.auth_mode: entra_id` mints a per-request bearer through azure-identity, so
no `AZURE_FOUNDRY_API_KEY` ever exists; `_overlay_has_env_creds` only counted an
env key as "configured" and hid the azure-foundry row from every picker (CLI, TUI,
Desktop, gateway) and from the prefetch scan. The row now counts as configured when
the runtime resolver's own inputs are present: `model.provider: azure-foundry`,
`model.auth_mode: entra_id`, and an endpoint (`model.base_url` or
`AZURE_FOUNDRY_BASE_URL`). No token is minted for the listing, so azure-identity
is not required to see the row; without an endpoint the row stays hidden, as the
runtime would refuse it anyway.
Reapplies the `_keyless_builtin_configured` hunk of #107991 onto the overlay
credential ladder (azure-foundry is a HERMES_OVERLAYS row, not a canonical one)
and adds the endpoint requirement. Part of #27989.
Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
Two invariant tests (both red on origin/main): `provider_model_ids("azure-foundry")`
probes the configured resource and fails soft to `[]`; a `providers.azure-foundry.models`
block extends the overlay picker row. Guide: the in-session `/model` picker now
lists the resource catalog; how to pin deployments via `providers.azure-foundry.models`.
Two invariant tests (red on base): the task-level `auxiliary.vision.api_mode:
responses` route resolves to `CodexAuxiliaryClient` with the azure-foundry
identity intact, and an explicit `api_mode="responses"` kwarg to
`resolve_provider_client` does the same. The Azure Foundry guide now lists
`responses` as an accepted spelling for model, fallback and auxiliary routes.
Auxiliary tasks on `provider: auto` (title generation, context compression,
smart approval) forward the main runtime's api_key into
`_resolve_azure_foundry_runtime` as `explicit_api_key`. Under
`auth_mode: entra_id` that api_key is the Entra token-provider callable;
the resolver's first line stringified it, so the truthy function repr took
the "explicit string" escape hatch, was relabelled `auth_mode: api_key` and
sent to Azure as a static key -> HTTP 401 on every aux call while the main
conversation worked.
A forwarded value recognised by `is_token_provider()` now stays the runtime
api_key with `auth_mode: entra_id` and the config's Entra metadata. The
explicit STRING escape hatch (`--api-key` while config says entra_id) is
unchanged, and api_key mode remains string-only: a callable there falls
through to the env/.env key as before.
Semantic hunk ported from #108039 (reformat/bloat stripped: unrelated
callable handling in models.py / runtime_provider_custom.py /
tui_gateway.model_switch left out); #72463 by kyssta-exe filed the same
fix first against the pre-split runtime_provider.py.
Fixes#72421
Co-authored-by: kyssta-exe <kyssta-exe@users.noreply.github.com>
The generic API-key connectivity check sent `GET <base>/models` with Bearer
auth to every provider. Azure Foundry's `/anthropic` route has no models
listing, so a fully working Claude deployment showed `Azure Foundry (HTTP
404)`; when the base URL lived only in `model.base_url` (not the env var)
the probe had no URL at all and printed an httpx type error instead.
Now an Anthropic-only base (still `/anthropic` after the dual-surface
rewrite) is probed with a one-token `POST /v1/messages` built from the
adapter's own helpers — `_base_client_kwargs` for the normalized base and
the Azure `api-version` default_query, `_requires_bearer_auth` for the auth
header family — so doctor exercises exactly the request the runtime sends.
200/400 (Messages-API-shaped) = reachable, 401/403 = auth failure, 404 still
reported. OpenAI-style Foundry bases keep the existing `GET /models` probe.
The Azure row also falls back to `model.base_url` when
`AZURE_FOUNDRY_BASE_URL` is unset, matching runtime resolution.
Hand-reapplied from PR #66798 (reformat/bloat stripped: no dedicated
probe function, no models-then-messages double request, no `api-key`
header, one shared status mapping).
Fixes#66756
Salvages #66798
The cherry-picked #114482 never resolved a profile in production: the only
caller, agent/model_metadata.py::_resolve_bedrock_context_length, invokes
get_bedrock_context_length(model, probe=False) with no region, and the
resolver was gated on `region`; and it called
get_inference_profile(inferenceProfileId=...) where botocore requires
`inferenceProfileIdentifier`, so even with a region the call raised
ParamValidationError, was swallowed, and the 128k default applied with only
the new warning. Live against a botocore Stubber: 128000 before, 1000000 after.
- resolve in the ARN's own region (field 4), then the passed region, then
the standard AWS chain; the runtime region / base_url may differ
- inferenceProfileIdentifier is the only request parameter GetInferenceProfile has
- match only `application-inference-profile/`: system-defined
`inference-profile/us.anthropic...` ARNs embed the model id and need no call
- no nested-profile recursion: GetInferenceProfile lists foundation-model ARNs
and the wrapped ARN itself satisfies the static-table substring match
- cache both outcomes per process (runs on every context-length resolution),
cleared by reset_client_cache()
- tests trimmed to two invariants on the production call shape; restore the
TestBedrockContextProbe class header the cherry-pick clobbered
- docs: bedrock:GetInferenceProfile IAM permission and the fallback WARNING
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.
One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.
Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.
Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.
Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).
guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
Follow-up on the picked #113768 commit so it clears the constraint that got
e9a54c48f2 reverted (a9fabe43c4): `/model <direct-alias>` resolves the alias
LABEL first and applied the alias endpoint only afterwards, so any guard that
fires on the label alone turns a working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery went red again
with the PR as-is).
- hermes_cli/runtime_provider.py::_raise_if_local_alias_missing_endpoint keys
on the class (anything auth.resolve_provider maps to `custom` without a rung
of its own; llamacpp keeps its managed-server fail-fast) instead of a second
hardcoded alias set, requires the providers.<alias> block to actually carry a
base_url (an entry without one falls through to OpenRouter too), and names
the alias plus where to set its endpoint. OPENROUTER_BASE_URL never counts as
the alias endpoint; an explicit api_key does not lift the guard.
- hermes_cli/model_switch.py::_creds_for_switched_provider hands a URL-bearing
direct alias's base_url to the resolver as explicit_base_url, so the alias
has an endpoint at the moment the guard runs (the same URL
_apply_direct_alias_endpoint installs later).
- Tests trimmed to two invariants (raise-and-name incl. the OPENROUTER_BASE_URL
/ explicit-api_key non-lifts + bare-custom control; every endpoint source
resolves to its URL). Docs: troubleshooting entry in the local Ollama guide.
Cron replay snapshots that store the resolved `custom` instead of the alias
are #109765's atom and untouched here.
Two invariant tests: the handler returns 0/1/3 for connected / connection
failed / not in config, and the `hermes mcp` CLI dispatcher forwards that code
to `main()` (which already exits on an int return). The MCP guide documents the
codes so probes and watchdogs can stop parsing the output.
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.
The error now names the resolved home's `.env` (and whether the platform's
token key is defined there), `config.yaml` (block absent / `enabled: false` /
no token) and the environment variable(s) checked, so a Windows or profile
home user can fix the file this process actually read. When a gateway started
from the same home already has the platform connected, the message says its
token lives only in that process's environment and which key to add to `.env`.
Docstrings in `gateway/channel_directory.py` and `send_cmd._load_hermes_env`
stop naming `~/.hermes`; the pipe-script-output guide documents the message.
Google issues two Gemini key families: AI Studio keys (AIza...) and Vertex AI
express-mode keys (AQ....). Express keys only authenticate against
aiplatform.googleapis.com; the native adapter hardcoded the Studio host, so an
express key had no working path (403), and an explicit aiplatform base URL was
not even recognised as native Gemini and used the wrong model path.
- normalize_gemini_base_url(base_url, api_key="") routes an AQ. key that would
land on generativelanguage to
https://aiplatform.googleapis.com/v1beta1/publishers/google; an explicit proxy
base is never rewritten. The express base carries the publishers/google
prefix so every {base}/models/{model}:... builder (chat, tier probe, Gemini
TTS) needs no path branching; an explicit aiplatform host root / v1beta1 base
is completed to that form.
- is_native_gemini_base_url accepts the express host but NOT the OAuth Vertex
provider's .../projects/{p}/locations/{r}/endpoints/openapi base, which is
OpenAI-compatible and must stay off the native adapter.
- GeminiNativeClient, probe_gemini_tier and both Gemini TTS call sites pass the
key through.
Cherry-picked from the reporter's earlier PR #96587 and reshaped onto current
main; #114343 and #101918 proposed the same routing.
Fixes#114335
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: cloim <cloimism@gmail.com>
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.
SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).
The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.
Fixes#112570
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.
normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).
Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Direct package installs bypass PM's dependency selection and do not survive
new generations. Document explicit runtime extras, recorded repair, fresh
build outputs, and lock generation through PM instead.
Keep the prepared-interpreter prerequisite and explicit removal requirement
for disposable test environments. Preserve Nix and external-project package
manager ownership. Correct platform and Python-marker claims where the
manifest contradicts the installation hints.
Checked the public CLI help, literal PM calls against public signatures and
extra declarations, fenced blocks, and whitespace. No dependency build or
site build ran. The docs toolchain is not installed in this checkout.
bedrock.guardrail was only attached on the Converse route (guardrailConfig in the
body). Claude on Bedrock goes through the AnthropicBedrock SDK, i.e. InvokeModel,
whose body has no guardrailConfig, so the default Claude route ran with no guardrail
at all (#52179; live-verified by JiaDe-Wu: the blocked word came back through Hermes).
Bedrock reads the guardrail for InvokeModel from X-Amzn-Bedrock-GuardrailIdentifier /
-GuardrailVersion / -Trace headers. Attach them as default_headers in
build_anthropic_bedrock_client so every AnthropicBedrock client Hermes builds
(primary init, /model switch, fallback, per-request rebuild, auxiliary) enforces the
same guardrail, with prompt caching / thinking / 1M context kept (the reason Claude is
not routed through Converse).
InvokeModel blocks do NOT change stop_reason (stays end_turn) and return the guardrail's
canned text as an ordinary assistant reply, flagged only by
amazon-bedrock-guardrailAction=INTERVENED in the body (SDK: response.model_extra).
AnthropicTransport.response_finish_reason maps that to content_filter so the loop runs
its refusal handling instead of reasoning over the canned text; _derive_finish_reason
uses it for the anthropic_messages branch.
Mantle (openai.gpt-5.x) is documented by AWS as not supporting Guardrails on the
Responses endpoint; the docs now say so instead of promising "all model invocations".
Header mechanism proposed in #52312 by @JoaoMarcos44 (stale base, 7-file conflict,
detection keyed on a Converse-only stopReason); reimplemented on current main.
Live probe (local sink, SigV4 fake creds): before, no X-Amzn-Bedrock-* header on the
InvokeModel request; after, headers present, SigV4 intact, INTERVENED → content_filter.
The delegation page and the delegation-patterns guide still stated the pre-v2026.8.31 defaults (50 iterations, 3 concurrent subagents). Shipped values: DEFAULT_MAX_ITERATIONS = 250 (tools/delegate_tool.py) and max_concurrent_children: 10 (hermes_cli/config_defaults.py). The per-task output_schema contract (one bounded correction retry, schema_valid / schema_errors on the result) was not documented anywhere on the page. The Max Iterations section also showed max_iterations as a per-call argument; delegate_task ignores caller-supplied values and reads delegation.max_iterations from config.
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.
Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
Follow-up to the salvaged contributor commit: the Gemini and Vertex guide model
tables now list both Flash generations the pickers offer, and an invariant test
ties the OpenRouter/Nous curated google/gemini-*-flash entries to (a) a Google
official-docs pricing row on the direct Gemini and Vertex routes and (b)
membership in the direct gemini/vertex picker lists. Red on main
(gemini-3.7-flash -> unknown), green with the contributor commit.
Campaign: https://github.com/NousResearch/hermes-agent/issues/104154
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.
- delegate_tool: result.failed now forces status 'failed' (with the
error carried on the entry); new shared format_subagent_failure_line()
renders one clean human-readable line (traceback -> exception message,
length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
terminal failure status deliver that line via _deliver_platform_notice
BEFORE all progress-queue gates; tool_progress_callback is now always
attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.