Commit Graph

198 Commits

Author SHA1 Message Date
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
41a0c3f17c fix(picker): Entra-only Azure Foundry stays listed in /model without an API key
`model.auth_mode: entra_id` mints a per-request bearer through azure-identity, so
no `AZURE_FOUNDRY_API_KEY` ever exists; `_overlay_has_env_creds` only counted an
env key as "configured" and hid the azure-foundry row from every picker (CLI, TUI,
Desktop, gateway) and from the prefetch scan. The row now counts as configured when
the runtime resolver's own inputs are present: `model.provider: azure-foundry`,
`model.auth_mode: entra_id`, and an endpoint (`model.base_url` or
`AZURE_FOUNDRY_BASE_URL`). No token is minted for the listing, so azure-identity
is not required to see the row; without an endpoint the row stays hidden, as the
runtime would refuse it anyway.

Reapplies the `_keyless_builtin_configured` hunk of #107991 onto the overlay
credential ladder (azure-foundry is a HERMES_OVERLAYS row, not a canonical one)
and adds the endpoint requirement. Part of #27989.

Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
2026-09-19 09:42:49 -07:00
teknium1
cfd4921be6 test(models): azure-foundry picker invariants; docs: /model lists Foundry deployments
Two invariant tests (both red on origin/main): `provider_model_ids("azure-foundry")`
probes the configured resource and fails soft to `[]`; a `providers.azure-foundry.models`
block extends the overlay picker row. Guide: the in-session `/model` picker now
lists the resource catalog; how to pin deployments via `providers.azure-foundry.models`.
2026-09-19 09:42:49 -07:00
teknium1
92fa970e20 test(aux): pin responses alias routing for azure-foundry vision; document the alias
Two invariant tests (red on base): the task-level `auxiliary.vision.api_mode:
responses` route resolves to `CodexAuxiliaryClient` with the azure-foundry
identity intact, and an explicit `api_mode="responses"` kwarg to
`resolve_provider_client` does the same. The Azure Foundry guide now lists
`responses` as an accepted spelling for model, fallback and auxiliary routes.
2026-09-19 09:33:37 -07:00
Tudor Pastor
fc9f40ac2c fix: keep the Entra token provider intact when auxiliary tasks re-resolve Azure Foundry
Auxiliary tasks on `provider: auto` (title generation, context compression,
smart approval) forward the main runtime's api_key into
`_resolve_azure_foundry_runtime` as `explicit_api_key`. Under
`auth_mode: entra_id` that api_key is the Entra token-provider callable;
the resolver's first line stringified it, so the truthy function repr took
the "explicit string" escape hatch, was relabelled `auth_mode: api_key` and
sent to Azure as a static key -> HTTP 401 on every aux call while the main
conversation worked.

A forwarded value recognised by `is_token_provider()` now stays the runtime
api_key with `auth_mode: entra_id` and the config's Entra metadata. The
explicit STRING escape hatch (`--api-key` while config says entra_id) is
unchanged, and api_key mode remains string-only: a callable there falls
through to the env/.env key as before.

Semantic hunk ported from #108039 (reformat/bloat stripped: unrelated
callable handling in models.py / runtime_provider_custom.py /
tui_gateway.model_switch left out); #72463 by kyssta-exe filed the same
fix first against the pre-split runtime_provider.py.

Fixes #72421

Co-authored-by: kyssta-exe <kyssta-exe@users.noreply.github.com>
2026-09-19 09:31:28 -07:00
Bartok9
9bd66000a7 fix: hermes doctor probes Azure Foundry Anthropic endpoints like the runtime
The generic API-key connectivity check sent `GET <base>/models` with Bearer
auth to every provider. Azure Foundry's `/anthropic` route has no models
listing, so a fully working Claude deployment showed `Azure Foundry (HTTP
404)`; when the base URL lived only in `model.base_url` (not the env var)
the probe had no URL at all and printed an httpx type error instead.

Now an Anthropic-only base (still `/anthropic` after the dual-surface
rewrite) is probed with a one-token `POST /v1/messages` built from the
adapter's own helpers — `_base_client_kwargs` for the normalized base and
the Azure `api-version` default_query, `_requires_bearer_auth` for the auth
header family — so doctor exercises exactly the request the runtime sends.
200/400 (Messages-API-shaped) = reachable, 401/403 = auth failure, 404 still
reported. OpenAI-style Foundry bases keep the existing `GET /models` probe.
The Azure row also falls back to `model.base_url` when
`AZURE_FOUNDRY_BASE_URL` is unset, matching runtime resolution.

Hand-reapplied from PR #66798 (reformat/bloat stripped: no dedicated
probe function, no models-then-messages double request, no `api-key`
header, one shared status mapping).

Fixes #66756
Salvages #66798
2026-09-19 09:29:18 -07:00
teknium1
11445cc54e fix(bedrock): application inference profile ARNs size from the wrapped model on the real path
The cherry-picked #114482 never resolved a profile in production: the only
caller, agent/model_metadata.py::_resolve_bedrock_context_length, invokes
get_bedrock_context_length(model, probe=False) with no region, and the
resolver was gated on `region`; and it called
get_inference_profile(inferenceProfileId=...) where botocore requires
`inferenceProfileIdentifier`, so even with a region the call raised
ParamValidationError, was swallowed, and the 128k default applied with only
the new warning. Live against a botocore Stubber: 128000 before, 1000000 after.

- resolve in the ARN's own region (field 4), then the passed region, then
  the standard AWS chain; the runtime region / base_url may differ
- inferenceProfileIdentifier is the only request parameter GetInferenceProfile has
- match only `application-inference-profile/`: system-defined
  `inference-profile/us.anthropic...` ARNs embed the model id and need no call
- no nested-profile recursion: GetInferenceProfile lists foundation-model ARNs
  and the wrapped ARN itself satisfies the static-table substring match
- cache both outcomes per process (runs on every context-length resolution),
  cleared by reset_client_cache()
- tests trimmed to two invariants on the production call shape; restore the
  TestBedrockContextProbe class header the cherry-pick clobbered
- docs: bedrock:GetInferenceProfile IAM permission and the fallback WARNING

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 15:15:02 -07:00
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
liuhao1024
a0562d17d8 fix(auth): missing-credential hints name the real env var or the OAuth login (#114405, #78996)
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.

One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.

Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.

Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
2026-09-18 11:11:44 -07:00
teknium1
b9d3c70303 docs(portal): name the exact per-profile sign-in path and fix the guide page too
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.

Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).

guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
2026-09-18 10:28:04 -07:00
teknium1
fea2396981 fix(runtime): local alias guard keys on a missing endpoint, /model passes the alias URL
Follow-up on the picked #113768 commit so it clears the constraint that got
e9a54c48f2 reverted (a9fabe43c4): `/model <direct-alias>` resolves the alias
LABEL first and applied the alias endpoint only afterwards, so any guard that
fires on the label alone turns a working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery went red again
with the PR as-is).

- hermes_cli/runtime_provider.py::_raise_if_local_alias_missing_endpoint keys
  on the class (anything auth.resolve_provider maps to `custom` without a rung
  of its own; llamacpp keeps its managed-server fail-fast) instead of a second
  hardcoded alias set, requires the providers.<alias> block to actually carry a
  base_url (an entry without one falls through to OpenRouter too), and names
  the alias plus where to set its endpoint. OPENROUTER_BASE_URL never counts as
  the alias endpoint; an explicit api_key does not lift the guard.
- hermes_cli/model_switch.py::_creds_for_switched_provider hands a URL-bearing
  direct alias's base_url to the resolver as explicit_base_url, so the alias
  has an endpoint at the moment the guard runs (the same URL
  _apply_direct_alias_endpoint installs later).
- Tests trimmed to two invariants (raise-and-name incl. the OPENROUTER_BASE_URL
  / explicit-api_key non-lifts + bare-custom control; every endpoint source
  resolves to its URL). Docs: troubleshooting entry in the local Ollama guide.

Cron replay snapshots that store the resolved `custom` instead of the alias
are #109765's atom and untouched here.
2026-09-18 10:20:32 -07:00
teknium1
23f4db708f test(mcp): pin hermes mcp test exit codes; document them
Two invariant tests: the handler returns 0/1/3 for connected / connection
failed / not in config, and the `hermes mcp` CLI dispatcher forwards that code
to `main()` (which already exits on an int return). The MCP guide documents the
codes so probes and watchdogs can stop parsing the output.
2026-09-18 10:11:46 -07:00
teknium1
2164221844 fix(send): name the default-root gateway and external secret sources in the 'not configured' error
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.
2026-09-18 10:11:14 -07:00
teknium1
7a99c4d13c fix(send): the 'not configured' error lists the home and sources it consulted
The error now names the resolved home's `.env` (and whether the platform's
token key is defined there), `config.yaml` (block absent / `enabled: false` /
no token) and the environment variable(s) checked, so a Windows or profile
home user can fix the file this process actually read. When a gateway started
from the same home already has the platform connected, the message says its
token lives only in that process's environment and which key to add to `.env`.

Docstrings in `gateway/channel_directory.py` and `send_cmd._load_hermes_env`
stop naming `~/.hermes`; the pipe-script-output guide documents the message.
2026-09-18 10:11:14 -07:00
lepetitprince716-prog
6293fca019 fix(gemini): route Vertex AI express keys (AQ.) to aiplatform instead of 403ing on AI Studio
Google issues two Gemini key families: AI Studio keys (AIza...) and Vertex AI
express-mode keys (AQ....). Express keys only authenticate against
aiplatform.googleapis.com; the native adapter hardcoded the Studio host, so an
express key had no working path (403), and an explicit aiplatform base URL was
not even recognised as native Gemini and used the wrong model path.

- normalize_gemini_base_url(base_url, api_key="") routes an AQ. key that would
  land on generativelanguage to
  https://aiplatform.googleapis.com/v1beta1/publishers/google; an explicit proxy
  base is never rewritten. The express base carries the publishers/google
  prefix so every {base}/models/{model}:... builder (chat, tier probe, Gemini
  TTS) needs no path branching; an explicit aiplatform host root / v1beta1 base
  is completed to that form.
- is_native_gemini_base_url accepts the express host but NOT the OAuth Vertex
  provider's .../projects/{p}/locations/{r}/endpoints/openapi base, which is
  OpenAI-compatible and must stay off the native adapter.
- GeminiNativeClient, probe_gemini_tier and both Gemini TTS call sites pass the
  key through.

Cherry-picked from the reporter's earlier PR #96587 and reshaped onto current
main; #114343 and #101918 proposed the same routing.

Fixes #114335

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: cloim <cloimism@gmail.com>
2026-09-18 09:28:39 -07:00
teknium1
7bb4336811 fix(prompt_builder): load the user's own SOUL.md on a scanner hit instead of blocking it
A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.

SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).

The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.

Fixes #112570
2026-09-16 17:17:47 -07:00
teknium1
123db98635 docs(gemini): scope the base-URL normalization claim to the Google host and TTS
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
2026-09-15 03:54:01 -07:00
Teknium
f030c03970 Port from cline/cline#13329: normalize host-root Gemini base URLs to /v1beta
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.

normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).

Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
2026-09-15 03:54:01 -07:00
Teknium
0c2e66ea8b test(auth): OpenRouter PKCE invariants, fake-authority A/B harness, docs, contributor map
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
2026-09-12 22:07:41 -07:00
Teknium
f8456248d9 fix: Bedrock Guardrails enforced on the Claude route and blocks surface as refusals
bedrock.guardrail was only attached on the Converse route (guardrailConfig in the
body). Claude on Bedrock goes through the AnthropicBedrock SDK, i.e. InvokeModel,
whose body has no guardrailConfig, so the default Claude route ran with no guardrail
at all (#52179; live-verified by JiaDe-Wu: the blocked word came back through Hermes).

Bedrock reads the guardrail for InvokeModel from X-Amzn-Bedrock-GuardrailIdentifier /
-GuardrailVersion / -Trace headers. Attach them as default_headers in
build_anthropic_bedrock_client so every AnthropicBedrock client Hermes builds
(primary init, /model switch, fallback, per-request rebuild, auxiliary) enforces the
same guardrail, with prompt caching / thinking / 1M context kept (the reason Claude is
not routed through Converse).

InvokeModel blocks do NOT change stop_reason (stays end_turn) and return the guardrail's
canned text as an ordinary assistant reply, flagged only by
amazon-bedrock-guardrailAction=INTERVENED in the body (SDK: response.model_extra).
AnthropicTransport.response_finish_reason maps that to content_filter so the loop runs
its refusal handling instead of reasoning over the canned text; _derive_finish_reason
uses it for the anthropic_messages branch.

Mantle (openai.gpt-5.x) is documented by AWS as not supporting Guardrails on the
Responses endpoint; the docs now say so instead of promising "all model invocations".

Header mechanism proposed in #52312 by @JoaoMarcos44 (stale base, 7-file conflict,
detection keyed on a Converse-only stopReason); reimplemented on current main.

Live probe (local sink, SigV4 fake creds): before, no X-Amzn-Bedrock-* header on the
InvokeModel request; after, headers present, SigV4 intact, INTERVENED → content_filter.
2026-09-11 01:37:22 -07:00
witcheer
bc1d1a2879 docs(delegation): update shipped defaults (250 iterations, 10 concurrent children) and document output_schema
The delegation page and the delegation-patterns guide still stated the pre-v2026.8.31 defaults (50 iterations, 3 concurrent subagents). Shipped values: DEFAULT_MAX_ITERATIONS = 250 (tools/delegate_tool.py) and max_concurrent_children: 10 (hermes_cli/config_defaults.py). The per-task output_schema contract (one bounded correction retry, schema_valid / schema_errors on the result) was not documented anywhere on the page. The Max Iterations section also showed max_iterations as a per-call argument; delegate_task ignores caller-supplied values and reads delegation.max_iterations from config.
2026-09-08 18:58:00 +05:30
Teknium
bbcf1ee180 fix: preserve native Gemini union constraints
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.

Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
2026-09-07 21:10:50 -07:00
Teknium
34e512ae58 fix(sessions): remove remaining timer migration and guidance 2026-09-07 06:10:54 -07:00
Teknium
1481a0de96 docs(gateway): describe persistent sessions without reset timers 2026-09-07 06:10:54 -07:00
Teknium
f914c9b070 fix(mcp): enforce profile ownership throughout OAuth sessions 2026-09-07 06:10:28 -07:00
Teknium
b706529476 docs+test(models): Gemini 3.7/3.8 Flash in the Gemini and Vertex guides; curated Google Flash pickers must bill via the official snapshot
Follow-up to the salvaged contributor commit: the Gemini and Vertex guide model
tables now list both Flash generations the pickers offer, and an invariant test
ties the OpenRouter/Nous curated google/gemini-*-flash entries to (a) a Google
official-docs pricing row on the direct Gemini and Vertex routes and (b)
membership in the direct gemini/vertex picker lists. Red on main
(gemini-3.7-flash -> unknown), green with the contributor commit.

Campaign: https://github.com/NousResearch/hermes-agent/issues/104154
2026-09-06 05:42:22 -07:00
Victor Kyriazakos
c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
emozilla
43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium
5a134383fe fix: failed subagents now surface a clean error to the user (CLI + gateway)
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.

- delegate_tool: result.failed now forces status 'failed' (with the
  error carried on the entry); new shared format_subagent_failure_line()
  renders one clean human-readable line (traceback -> exception message,
  length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
  terminal failure status deliver that line via _deliver_platform_notice
  BEFORE all progress-queue gates; tool_progress_callback is now always
  attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
2026-08-30 20:40:14 -07:00
Teknium
99c3cad857 docs: relay-fronted cron delivery + manual-run gateway forward (troubleshooting) 2026-08-27 20:53:02 -07:00
Teknium
6a6e16fa5d feat(desktop): OS-keychain encryption for stored secrets is now opt-in — no more macOS Keychain password prompt on every launch
Electron safeStorage parks a per-app key ('Hermes Key') in the macOS login
keychain; on machines with a locked/missing/corrupted default keychain that
turned every Hermes Desktop launch into a blocking 'Keychain Not Found' /
password dialog. Keychain-backed encryption is now an explicit opt-in:

- electron/secret-storage-policy.ts: standalone policy seam (default OFF,
  strict === true coercion, one-shot migration flag) + unit tests
- default path never calls any safeStorage API (including
  isEncryptionAvailable, which itself touches the keychain)
- one-shot legacy migration decrypts existing safeStorage blobs to plain
  0600 files at first launch; undecryptable blobs are kept but read as
  absent afterward (classify 'drop') so a dead keychain prompts at most once
- Settings -> Gateway toggle (all 5 locales) re-encodes every stored secret
  store in place when flipped (v1 connection.json, v2 connections.json,
  native-oauth-tokens.json)
- e2e: at-rest spec now covers both postures (opted-in unchanged contract,
  default saves without secure storage, owner-only bits, restart round-trip)
- docs: multi-connection-desktop + desktop-native-signin updated
2026-08-25 14:10:19 -07:00
Mark McDonald
eaf6545ab4 docs(gemini): update to use latest gemini models 2026-08-25 13:14:30 +05:30
Teknium
d9a48f656a fix(desktop): scheduled jobs on sleeping profiles keep firing
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
2026-08-24 03:14:30 -07:00
Vinay Shah
41ca67c5b1 fix(bedrock): align auxiliary region resolution with runtime + document Mantle route
Address review feedback on #65076:

- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
  config-first region resolution (bedrock.region in config.yaml, then
  AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
  runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
  branch) to the new helper. Previously it derived its region with bare
  resolve_bedrock_region() (env-first), so when config.yaml pinned
  bedrock.region to a different region than the ambient AWS env, auxiliary
  calls (compression, memory, vision) left the primary runtime's region.
  Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
  Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
  for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
  uses the OpenAI-compatible endpoint, which the Mantle route made stale.
  Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
  Converse), the Mantle auth model (bearer token or SigV4), and add the
  GPT-5.5/5.6 model IDs to the models table.
2026-08-21 15:02:29 -07:00
Teknium
a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium
ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Ben Barclay
46e1a59d84 Merge pull request #69829 from NousResearch/docs/hermes-cloud-mcp
docs: guide for managing Hermes Cloud via the Portal MCP server
2026-08-21 14:51:00 +10:00
Teknium
ac0a8cd281 feat(image-gen): xAI Grok image catalog goes live-driven; grok-imagine-image-2.0 selectable
- plugins/image_gen/xai: merge the live /v1/image-generation-models catalog
  (5-min cache, 10s timeout, static-table fallback when offline/unauth)
  into the picker so new xAI Imagine models appear automatically the day
  they launch, with generic metadata until curated text is added.
- Add grok-imagine-image-2.0 to the curated static table (typography/
  layout-aware model, API-available since Aug 8 2026).
- Edits honor an explicitly selected image-input-capable model
  (e.g. grok-imagine-image-2.0) instead of always forcing
  grok-imagine-image-quality; quality remains the default edit baseline.
- Tests: hermetic autouse fixture keeps unit runs offline; new coverage
  for live-merge, unknown-future-model selection, offline fallback, and
  edit-model resolution. Docs model table updated (en + zh-Hans).

Live-verified: /image-generation-models returns grok-imagine-image,
grok-imagine-image-2.0, grok-imagine-image-quality; real generation with
2.0 succeeded end to end.
2026-08-19 01:19:37 -07:00
Teknium
c820a5d383 docs(teams-pipeline): document fetch --organizer-user-id
Follow-up to #89382: the operator runbook, bundled-skill docs page, and the
bundled SKILL.md now cover the organizer-scoped lookup flag and note that
/meet/ short URLs require it while webhook jobs derive the organizer
automatically.
2026-08-18 13:03:16 -07:00
Buff Pesos
56f1afc834 feat(dashboard-auth): extend RFC 8252 native sign-in to password providers (system-browser autofill) (#75808)
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers

The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.

The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.

Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:

* /auth/native/authorize now accepts a supports_password provider:
  register the pending broker authorization as usual, then 302 the
  system browser to the interactive /login form with the opaque
  broker_state in the gateway's PKCE cookie (the same server-controlled
  channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
  broker handle, a successful credential check completes the pending
  authorization exactly like the /auth/callback native branch — mint
  the one-time loopback code, return the loopback redirect (validated
  loopback-only at authorize time) as `next`, clear the PKCE cookie,
  and set NO session cookies. A lapsed broker is a clean 400 telling
  the user to restart sign-in; a failed credential attempt leaves the
  pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
  session provider is registered (previously only for non-password
  providers), so the desktop selects the system-browser strategy for
  password-only gateways.

Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.

Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dashboard-auth): bind native password completion to the authorize-time provider

Review follow-ups for #75808:

* /auth/password-login now enforces that body.provider matches the
  provider recorded in the server-set PKCE cookie by
  /auth/native/authorize before completing a pending native
  authorization. /login renders a form for every session provider, so
  without this a native flow started for provider A could be completed
  with provider B's credentials, binding B's session into A's pending
  entry. The mismatch is rejected BEFORE credential verification (no
  session minted, no oracle) and preserves both the pending entry and
  the cookie, so the user can still submit the correct provider's form.
  Covered by a two-password-provider E2E regression test.

* Update the two docs spots that still said password-only providers do
  not advertise native_pkce (website desktop-native-signin guide and the
  auth_flows type comment in web/src/lib/api.ts).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: map contributor email for #75808 (buffpesos)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
2026-08-14 21:39:01 +00:00
Jaaneek
9c15f0191c fix(models): refresh xAI picker via models.dev; pin grok-4.6
Stop freezing the xAI/xAI-OAuth catalog at import so /model and setup
pick up new Grok IDs after the models.dev cache refreshes. Put xai and
xai-oauth on the shared picker-time models.dev merge path and pin
grok-4.6 as the default headline model.
2026-08-13 22:09:50 -07:00
Teknium
e4aeb65599 feat(webhook): per-route toolset overrides for webhook agent runs
Webhook agent runs default to the constrained hermes-webhook toolset
(web/vision/clarify) because payloads can carry untrusted third-party
content. That default is right for public webhooks but wrong for trusted
local pushes (e.g. an OOM monitor daemon that needs the agent to run
ps/free/py-spy): the only workaround was widening platform_toolsets.webhook,
which elevates EVERY webhook route at once.

This adds a 'toolsets' key on individual webhook route configs (static
routes in config.yaml and dynamic subscriptions in
webhook_subscriptions.json) that replaces the platform-level resolution
for that route only:

- BasePlatformAdapter.toolsets_for_source(): per-source override hook,
  default None (no behavior change for any other platform).
- WebhookAdapter.toolsets_for_source(): maps the session chat_id
  (webhook:{route}:{delivery_id}) back to its route config and returns
  the route's toolsets list.
- GatewayRunner._resolve_enabled_toolsets_for_source(): shared resolver
  used by both agent-run call sites; validates the override through the
  SAME _get_platform_tools path as platform config, so unknown names and
  platform-restricted toolsets (e.g. discord_admin) are dropped rather
  than trusted.

Deliberately NOT exposed via 'hermes webhook subscribe': granting elevated
tools is a manual config edit only, so an agent-created subscription
cannot self-grant terminal at runtime.

Cache-safe: the toolset list is resolved before agent construction and is
constant for a route, so the per-session agent signature and frozen system
prompt are unaffected mid-conversation.
2026-08-13 01:51:19 -07:00
khanhngoo
f1c45f5727 feat(voice): add configurable TUI draft submission
Add voice.submit_mode=direct|draft without model-refine hooks or callbacks. Validate the config, preserve direct-submit compatibility, render editable drafts in the Ink composer, and document both locales.

Co-authored-by: BELIVIN MEDIA <212580280+KarateWilly@users.noreply.github.com>
2026-08-12 16:42:07 -07:00
mzkarami
9eec86923c docs: align Ollama tool-calling guidance 2026-08-08 20:21:30 -07:00
Teknium
5396da844a docs: DX sweep — 7 verified-absent documentation items
- developer-guide/codebase-ownership.md (new): subsystem -> source dirs ->
  docs entry point map; complements the narrow CODEOWNERS proposal in #23751
  (docs table only, no .github/CODEOWNERS).
- contributing.md: document the .agents/checks/*.md repo-local review
  checklist convention (idea from goose, Apache-2.0).
- integrations/index.md: "Quick connect links" table with prefilled
  create-your-app deep links (Telegram BotFather, Discord
  ?new_application=true, Slack ?new_app=1, LINE, Feishu). Poke-inspired.
- guides/agent-email-address.md (new): dedicated agent mailbox via the
  bundled himalaya skill — setup, cron polling pattern, prompt-injection
  safety notes. Poke-inspired.
- user-guide/features/browser.md: Chrome 136+ silently refuses
  --remote-debugging-port on the default user-data-dir; dedicated profile is
  now mandatory (diagnosis from oh-my-pi, MIT).
- developer-guide/adding-providers.md: "Tool-call wire format" section
  linking the OpenAI chat-completions reference as the canonical shape for
  convert_messages/convert_tools.
- user-guide/features/tools.md: shell-init pitfall — heavy/interactive rc
  files (nvm, TTY-expecting blocks) break non-interactive agent terminal
  calls; interactive-guard pattern documented (from cline, Apache-2.0).

Both new pages registered in sidebars.ts. Validated with npx docusaurus
build (en + zh-Hans green; zh-Hans relative-link warnings are the known
pre-existing untranslated-page noise).
2026-08-07 08:58:08 -07:00
witcheer
7ce6f97945 docs: explain the slow silent first turn (prefill) on local hardware 2026-08-05 21:33:44 +05:30
witcheer
a5ab9b2e5a docs: add troubleshooting checklist for perceived agent-quality regressions 2026-08-05 21:33:44 +05:30
witcheer
8618eba7c8 docs: add security-posture guide for running Hermes on a personal or work machine 2026-08-05 21:33:44 +05:30