Commit Graph

796 Commits

Author SHA1 Message Date
teknium1
86a599cbae fix(gateway): human-delay pacing comes from each profile's human_delay config, not process env
BasePlatformAdapter._get_human_delay read HERMES_HUMAN_DELAY_MODE/_MIN_MS/_MAX_MS from the
process environment at every send, so under multiplexing the launch profile's pacing applied
to every served profile, and the documented `human_delay:` config section (mode/min_ms/max_ms,
already in DEFAULT_CONFIG) was never consulted. The runner now resolves `human_delay` per
profile through the same seam as the busy-text timings (`_human_delay_from_config`,
snapshotted in `_snapshot_profile_busy_modes`, installed by `_wire_adapter_handlers`) and the
adapter only consumes the installed range. Invalid `custom` bounds (non-integer, negative,
inverted) warn naming the key and fall back to the natural range.

Fixes #116895
2026-09-20 17:01:05 -07:00
teknium1
522e121e90 fix(desktop): pin the update-check proxy deps exactly and document the proxy variables
The salvaged commit added https-proxy-agent, proxy-from-env and
@types/proxy-from-env as caret ranges; the repo pins every dependency to an
exact version so `npm ci` resolves the same tree everywhere. Pins are the
versions the lockfile already resolved (7.0.6 / 2.1.0 / 1.0.4), regenerated
with `npm install --package-lock-only`.

Documents that the Desktop update check now follows HTTPS_PROXY / HTTP_PROXY /
NO_PROXY in the environment-variables reference.
2026-09-20 15:08:40 -07:00
teknium1
dbcbd9d9db feat(gateway): /branch opens a sibling thread by default; --here keeps this chat (#66023)
On Discord, Telegram, Slack and Matrix a plain `/branch` used to rebind the
CURRENT chat/thread's session key to the clone, ending the original session
on that surface. The user could not keep the original path live while
exploring an alternate one — the opposite of what a branch is for.

Now the handler opens a sibling thread through the adapter's existing
`create_handoff_thread` BEFORE cloning (a failed create never orphans a
branch row), binds the thread's own session key to the clone with the
thread's routing columns written at create time, and leaves the origin key
untouched. `/branch --here` keeps the legacy in-place switch; platforms
without threads, DMs, unknown Discord parents and adapters that cannot open
a thread fall back to in-place with a one-line note. The CLI strips the
flag through the same parser so `--here` never becomes a session title.

Destination source shapes mirror each adapter's inbound key (Discord keys
threads on their own id; Telegram/Slack/Matrix on the parent chat), the
same rules the CLI->platform handoff uses.

Live repro (real gateway + real Slack adapter against a stand-in Slack
Socket Mode/Web API): base ends the origin session and rebinds its key;
fixed posts the thread seed, replies "this chat stays on it", the origin
thread keeps its session and the follow-up typed in the new thread lands on
the branch (parent_session_id = origin).

Design and first implementation by Angello Picasso (#66014, #66024);
this is a slim port onto the split slash_commands_* layout.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-20 13:37:10 -07:00
treatux
6ae4cb88d5 fix(docs): reference pages name tools the registry never registered
The shipped tool-surface references still document the pre-consolidation
surface: tools-reference.md lists cronjob/todo/process/project_create/
project_list/project_switch/open_preview/close_preview/read_preview/tour/tip
(6 uncallable, 5 hidden dispatch-only aliases), and toolsets-reference.md
still claims web_search is a member of the browser toolset — membership
decacbac3 deliberately removed (#64503) with a regression test. Both rename
commits (e16ad33a9, 217ab2f8d) left website/ untouched.

Pin the pages to the live registry with a contract test (real
discover_builtin_tools()/resolve_toolset queries over the shipped .md data);
rename the rows to the registered surface (cronjob_manage, todo_list,
process_manage, desktop_project enum, desktop_preview, gui_tour, show_tip)
and add the browser row's actual members (browser_vault_*, browser_exec,
apply_layout) that the docs never mentioned.
2026-09-20 12:56:25 -07:00
treatux
2a1ec15a6f docs(website): regenerate the skill docs from the shipped skill tree
website/scripts/generate-skill-docs.py is the documented source of both catalogs
and of the per-skill pages, but nothing compares its output with what is
committed, so the committed copies drifted:

- optional-skills-catalog.md was missing agent-merge-conflict-arbiter and listed
  pr-lens under blockchain (it lives in software-development);
- skills-catalog.md still listed merge-reconciler, which #98539 moved out of the
  bundled set;
- 196 pages and both catalogs carried Windows path separators — in the Path
  column and inside GitHub blob links, where a backslash is a broken URL.

This commit is the generator's output (re-running it on this branch is a no-op),
plus the orphan page #98539 left behind for merge-reconciler: no skill backs it,
its catalog row is gone, and no page links to it.

tests/skills/test_skill_docs_contract.py is the guard: the next skill that ships
without a catalog row, or a page regenerated on Windows, fails there instead of
on the published page.
2026-09-20 12:54:07 -07:00
chelsealong
e23a033e85 docs(matrix): surface the DM-classification bypass in the operator-facing docs
The prior commit only documented the <=2-member DM auto-classification
in the adapter.py module docstring — source an operator configuring
MATRIX_REQUIRE_MENTION etc. would never read. Issue #114733 explicitly
asks for the callout wherever MATRIX_ALLOWED_ROOMS,
MATRIX_FREE_RESPONSE_ROOMS, MATRIX_REQUIRE_MENTION, and MATRIX_AUTO_THREAD
are documented for operators, i.e. the env-var reference table and the
Matrix user guide. Add the same rule + bypassed vars + escape hatch to
both, matching the precedent set by 08aa67e473 for this kind of env-var
clarification, and extend the doc-content test to cover them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 10:42:25 -07:00
teknium1
d5c2e0fb16 fix(env): a parent-injected dashboard session token survives the .env reload
A launcher that spawns `hermes dashboard` (Desktop shell, link-style
integrations) mints HERMES_DASHBOARD_SESSION_TOKEN into the child's
environment and keeps the same token for its own /api probes. Every
dotenv layer in load_hermes_dotenv() loads with override=True, so a
persisted HERMES_DASHBOARD_SESSION_TOKEN line in ~/.hermes/.env replaced
the injected token: the child authenticated with the persisted value and
the parent got HTTP 401 from its own child.

Treat the token as a spawn credential in the dotenv publisher: a value
that dotenv did not put into os.environ (tracked by _DOTENV_PUBLISHED)
is left alone, while a value an earlier pass published still reloads,
so .env edits and home switches behave as before. Other keys, including
the documented HERMES_DASHBOARD_PUBLIC_URL, keep .env-wins precedence.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-20 10:39:01 -07:00
teknium1
c7e165a068 docs: list HERMES_LANGFUSE_MAX_DEPTH in the environment-variable reference
The knob is user-facing; the README alone is not where operators look.
2026-09-19 23:56:34 -07:00
teknium1
05c4c40219 docs(mcp): oauth.user_agent default, redacted token-error excerpt, login keeps server metadata 2026-09-19 23:46:42 -07:00
teknium1
0818892db3 fix(mcp): accept an origin-issued metadata document for a path-scoped OAuth authorization server
A protected resource may advertise a path-scoped authorization server
(`https://www.strava.com/mcp-issuer`) whose RFC 8414 document, served from
`/.well-known/oauth-authorization-server/mcp-issuer`, declares the origin
(`https://www.strava.com`) as its issuer. The SDK's exact-string check
(`validate_metadata_issuer`, RFC 8414 §3.3) rejected that document with
"Authorization server metadata issuer mismatch" and the connection parked
before registration or login (#116233).

`metadata_issued_by_origin` accepts exactly that shape and nothing else: the
document must have been read from the well-known URL derived from the
advertised identifier (so a redirect target, the root document or an OIDC
fallback never qualify) and its issuer must be the advertised identifier's
origin. Only the origin's operator controls that location, so a party
controlling a path or a sibling host cannot use it to make the client accept
another server's endpoints; whoever could would already control the
exact-match document too. The browser flow applies it in the mixin's request pump
(installs the document, hands the SDK a 204 so its loop stops) without touching
`auth_server_url`, so SEP-2352 credential binding keeps the advertised
identifier while the RFC 9207 `iss` check and refresh-token issuer binding use
the document's issuer. The device flow applies the same rule in its own
discovery and now binds registered credentials the way the SDK's Step 4 does,
so the runtime flow reuses them instead of discarding them on the next 401.

Supersedes #116359: its rule accepted the origin issuer from any discovery URL
(root and OIDC fallbacks, redirect targets) and fabricated a 500 response.

Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
2026-09-19 23:40:35 -07:00
teknium1
30de041b01 docs: explain the three gateway connection-failure replies
The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
2026-09-19 18:24:11 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
54c01bc19a chore: merge origin/main (resolve hermes_cli/runtime_provider.py, website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:51:04 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
b682a98ab8 test(codex): pin proxy override on rotation and model.base_url; document HERMES_CODEX_BASE_URL
Two invariant tests (red on origin/main): a 401 rotation onto a Codex pool
row keeps the HERMES_CODEX_BASE_URL target, and model.base_url under
model.provider: openai-codex resolves for pool credentials. Adds the
previously undocumented HERMES_CODEX_BASE_URL row to the environment
variables reference so proxy users can find the knob and its reach.
2026-09-19 10:33:08 -07:00
teknium1
85b2a3df6c feat(cli): hermes usage [--json] prints the /usage account limits without a session
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.

Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).

Co-authored-by: himanusia <himanusia@users.noreply.github.com>
2026-09-19 10:31:14 -07:00
teknium1
03d40f8939 fix(gateway): a --global /model keeps the session override under a channel_overrides model; report a failed stale-override cleanup
Precedence is session /model > channel_overrides > config.yaml, so dropping the session
override after a --global write regressed chats whose channel_overrides names a model: the
confirmation said "switched to gpt-5.5" while the next turn resolved 'channel-model'. The
cleanup now runs only when no channel_overrides entry applies to the source; otherwise the
override stays (and is written through) so the confirmation stays true.

A failed set_model_override(key, None) was logger.debug only while the reply claimed a clean
save and memory had already popped the override — the durable stale copy would shadow
config.yaml on the next restart (the original #100314 symptom). The cleanup failure is now
returned like the config-write error: the in-memory override is kept, the confirmation carries
a warning line instead of "Saved to config.yaml", and a riding --reasoning stays session-scoped.

Tests (tests/gateway/test_model_picker_persist.py, both red on the previous head): the
channel_overrides case drives _handle_model_command with a real JSONL SessionStore and asserts
_resolve_session_agent_runtime(source) yields gpt-5.5; the cleanup-failure case asserts the
warning and the kept in-memory override. Docs note the channel exception and that the CLI/TUI
keep their per-session pin by design (resume restores the model that chat used).
2026-09-19 09:59:32 -07:00
teknium1
1a59519244 fix(gateway): a --global /model pick leaves config.yaml as the only durable model authority
A `/model <m> --global` switch (typed or picker) used to persist twice: the profile
config.yaml AND a per-session `model_override` in the session store. The session copy
has higher precedence on rehydration, so after a later global change (CLI `hermes model`,
another chat's `--global`) and a gateway restart the stale override silently won —
#100314 saw an explicit `gpt-5.6-sol-900k` resume as the base 272K `gpt-5.6-sol`.

Now `_record_model_switch` writes config.yaml FIRST and, on success, drops the redundant
session override from memory and the store. If the config write fails the switch stays a
truthful session override and the confirmation says "config.yaml not updated (...)" plus
the session-only hint instead of claiming "Saved to config.yaml"; a riding `--reasoning`
pin follows the same effective scope. `--session` and `--once` semantics are unchanged.

Tests: two invariants in tests/gateway/test_model_picker_persist.py (typed+picker clear
the durable override and a fresh SessionStore rehydrates nothing; failed config write keeps
the override and an honest reply), red on base. test_model_command_request_overrides now
points get_hermes_home at its own config so the --provider switch resolves session-scoped
as intended instead of the sandbox's fresh-install first-pick rule.

Fixes #100314
Supersedes #99825 (slim redo; the original wrapped the commit boundary through a
sys.modules-swapped mixin).

Co-authored-by: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-19 09:59:32 -07:00
teknium1
393ffff04f fix(providers): resolve the chatgpt alias in the /model parser and hermes auth login too
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.

Part of #95794
2026-09-19 09:38:01 -07:00
Teknium
99a6ed741a Merge pull request #115757 from NousResearch/fix/boa-codex-app-server-lifecycle-migration-mcp
fix(codex-runtime): same-name MCP servers no longer make ~/.codex/config.toml unloadable; adds `hermes codex-runtime migrate` (#79023, salvage #79182)
2026-09-19 09:35:41 -07:00
kshitijk4poor
1536e1d242 docs(discord): finish the free-response auto-thread surfaces
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
2026-09-19 21:43:59 +05:30
Xunjin ZHENG
774e070731 feat(discord): add free_response_auto_thread opt-in
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.

Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.

Default behavior is unchanged; all 291 existing discord tests pass.
2026-09-19 21:43:59 +05:30
Siddharth Balyan
8d4abc3ea6 fix(mcp): retire the n8n bridge catalog entry (#116048)
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.

Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
2026-09-19 17:51:30 +05:30
Kevin
8b7caf226f feat(codex): named custom providers work with the codex_app_server runtime
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.

- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
  admits provider="custom" only when `codex_model_provider_id()` finds a
  configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
  and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
  for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
  `.modelProvider` (fields verified against the codex 0.147 app-server
  schema). `_ensure_codex_session` fills them only for custom agents; codex
  resolves base_url/env_key from its own `[model_providers.<name>]`, so the
  Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
  Desktop/TUI agent knows the provider id (it otherwise collapses to
  "custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
  bare-`custom` ineligibility.

Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).

Fixes #75186
2026-09-19 01:04:05 -07:00
teknium1
d09183d8d3 feat(cli): add hermes codex-runtime migrate [--dry-run] [--json]
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.

Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
2026-09-18 23:38:13 -07:00
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
5e95050608 fix(agent): a server context rejection the transcript cannot explain is no longer "conversation too long"
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).

_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.

Fixes #114644
2026-09-18 19:18:57 -07:00
kshitijk4poor
1250a3e4eb fix(backup): an incomplete archive never prunes the last complete ones
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
2026-09-19 03:43:58 +05:30
leomcamilo
ccb3d968ce fix(backup): exit non-zero when a full backup is incomplete
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.

`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.

Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
2026-09-19 03:43:58 +05:30
kshitijk4poor
8925c70a1c docs(whatsapp): group access section says what the gateway admits; env reference rows; trim bridge tests
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
2026-09-19 03:15:40 +05:30
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
b7b203cda0 fix(skills): built-in name collisions show a note in /skills, /help skills and the palette
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).

Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
  "slash command /<name> unavailable — name taken by built-in; use /skill <name>"

- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
  `hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
  skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
  carries the notice in its existing `warning` field (rendered by the Desktop
  `/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
  listing (Telegram/Discord/…) appends one ⚠ line per colliding skill

Fixes #113560
2026-09-18 12:51:22 -07:00
teknium1
250e12e760 fix(config): provider switch via config set drops the previous provider's route (#113719)
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.

Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):

- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
  a base_url is: the target's registry/plugin host, a `providers:` /
  `custom_providers:` entry resolving to the target, or bare custom/local
  aliases (configured BY base_url) -> owned; another known provider's host
  or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
  alone, with no base_url, is old-route wire state and goes too); keeps an
  owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
  prints what was cleared and why, or warns that an unrecognised base_url
  still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
  (openai-codex + chatgpt.com), a custom entry with that URL, and bare
  `custom` are untouched.

Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
2026-09-18 12:49:40 -07:00
teknium1
668e505ac7 fix(mcp): OAuth discovery/registration carry a User-Agent; cancelled login frees its callback port
Widen the salvaged fixes to the whole class and add the pieces they missed:

- The default User-Agent for SDK-built OAuth requests moves from the manager's
  bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
  provider gets it too, and the manager's pre-flight metadata discovery (its own
  client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
  registration and token requests through client.send(), which never merges the
  client's default headers; www.tradingview.com's WAF answers a header-less GET
  with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
  /register and /authorize on the MCP host, and the login died with
  "Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
  ... != None" because no issuer was ever discovered).

- The callback listener now runs serve_forever() and is shut down before
  server_close(). A thread parked in handle_request()'s select() keeps the
  closed listening socket alive (the kernel holds the file for the duration
  of the poll), so a flow cancelled mid-wait left the port bound and the
  retry on the same pinned/cached port raised "OAuth callback port N is
  already in use" with no external collider.

- When every authorization-server metadata fetch failed, a registration error
  is re-raised leading with those statuses ("Could not read
  authorization-server metadata (403 from ...); dynamic client registration
  then fell back to a guessed endpoint on the MCP host and failed: ...").
  humanize_oauth_registration_error leaves that message alone so the 403 in
  it is not mistaken for a DCR allowlist refusal.

Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
2026-09-18 12:47:49 -07:00
teknium1
76c640ff11 fix(web): list, activate and delete legacy custom_providers entries on Custom Endpoints
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.

The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
2026-09-18 12:44:32 -07:00
teknium1
3e182f46c3 fix(doctor): flag a non-list custom_providers and legacy list entries with no providers: twin; name the edited profile on Custom Endpoints / Local Models
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.

Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).

Part of #114471 (items 2, 5, 6).
2026-09-18 12:44:32 -07:00
teknium1
8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00
teknium1
09cf1b926f fix(state): reset forks leave the Python lineage walk too; /resume ranks lineages by activity
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.

Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.

Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.

Part of #114271
2026-09-18 10:39:13 -07:00
teknium1
4590ef8b58 fix(profiles): polled profile lists never walk skill trees; vanished skill dirs no longer abort enumeration
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.

- list_profiles(lazy_skill_count=True): skill_count is the last known value;
  a missing/aged entry schedules ONE background _count_skills per profile per
  60 s recheck window, so the request thread does zero skill-tree I/O and the
  refresh cadence is decoupled from the poll rate. The two polled callers and
  the REST fallback entry use it; the sync default (CLI, detail views) is
  unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
  os.walk with excluded/support dirs pruned) instead of rglob — half the
  fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
  uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
  projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
  skill add/remove immediately on the next refresh).

Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
  before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
          profiles.list 2189 + 2015, projects/tree 2177 + 2015;
          a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
  after:  GET /api/profiles 0 skill-tree stats/scandir on the request thread
          (counts land from the background refresh by the next poll),
          profiles.list 0, projects/tree 0; _count_skills (detail/control)
          still reports 200 with 203 scandir; the vanished-dir case returns
          199 and list_profiles() enumerates all 5 profiles.

Fixes #114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:18:40 -07:00
teknium1
72360ae1d2 fix(doctor): detect a dead IPv6 route and name network.force_ipv4
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.

Docs: doctor reference and the network config section describe the check.
2026-09-18 09:56:28 -07:00
teknium1
fc94ff56e1 fix(cli): expand HERMES_HOME once at process entry; sweep raw readers
Follow-up to the salvaged hermes_constants expansion (#109212): the CLI has
~30 raw `os.environ["HERMES_HOME"]` readers (fast --version path, early
display.interface probe, profile re-home, dotenv loader, subprocess-home
helper) that never go through get_hermes_home(). A literal `~` (fish, or
any quoted value) left them resolving `~/.hermes` against cwd while the
resolver now expands it, so the process would disagree with itself.

Normalize the env var once, at the earliest point of hermes_cli/main.py
(stdlib-only, before the fast paths), and expand it in the one raw
hermes_constants reader (`_profile_home_path`). A relative value that is
not tilde/variable-shaped is deliberately left alone: rejecting it would
break `HERMES_HOME=./tmp-home` in tests and CI for no user-facing gain.

Tests: real-CLI subprocess with HERMES_HOME='~/.x' under a fake HOME asserts
`config path` lands in the fake home and that no literal `~` directory
appears under cwd (the reporter's acceptance criterion, #114353), plus a
unit test for the normalizer. Docs: environment-variables reference.
2026-09-18 09:49:46 -07:00
teknium1
b65ec6a3f0 docs: list the DISCORD_MISSED_MESSAGE_BACKFILL_* env fallbacks
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
2026-09-18 09:44:59 -07:00
teknium1
9c088a7a78 fix(config): fix container type for unseeded roots, top-level lists and bare-name list slots
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.

agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
2026-09-18 09:41:02 -07:00
teknium1
0972b4819a fix(config): config set refuses a wrong-shaped value for a list/mapping key
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.

Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.

Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.

Config-set atom of #114471.
2026-09-18 09:41:02 -07:00
teknium1
46500f9c79 fix(dashboard): --stop and the update sweep resolve backend ownership tri-state, never guess
Composes the two contributor halves into the scoped root set the issue asks for:
ownership-scoped roots (#113982, @KoNit-K) ∩ caller-chain-spared roots (#113994, @kokhlo).

What this commit adds on top of the picks:
- `_hermes_home_for_pid` is tri-state. `None` now means ONLY "environment unreadable"
  (spared). A readable environment always resolves to a home: exec-time `HERMES_HOME`,
  else the process's own platform default home (`HOME/.hermes`, `LOCALAPPDATA/hermes`),
  with a `--profile X`/`-p X` argv flag selecting `<default>/profiles/X` because
  `_apply_profile_override` writes HERMES_HOME into os.environ AFTER exec and
  `/proc/<pid>/environ` never reflects it. Without this, the common default-home install
  (nothing exported) got a `--stop` that matched nothing — the regression flagged on
  #113982 and #113991.
- The home filter lives in `_find_stale_dashboard_pids(scope_home=...)` (the shape
  #113991 by @kvnloo used), so the `--stop` pre-check is scoped as well: a machine with
  only a foreign backend prints "No hermes dashboard processes running for this profile."
  instead of a silent exit.
- `_kill_stale_dashboard_processes` forwards `scope_home`; both call paths (`--stop` in
  hermes_cli/main.py and `_finish_dashboard_update_cleanup` in update_cmd_maint.py) pass
  their own `get_hermes_home()`.
- Test doubles widened for the new kwarg; #113982's test patches the scan seam instead of
  the (now filtering) finder. Docs: `--stop` row in the CLI reference.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 09:38:01 -07:00