Commit Graph

940 Commits

Author SHA1 Message Date
teknium1
c849bc383a refactor(env): agent.secret_scope.load_env_file is the only .env tokenizer; six hand parsers collapse onto it
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.

Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.

Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.

Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
2026-09-13 05:07:50 -07:00
teknium1
de2d6a1b93 fix(config): a fresh process recovers the last good config.yaml instead of running on defaults
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).

Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.

Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
2026-09-12 16:17:04 -07:00
teknium1
b78abb4710 fix(config): stop forcing 0700 on HERMES_HOME inside containers
Every start (and every `ensure_hermes_home` from a sibling CLI invocation)
chmod'd the data directory to 0700, wiping group/other bits and the ACL
mask on a bind mount shared with other containers (hermes-webui, Nix
desktop + dashboard). _secure_file already skipped containers for this
reason; _secure_dir did not.

In a container the directory mode is now left to the operator unless
HERMES_HOME_MODE is set explicitly, which is still applied. cron/jobs.py
had its own 0700/0600 copies that bypassed the managed/container rules;
they now delegate to the shared helpers so cron/output stops re-locking
the mount as well.

Fixes #10757
2026-09-12 08:43:43 -07:00
周鹤0668001310
92a8398087 fix(config): treat explicit false values in HERMES_MANAGED as unmanaged
Previously, setting HERMES_MANAGED=false (or 0, no, off) would be
interpreted as a literal managed-system name, causing is_managed()
to incorrectly return True and block update/config commands.

- Add _MANAGED_FALSE_VALUES tuple for canonical false strings
- Check false values before true values in get_managed_system()
- Add parametrized regression tests for all false variants

Fixes #12864
2026-09-12 08:41:00 -07:00
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
Siddharth Balyan
b4d04eb8fd Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding

* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg

The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.

On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.

* docs(tool-search): connectors section — remote tools through the bridge

The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.

* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots

dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.

The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.

The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).

Live, 311 local tools + gateway, before -> after:
  "send gmail email":           5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
  "read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
  "linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.

* refactor(tool-search): connector leg into tools/connector_search.py

tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.

No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.

* fix(tool-search): at most 7 queries per call, the gateway's search limit

One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.

The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.

* fix(tool-search): the model is told that connectors__ names are manage_connections accounts

tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.

The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.

Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.

Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.

* fix(connectors): /stop halts a connector batch before the next remote call

dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.

The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.

Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.

* test(connections): schema assertions become dispatch contracts

test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.

Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.

Test count in the file goes from 26 to 25.

* docs(tool-search): connector batches are one gateway request per entry

The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.

Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.

Docs only, no test.

* fix(tools): the between-turns refresh never rewrites the bridge tools

The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.

The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.

Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.

* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation

The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.

* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote

bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.

The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.

Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.

* fix(connectors): search keeps the twin a colliding name reaches, and says so

format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.

Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
2026-09-10 02:21:16 +05:30
Teknium
e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Teknium
7e4d02fef5 fix(cron): unpinned jobs run on their creation-snapshot model instead of failing closed
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).

The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.

Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".

Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).

The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
2026-09-09 04:32:13 -07:00
Teknium
d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
Teknium
bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium
1d5d059410 fix(gateway): stop time-triggered conversation rotation 2026-09-07 06:10:54 -07:00
Teknium
e6df8675f5 fix(config): diagnose unavailable home links without claiming YAML corruption
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.

Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.

Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.

Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
2026-09-07 04:46:37 -07:00
kshitijk4poor
256bd1adc9 feat(docker): terminal.docker_snap_compat opt-out for snap-packaged Docker under AppArmor (#9730)
On hosts where Docker ships as a snap (Ubuntu cloud images / Azure VMs), the
snap's AppArmor confinement turns two sandbox hardening flags into a dead
container at start: `--init` fails with "exec /sbin/docker-init: operation not
permitted" and `--security-opt no-new-privileges` then fails every exec the
same way ("exec /usr/bin/sleep: operation not permitted"). This is snapd
LP#1908448 — not probeable from the client, and docker_extra_args cannot remove
flags we add.

`terminal.docker_snap_compat: true` drops exactly those two flags; cap-drop ALL,
the tmpfs hardening, PID limits and the privdrop caps are unchanged, and a
warning is logged at container start. Bridged everywhere the other docker_*
keys are (CLI env map, gateway env map, `hermes config set` sync, terminal_tool
env read, the shared container_config shaper, DEFAULT_CONFIG).
2026-09-05 21:00:19 +05:30
Teknium
f1ccf436a2 feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:

- search: POST https://api.perplexity.ai/search (documented Search API),
  search_context_size=low so `snippet` stays description-sized;
  results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
  route behind `pplx content snippets` (the CLI's `content fetch` is
  deprecated upstream). web_extract has no query, so the URLs' path words
  serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
  backend set + credential ladder + availability probe (web_tools),
  registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
  dump key lists, nous_subscription direct-credential detection, setup
  summary, test conftests, docs.

Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
2026-09-04 07:17:00 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium
c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
fb14bc4e11 review-fix(whitespace): strip trailing whitespace and EOF blank lines introduced by this PR
Trailing-whitespace-only edits so 'git diff --check BASE HEAD' is clean
(16 diagnostics across 11 files). No code changes.
2026-09-03 09:31:54 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
kshitijk4poor
68cbe484a4 refactor(config): close empty-list root gap in strict check; assert tolerant returns
Review follow-ups on the final salvage stack:
- `fast_safe_load(f) or {}` collapsed a falsy non-mapping root (`[]`) to `{}`
  before the strict isinstance check, so an empty-list config evaded the
  raise the previous commit added. Only map a None document (empty file) to
  `{}`; every other non-mapping root now hits the strict branch.
- Docstring: the strict flag covers non-mapping roots too, not just parse
  failures.
- Tests: assert the tolerant call's return per shape instead of merely
  calling it; add the empty-list-root case.
2026-09-03 12:38:46 +05:30
kshitijk4poor
4d24f357a0 fix(config): strict version check also refuses non-mapping config roots
A config.yaml whose top-level value is a list or scalar parses fine, so the
strict check_config_version(raise_on_parse_error=True) from #101778 still
returned (0, latest) and migrate_config() proceeded: sanitize_env_file()
rewrote .env, then save_config()'s fail-closed guard raised RuntimeError.
Raise InvalidUserConfigError up front for that shape too, so the
"no side effect before the invalid config is surfaced" guarantee holds
for both invalid-config shapes. Tolerant callers are unchanged.

Tests: parametrize the two #101778 regression tests over malformed-yaml
and list-root; assert the tolerant call still does not raise. Also make
the test_update_autostash check_config_version mock kwarg-tolerant,
matching the author's fix in test_config.py.
2026-09-03 12:38:46 +05:30
EmanueleCornaggia
a3ceee4333 fix(config): refuse migration on malformed YAML
Make validation and migration paths distinguish parse failures from current configs so invalid YAML cannot trigger .env or config-side effects.
2026-09-03 12:38:46 +05:30
Teknium
94816e9ed7 refactor(hermes_cli): config.py compact strip/merge/env-load walkers, single-line section banners, guard flattening in show/migrate helpers 2026-09-03 00:07:31 -07:00
Teknium
696d6a165d refactor(hermes_cli): config.py AST-neutral layout pass (hug closers, reflow string literals) 2026-09-03 00:01:24 -07:00
Teknium
2f429b5533 refactor(hermes_cli): config.py _issue helper for validation appends, compact nested-key walkers, cron/coercion guard flattening 2026-09-02 23:59:04 -07:00
Teknium
adffcbf701 refactor(hermes_cli): config.py table the config get/set/unset usage text, merge check/migrate listing loops, compact manifest/profile injection 2026-09-02 23:54:14 -07:00
Teknium
e7034fea17 refactor(hermes_cli): config.py flatten guard ladders and redundant locals (unset walk, deep-merge, env save, terminal bridge) 2026-09-02 23:48:57 -07:00
Teknium
16d700646d refactor(hermes_cli): config.py inline single-use helpers (_run_once, managed-home, floor gate, cron fleet cover), collapse boolean ladders 2026-09-02 23:41:36 -07:00
Teknium
2f7be88154 refactor(hermes_cli): config.py shared _managed_source, packed re-export list 2026-09-02 22:14:28 -07:00
Teknium
e7a01c13d5 refactor(hermes_cli): config.py unify list-entry validation, offer-list prompts, banner/tools-suffix helpers 2026-09-02 22:07:57 -07:00
Teknium
6f7c22b9a7 refactor(hermes_cli): config.py docstring compaction and closer hugging (AST-neutral) 2026-09-02 21:37:58 -07:00
Teknium
da9a0551e4 refactor(hermes_cli): config.py hoist managed_scope/atomic_yaml_write imports (8 lazy import sites removed) 2026-09-02 21:06:53 -07:00
Teknium
f5699c6f60 refactor(hermes_cli): config.py AST-neutral layout compaction (hug short brackets, pack argument lists) 2026-09-02 20:51:13 -07:00
Teknium
8345581dbc refactor(hermes_cli): config.py env-file, display, cron-drift and config-set sections: shared helpers, dispatch tables, docstring compaction 2026-09-02 20:47:01 -07:00
Teknium
cd49358836 refactor(hermes_cli): config.py validation/migration/load/save phase helpers, shared _canonicalize_config and refuse-overwrite builder 2026-09-02 20:35:53 -07:00
Teknium
04dfa0a549 refactor(hermes_cli): config.py header/install-method/home-skeleton compaction, dispatch tables for parse-failure and update-command messages 2026-09-02 20:28:28 -07:00
Teknium
55f47d8e95 refactor(config): extract custom-provider resolution to hermes_cli/config_providers.py 2026-09-02 15:31:26 -07:00
Teknium
5d4b97939e refactor(hclib): config — config/config_migrations/tools_config/toolset_* dispatch tables and dedupe 2026-09-02 14:42:18 -07:00
Teknium
2475335443 fix(config): keep .env publishes inside the routed profile scope under multiplex (#88441)
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.

Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
2026-09-02 06:19:03 -07:00
Teknium
b53c50cf8a fix(config): resolve ${VAR} config refs through the profile secret scope (#84079)
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).

Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
2026-09-02 06:19:03 -07:00
Patryk Kopycinski
1eb71756af toolchain-self-improve: route API_SERVER_KEY to profile env 2026-09-02 06:19:03 -07:00
Teknium
552159d222 feat(cli,tui): collapse bell_on_clarify/approval into display.bell_on_prompt
One key covers every blocking prompt modal: clarify (single + batch),
dangerous-command approval (incl. computer_use), sudo password, and
secret capture. CLI gets a _ring_bell() helper shared with
bell_on_complete; TUI rings on clarify/approval/sudo/secret .request
events (isTTY-gated). 'hermes config' Bell summary shows both flags.
2026-09-02 05:34:35 -07:00
João Vitor Cunha
384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
João Vitor Cunha
ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
Vibe Coder
af709b31ca fix(config): redirect platforms.<name>.<display_setting> to display.platforms.<name>.<setting>
Problem A of #71047: 'hermes config set platforms.telegram.streaming false'
wrote to a key the gateway never reads. The connection config
(gateway/config.py) reads only token/extra/overrides from the top-level
platforms.<name> block, while per-platform display settings (streaming,
show_reasoning, tool_progress, ...) are resolved from
display.platforms.<name>.<setting> (gateway/display_config.py).

Redirect a platforms.<name>.<setting> key to
display.platforms.<name>.<setting> only when <setting> is a known per-platform
display setting (gateway.display_config.OVERRIDEABLE_KEYS), leaving real
connection keys (token, extra, channel_overrides, ...) untouched. The
gateway.display_config import is lazy/try-guarded to avoid a circular import
and to keep the CLI working where gateway is not importable.

Adds tests/hermes_cli/test_config_set_platforms_redirect.py covering the
redirect, connection-key non-redirect, and the no-stray-top-level-platforms
case.
2026-09-02 00:07:15 -07:00
Lakshya Agarwal
428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium
043c258ac2 fix(web): stale removed-backend config warns at startup and errors by name
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).

- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
  removed_backend_note(); selection_error() swaps in the specific
  removal explanation (removed in v0.21.0, keyless alternatives) while
  keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
  search_backend / extract_backend against the registry and emits a
  startup warning (deduped per stale value), surfaced by the existing
  print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
  per-capability keys, dedupe, healthy-config negative, live-backend
  failure text preserved.
2026-09-01 10:43:20 -07:00
Teknium
51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00