Commit Graph

3815 Commits

Author SHA1 Message Date
teknium1
9267706d6e fix(kanban): board identity follows the bound profile override, not the launch env
One resolver, hermes_cli.profiles.current_profile_name(): the HERMES_HOME override
(a multiplexed cron tick or routed gateway turn) names the profile first; the
dispatcher's HERMES_PROFILE pin is consulted only when no override is bound; the
process home last. #112888 added the HERMES_HOME-derived fallback but kept the env
pin FIRST, so under a multiplexer whose launch process carries a HERMES_PROFILE the
served profile's writes were still re-labelled with the host's name. The same
resolver replaces the per-module copies in hermes_cli/kanban.py, kanban_specify.py,
cron/lifecycle_guard.py and the kanban notify-target default.

Tests trimmed to two invariants (A->B->A under the override; control pin + generic
absence).

Closes #119859
Supersedes #112888
2026-09-23 08:09:47 -07:00
max
253a956144 fix(kanban): persist active-profile identity in comment author / created_by when HERMES_PROFILE unpinned
Board records (comment author, task creator) were written as the generic
"worker" whenever the dispatcher did not pin HERMES_PROFILE, even though
the active Hermes profile was resolvable. Add _persisted_identity():
environment (HERMES_PROFILE_NAME / HERMES_PROFILE) first, else the active
profile derived from HERMES_HOME via hermes_cli.profiles, else "worker".
Wire it into _handle_comment author, _handle_create created_by, and the
own-comment skip filter in inject_new_comments_from_env so a worker's own
notes never re-enter its live turn as fake operator steering.

Identity is never taken from tool args: board records are injected into
future workers' prompts, so a caller-supplied author override could forge
an authoritative-looking directive (see #19713).

Tests: regression tests for env present, env absent with active profile,
and no profile; injection echo guard without env profile.
2026-09-23 08:09:47 -07:00
teknium1
6b6c7f4a99 fix(web): a routed profile without an Exa/Parallel key is refused, not served on the launch key
Under multiplexing `get_provider_env` resolved a scoped miss (`get_env_value` -> None) by falling
through to a bare `os.getenv`, which is the LAUNCH profile's `.env`: a routed profile with no
EXA_API_KEY/PARALLEL_API_KEY of its own searched on another profile's key, and the (key, client)
cache in `cached_sdk_client` faithfully built it a client on that borrowed key. Independent-review
major on #120092 (pre-existing on main, but this PR's claim covered only the both-have-keys case).

The bare `os.getenv` rung now runs only when no profile secret scope is bound (stripped installs,
plain CLI, systemd-injected env keep working); with a scope bound a miss is "" and the provider
raises its usual missing-key error. Isolation is between profiles; no environ fallthrough on a
scoped miss.

Tests: the A->B->A row gains a no-key profile (refused, launch key never used) plus a
single-profile control; the openrouter video sibling gets the same absence row (already correct
on this PR's head, red on main).
2026-09-23 07:49:35 -07:00
John Paul Soliva
707852d92f fix(web): rebuild the Exa and Parallel clients when their API key changes
cached_sdk_client returned the client cached on tools.web_tools before it
read the key, so the Exa, Parallel and AsyncParallel clients kept the key
they were first built with for the life of the process. A key fixed in .env
and applied with /reload still sent the old key (401s until a restart), and
on a gateway serving multiplexed profiles every profile's Exa and Parallel
calls went out on whichever profile's key built the client first, billed to
that account. A key removed from the environment also kept being used.

Resolve the key on every call and reuse the cached client only when it was
built with that key; the slot now holds (key, client) as one value so two
builds racing under different keys cannot record one key beside the other
key's client. Firecrawl already compares its credential before reusing its
client; this brings the two SDK-backed providers in line.
2026-09-23 07:49:35 -07:00
teknium1
ebb9bc5699 fix(env-policy): a profile gate is a platform prefix plus a gate suffix, not a bare name shape
is_profile_gate_env matched any name containing _ALLOWED_ / _ALLOW_ALL_ / ... so an operator's own
DEMO_ALLOWED_SENDER (script data, not a Hermes authorization gate) was deleted from every routed
child env built with strip_launch_profile=True, breaking no_agent cron scripts under the host
gateway. The owner prefix now has to be a platform: built-in Platform values, bundled platform
plugins (directory names and manifest aliases), runtime-registered plugin adapters, plus GATEWAY_
(pairing) and QQ_ (qqbot). Every gate the adapters read still strips; operator variables survive.

Fixes #119539
2026-09-23 06:46:02 -07:00
beardthelion
9cae47fd90 fix(mcp): scope late OAuth attempts by profile
_LATE_ATTEMPTS was keyed by session key alone while live.py keys the
open-operation table by (profile_key, session_key) for exactly this
collision: two multiplexed profiles can carry the same session key (the
api_server binds X-Hermes-Session-Key verbatim, and it is client-chosen).

A parked OAuth attempt could therefore be adopted by a different
profile's next turn: adopt_late_connections would poll it, register the
same-named server from that profile's config, and enable it - a grant
authorized under profile A materializing under profile B, or being
discarded so the owning profile never adopts it.

Key the table by (profile_key, session_key), matching live.py's
documented identity pairing. The detached no-card path never opens an
operation, so its profile_key is empty; fall back to the calling
thread's home, which close() runs under the turn's profile scope.
2026-09-23 06:46:02 -07:00
beardthelion
6aaa1c480b fix(profile-scope): bind home and secret tokens inside the try that resets them
Several profile-scope binders called set_hermes_home_override and
set_secret_scope before the try/finally that releases them. A raise in
scope setup (a corrupt or removed profile home) propagated with the
foreign override still bound to the caller's context, silently re-homing
every later read — and in _reregister_orphaned_adopters it also skipped
every remaining adopter. The set calls now run inside the try with
None-guarded resets; the same shape is fixed in the routed-turn scope,
the cron external worker (which also leaked the multiplex flag), the
kanban worker scope, the MCP OAuth paths, the launch-profile policy,
and model_switch, which releases partially-bound scopes on raise.
2026-09-23 06:46:02 -07:00
teknium1
6ba45d4adb test(process): a kill reaps children the shell spawns during the grace window
Real-process invariant for the re-snapshot fix: kill a background job while its
bash -lic shell (which ignores SIGTERM) is about to start three children; all three
must be dead once kill_process returns. Red on origin/main (3/3 orphaned every run),
green with the fix. Needs live_system_guard_bypass because the late children are
SIGKILLed after their shell, when they are no longer in the pytest subtree.
2026-09-23 06:01:37 -07:00
teknium1
6d885a296a test(mcp): callback-latch tests keep the listener up until the browser stand-in is done
tests/tools/test_mcp_oauth_callback_latch.py failed on main four times in
two days (Sep 20-21) with ConnectionRefusedError / ConnectionResetError
on the stand-in's second request. The production waiter polls
_result_taken every 500 ms and shuts the listener down as soon as the
first terminal /callback lands; the test relied on all follow-up GETs
(favicon, duplicate callbacks) landing inside that poll window, which a
loaded runner does not guarantee.

Gate the waiter thread's _result_taken poll on a "requests sent" event
so the listener stays bound until every stand-in request has been
answered. The handler's own reads are untouched, so the first-wins latch
is still what is being exercised. Waiter timeout 4 s -> 30 s: the happy
path finishes in one poll, and a stuck waiter still fails the test.

Live repro: a 0.7 s gap between the stand-in's requests turns the CI
signature deterministic on base (2/2 ConnectionRefusedError); with the
fix the same gap passes 2/2, and the unmodified test passes 5/5.
2026-09-23 05:35:36 -07:00
teknium1
1393403dd7 fix(process_registry): strip bash startup noise until real output, not only from the first read
A tty-less `bash -lic` writes "cannot set terminal process group" and
"no job control in this shell" as two separate write() calls. The reader
cleaned shell noise from the FIRST chunk only, so whenever it woke between
the two writes (loaded CI runners) the second line landed in
output_buffer as the process's only "output".

Symptoms on main: tests/hermes_cli/test_process_dock.py painted
"last: bash: no job control..." instead of "starting" (3 main reds,
Sep 21-22) and tests/tools/test_process_registry_list_exit.py's probe
woke on that noise before the writer had printed, so the completion
event lacked "owner-output" (2 main FLAKY frames). The same leak reaches
users through process.list output_preview and the live-work dock.

Keep stripping leading noise from every chunk until the process has
produced non-blank output; after that, matching text is the process's
own. The list_exit probe now waits for the writer's marker rather than
for any bytes.

Live repro (stand-in shell reproducing bash's two writes with a 50 ms
gap, real spawn_local + select/read1 reader): base 10/10 buffers carry
"bash: no job control in this shell\n"; fixed 0/10.
2026-09-23 05:34:25 -07:00
teknium1
7ed6534c7e Merge origin/main: browser fence composed with the dispatch/retry split (#115184); server registration, i18n, contracts 2026-09-23 03:25:06 -07:00
teknium1
f287036ea3 test: purge low-value tests, lane py19 (370 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
f596ed1595 test: purge low-value tests, lane py18 (303 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
5f6b1d251f test: purge low-value tests, lane py17 (495 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
524c38a98a test: purge low-value tests, lane py16 (624 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
090b06c345 test: drop the bundled-hindsight tests, retarget the generic ones
Deleted: the seven files whose subject was plugins/memory/hindsight itself
(provider, config schema, env perms, local-runtime hint, templates, health
grace timeout, root guard) and the hindsight-only cases inside the multiplex
identity-scope, provider-thread, BOM-tolerance and session-switch suites.

Retargeted: the dashboard memory-provider config-surface tests used hindsight
as the only bundled provider with a flat `<home>/<name>/config.json` declared
schema; they now install a synthetic user plugin (`flatprov`) into the isolated
HERMES_HOME so the generic declared/live PUT + GET paths stay covered. The
lazy_deps plugin-owned-range test keeps mem0ai as its subject.
2026-09-23 01:16:41 -07:00
alt-glitch
31e59b9441 feat: setup agent can search the catalog and install plugins and skills through the approval card
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog:

- search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with
  whether each is already installed in the default profile. Read-only.
- install: opens the same connection operation manage_connections opens, with rows of kind
  plugin / skill. Nothing installs until the user approves a row. An approved row installs
  into `default` (or the profile the Advanced modal named) through dashboard_install_plugin /
  the hub's headless install, so the catalog pin, kill list, security scan and live
  activation (#119644) are the host's. The row settles with the live MCP tool names and the
  plugin's skill.

The model sends catalog ids and an action only; every other key is refused before anything
runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own
text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the
result is the `hermes plugins install` / `hermes skills install` pointer.

- contract: plugin/skill targets follow the MCP transitions.
- run.apply_answer / reissue route a card answer to the module that owns the operation.
- tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never
  deferred behind tool_search.
- docs: tools reference, toolsets reference, plugin catalog page.

Linear NS-964.
2026-09-23 08:04:33 +05:30
Siddharth Balyan
3e00a356a4 fix(mcp): one reader for mcp_servers.<name>.enabled (#119567)
The `enabled` key had four parsers. The MCP client (`_parse_boolish`) read
`enabled: 0` as on; the toolset resolver and editor (`_parse_enabled_flag`)
read it as off. The server list (`summarize_server`, `/api/mcp/servers`) read
any non-`False` value as on, so `enabled: "false"` showed on while the agent
skipped it. The catalog and `hermes mcp list` accepted only true/1/yes, so
`enabled: on` showed off while the server ran.

`tools/mcp_tool_common.py::mcp_server_enabled` is now the only reader, and
every surface calls it. `_parse_boolish` treats YAML numbers by truthiness
(0 off, other numbers on). Everything else keeps the client's semantics:
the off words are off, absent / null / junk stay on, with the existing
warning for junk.

The desktop MCP page mirrors the rule in `serverEnabled`
(`apps/desktop/src/lib/mcp-servers.ts`). One case table
(`mcp-enabled-cases.json`) drives the Python invariant test and the vitest
test, so the page and the runtime cannot drift apart again.
2026-09-22 22:00:24 +00:00
Siddharth Balyan
eb5d735e9a Application-backed MCP servers: live endpoint per attempt, tools follow app liveness, errors and the tool catalog say why (NS-943) (#119050)
* feat: expose interactive and live endpoint facts

* feat: hydrate application-backed MCP liveness

* feat: export interactive session fact

* fix: classify unsupported declared hosts

* test(mcp): a runtime file without a token connects without an Authorization header

* fix(mcp): server_json liveness fields default to http/token/pid; an invalid declaration degrades to static instead of failing the server task

* fix(mcp): a connected server is offerable regardless of the interactive-session rule; the rule only shapes the not-connected sentence

* test(mcp): follow main's ToolSearchConfig fields and patch _bump_server_error at its origin module
2026-09-22 17:54:42 +05:30
Siddharth Balyan
4094ab610d hermes_platform.resolver: locate/inspect/probe tiers, AppResolver, gh lookup migrated (NS-921) (#118065)
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates

Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.

* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources

Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.

* refactor(copilot): gh candidates through locate_command and the Homebrew table

First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.

* feat(platform): availability() over an application declaration

locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.

* feat(platform): application declarations parsed into AppDef per OS

The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.

* feat(mcp): check_fn honours a registered application declaration

_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).

* feat(skills): requires_apps gate through registered declarations

Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.

* docs: application declarations page

The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.

* test(platform): declaration parser, availability, gates

The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
2026-09-22 17:08:30 +05:30
Siddharth Balyan
de5bb5cf13 Merge pull request #117863 from NousResearch/sid/ns-920-hermes-platform-host
hermes_platform.host: one machine-facts package, WSL/container predicates, NVIDIA ARM64 SoC recognizer (NS-920)
2026-09-22 16:49:01 +05:30
kshitijk4poor
92dd332192 test(skill_ledger): gc test ages its orphan blobs past the in-flight grace
The blob GC contract is now "unreferenced AND older than
_BLOB_GC_GRACE_SECS"; the existing test stored its orphans and swept
them in the same second, which the grace window (correctly) refuses.

PROOF: test_gc_blobs_removes_only_unreferenced red on the previous head
(1 deleted expected, 0 observed), green now; 11 passed in the file.
2026-09-22 16:16:30 +05:30
kshitijk4poor
65d33f20ba test(skill_ledger): shorten the concurrency test's forced-race wait to 1 s
`b_done.wait(3.0)` always times out on correct code (the lock blocks writer B
until A's sweep finishes), so it was a fixed 3 s sleep. 1.0 s keeps the same
guarantee: on the broken-lock path B's append lands in milliseconds, and no
assertion depends on the timing.

PROOF: run_tests.sh ledger + AX files 29 passed; test_skill_ledger.py file
time 5.6 s -> 3.9 s.
2026-09-22 16:16:30 +05:30
kshitijk4poor
ed6b55d1d8 fix(skill_ledger): blob GC keeps unreferenced blobs younger than an hour
`snapshot_paths` / `_store_blob` write blobs OUTSIDE `_ledger_lock`, seconds
before the row that references them is appended. Since `_maintain_size` now
runs `gc_blobs()` on any append that trims, a sweep in another process during
that window deleted the in-flight capture: its row then landed with sha256
values `read_blob` cannot resolve and `rollback_entry` failed.

`_gc_blobs_locked` now skips any unreferenced blob file (incl. `.tmp-*`) whose
st_mtime is newer than `_BLOB_GC_GRACE_SECS = 3600`; no new config key.
Docs (curator.md) and the `gc_blobs` docstring say so.

PROOF: probes/s5_final_inflight_blob.py before -> `P1 blob survives P2's
sweep: False ... read_blob: None`; after -> `True ... read_blob: b'P1
in-flight skill file'`. test_concurrent_appends_never_lose_a_middle_row gained
a fresh + 2h-aged unreferenced blob pair: red with the grace set to 0 (fresh
blob deleted), green at 3600 (aged deleted, fresh kept). ruff clean,
check-windows-footguns --all clean, run_tests.sh ledger + AX files 29 passed.
2026-09-22 16:16:30 +05:30
kshitijk4poor
c86607124f fix(skill_ledger): hold .locks/ledger.lock across append + sweep and the CLI compact/GC path
`_maintain_size()` now runs inside every `append_entry`, and `compact_ledger`/
`_trim_oldest` do read-bytes -> `os.replace` over the ledger with no cross-writer
lock, so a concurrent appender's O_APPEND write landing on the replaced inode was
silently lost — and `gc_blobs` then deleted that row's blobs. `_ledger_lock()`
(same `skill_file_lock` idiom as `_skill_mutation_lock`; fcntl/msvcrt, re-entrant,
no-op where neither exists) is held across the append AND the sweep, and inside
`compact_ledger`/`_trim_oldest`/`gc_blobs` for the `curator ledger --compact` path.

PROOF: tests/tools/test_skill_ledger.py::test_concurrent_appends_never_lose_a_middle_row
forces writer B to append between writer A's sweep read and its os.replace; every
row must be present or counted as trimmed. Green 2/2 (3.5 s). With `_ledger_lock`
made `contextlib.nullcontext()`: red 4/4 ("none silently lost").
2026-09-22 16:16:30 +05:30
kshitijk4poor
25d240da11 perf(config): ledger_enabled() reads the config read-only too
`ledger_enabled()` is the first call on every `append_entry`, and it still
paid `load_config()` (a deepcopy of the cached config) after the rest of the
append path had moved to `load_config_readonly()`. It only reads one key via
`cfg_get`, so the read-only load is safe. The one test that patches
`load_config` to turn the ledger off now patches `load_config_readonly` too.

PROOF: PYTHONPATH=<wt> import + `ledger_enabled()` → True;
scripts/run_tests.sh tests/tools/test_skill_ledger.py
tests/tools/test_computer_use_ax_walk_bound.py → 28 passed, 0 failed.
2026-09-22 16:16:30 +05:30
kshitijk4poor
9fa284fa2a test(computer_use): prove capture() forwards the AX-walk bound onto the CaptureResult
`capture()` builds `CaptureResult(ax_max_elements=getattr(self,
"_ax_max_elements_sent", 0))` but no test reached that line: the invariant
test stopped at the private `_ax_max_elements_sent` attribute and the hint
tests build CaptureResult by hand, so replacing the getattr with `0` survived
every test. The invariant test now stubs `_capture_window_state` on the stub
backend (`(None, None, [], "")`, plus a no-op `_set_active_target`) and asserts
`stub.capture("ax").ax_max_elements == 350`; the private-attribute assert is
dropped.

PROOF: with `getattr(self, "_ax_max_elements_sent", 0)` replaced by `0` in
cua_backend_capture.py the test fails (`assert 0 == 350`); reverted, 5 passed.
2026-09-22 16:16:30 +05:30
kshitijk4poor
767e393c5b test(computer_use): collapse the three capped-walk hint cases into one parametrized test
The AX-walk PR carried 2 invariant tests + 3 `TestCappedWalkHint` tests, one
over the ≤2-per-PR shape gate. The three cases differ only in (n, bound,
expected) so they are one `@pytest.mark.parametrize` test:
[(8, 8, True), (8, 50, False), (8, 0, False)]. The AX PR now carries 2 + 1.

PROOF: same assertions as before, expressed as `is capped` / `is not capped`;
tests/tools/test_computer_use_ax_walk_bound.py 5 passed (2 + 3 params).
2026-09-22 16:16:30 +05:30
kshitijk4poor
dd6229fe11 fix(curator): ledger trim splits rows on the physical newline only
`_trim_oldest` used `str.splitlines()`, which also breaks on U+2028/U+2029/
U+0085. `append_entry` writes rows with `ensure_ascii=False`, so a row whose
evidence contains one of those code points was counted as two lines and
rewritten as two malformed physical lines — contradicting "lines in the
retained tail are never rewritten". Split on b"\n" (dropping the trailing
empty element), re-join with b"\n", and account sizes in bytes.

Sibling readers (`compact_ledger`, `gc_blobs`, `list_entries`) use the same
idiom pre-existing on main and are left for a follow-up.

PROOF: tests/tools/test_skill_ledger.py::test_trim_oldest_when_still_over_cap
gains one assertion (a U+2028 row survives a trim byte-for-byte); it fails
with the old `splitlines()` body ("At index 85 diff: b'\n' != b'\xe2'") and
passes with this change. 23 passed.
2026-09-22 16:16:30 +05:30
kshitijk4poor
76c3b0c735 fix(computer_use): the backend reports the AX-walk bound it actually sent
tool.py inferred 'walk capped' by re-reading cua's private config through
_cua_configured_ax_max_elements() — a leaky abstraction that also cost a
config load per capture. The backend now records the max_elements it put on
the get_window_state call and returns it on CaptureResult.ax_max_elements
(defaulted field, 0 = unbounded / vision); _capture_view compares
len(cap.elements) >= cap.ax_max_elements > 0 and the getattr guards go
since the field is always set. Hint text and line order unchanged.

PROOF: tests/tools/test_computer_use_ax_walk_bound.py 5 passed (hint tests
now build CaptureResult(ax_max_elements=bound)); the new
_ax_max_elements_sent assertion is red with cua_backend_capture.py
reverted to HEAD~1 ('AttributeError: ... no attribute _ax_max_elements_sent',
1 failed 4 passed). Probe: 200/200 -> capped hint + 'full ' dropped;
199/200 and 8/0 -> full tree promised. check-windows-footguns --all clean.
2026-09-22 16:16:30 +05:30
kshitijk4poor
559f7d170f perf(config): read-only config loads on the ledger append and capture hot paths
_max_ledger_bytes() runs on every ledger append (right after ledger_enabled()
already deep-copied the config) and _computer_use_cfg() now runs twice per
computer-use capture; neither mutates the result, so use
load_config_readonly() as hermes_cli/config.py documents for read-only hot
paths (skips the deepcopy, ~half the cache-hit cost). Tests that patched
load_config now patch load_config_readonly as well (test-only edit).

PROOF: real import under PYTHONSAFEPATH shows load_config() ==
load_config_readonly() (same dict shape; defaults 5242880 / 200 resolve);
tests/tools/test_skill_ledger.py + test_computer_use_ax_walk_bound.py +
test_skill_ledger_delta.py -> 39 passed, 0 failed.
2026-09-22 16:16:30 +05:30
kshitijk4poor
462af0ab65 test(curator): make the low-water assertions actually bite
The old test's direct _maintain_size() call was a no-op: the five seeding
appends had already swept the file to ~7.3 KB (under the 8 KB cap), so the
'<= 80%' assertion after the follow-up append held only because that 3 KB
append re-fired the whole sweep and landed under 6553 B by row-size
coincidence. Seed with the bound disabled, enable it for the direct call,
assert st_size <= int(cap*0.8) right after the sweep, and count
compact_ledger calls on the follow-up (1 KB) append: it must be zero.

PROOF: with _TRIM_LOW_WATER = 1.0 (temporary edit, reverted) ->
'AssertionError: assert 7319 <= 6553', 1 failed; at 0.8 -> 1 passed.
Gate S5-2ab mutation M3 previously SURVIVED 28/28.
2026-09-22 16:16:30 +05:30
kshitijk4poor
d2e492da66 docs(curator): trim drops the oldest lines regardless of shape
_trim_oldest never parses lines: it drops the OLDEST ones whatever they
are, so 'malformed lines always survive' was false. Reword docstring,
curator.md and the kept test's docstring to the true guarantee: lines in
the retained tail are never rewritten or parsed, so a malformed line there
survives verbatim. No behaviour change.

PROOF: probes/s5_trim_probe.py (cap 4096, malformed first line, 4 fat
appends) -> old malformed survived: False, contradicting the old wording;
the kept test asserts a LAST-line malformed row, which is what the new
wording promises.
2026-09-22 16:16:30 +05:30
kshitijk4poor
0f00cb1a9e test: the auto-compact test seeds legacy fat rows so it is red on main
Main's append-time _delta already empties identical manifests, so three appended fat entries never crossed the cap and the sweep never fired — the test was green without _maintain_size. Seed pre-delta rows directly and let the next append trigger the sweep.
2026-09-22 16:16:30 +05:30
kshitijk4poor
9fd07283d4 test: keep the invariant tests for the ledger bound and the AX-walk bound
Ledger (#118696): keep `test_auto_compact_triggers_at_threshold` and
`test_trim_oldest_when_still_over_cap`; drop the disabled-threshold and
failure-never-blocks tests — both pin the shape of a try/except, not a
contract. AX walk (#116512): keep `test_configured_value_reaches_the_driver_args`
and `test_zero_disables_the_bound`; drop the four constant/clamp
change-detectors that would fail on any retune of the default.
2026-09-22 16:16:30 +05:30
kshitijk4poor
87276637ed fix(computer_use): a capped accessibility walk says so in the capture hint
`_capture_view`'s hint promised a "full element tree saved to elements_file",
but once `computer_use.ax_max_elements` bounds the driver's walk the spill
holds only the first N nodes. When the capture returned at least the bound
(bound > 0), drop "full" and add "accessibility walk capped at N elements;
pass app= to narrow" so the model knows how to reach the rest.
2026-09-22 16:16:30 +05:30
kshitijk4poor
74069eb428 fix(curator): ledger trim stops at a low-water mark
At the cap every append ran compact + trim + gc_blobs: trimming to exactly
`skills.ledger_max_bytes` leaves the file over the cap again on the very next
append. Trim to 80% of the cap instead (one rule derived from the cap, no
second config key), and run `gc_blobs()` only when the trim actually dropped
rows — the only path that can orphan a blob. `_trim_oldest` now reports the
dropped-line count and reuses `_read_ledger`, so an unreadable/undecodable
ledger skips the trim the same way it aborts compaction and blob GC.
Malformed lines are still kept verbatim.

The low-water-mark idea credits @fangliquanflq (#118690).
2026-09-22 16:16:30 +05:30
vgarlu
a85b6f9cdd perf(computer_use): bound the driver's accessibility-tree walk per capture
`capture()` sent `get_window_state` with only `pid`, `window_id` and `session`, so the driver
walked the target's entire accessibility tree before Hermes trimmed the surfaced element list
to `_DEFAULT_MAX_ELEMENTS` (100) and spilled the rest to a cache file. Every node past the
first ~100 was paid for and discarded.

Measured on macOS (cua-driver 0.28.2, M-series), same window, bound fixed at 200:

| Target | Unbounded walk | Bounded |
|---|---|---|
| 1,444-node Chrome window | 540 ms | 83 ms |
| 456-node Finder window (pathologically slow AX surface) | 6.9 s | 0.6 s |

The bound is lossless for the response: the bounded element list is a *prefix* of the
unbounded walk (checked by role/label/depth at 200/400/600/1000), so nothing the model sees
changes.

The bound is internal and config-driven — `computer_use.ax_max_elements`, default 200,
`0` disables (driver default) — and deliberately NOT a model-facing schema parameter:
`max_elements` was removed from the tool schema on purpose, so a capture's surfaced window
stays at its fixed default.

(cherry picked from commit 9f07ff428bd1fcfbcc0e5089d834e1f33d8d1b76)
2026-09-22 16:16:30 +05:30
Konstantin Khlopkov
71932b150b fix(curator): keep the skill ledger size-bounded (dedup rewrite + oldest-entry trim)
(cherry picked from commit 42572ebf207d411edd534055b1d750d91efcd4f7)
2026-09-22 16:16:30 +05:30
kshitijk4poor
71cb99136a test(bot-mode): drop the unused cache fixture and profile write from the skill-scope test
`_fresh_cache` cleared `bot_mode_probe._cached`, which `capability_fingerprint` is
documented as deliberately never touching; the `home` fixture wrote a hermes-bots
managed `profile.yaml` that no assertion reads (every test compares digests before
and after edits to the skills tree only).

PROOF: `pytest tests/tools/test_capability_fingerprint_skill_scope.py` → 2 passed
with both removed; `ruff check` clean; footguns clean.
2026-09-22 15:55:09 +05:30
kshitijk4poor
e1760ee4c9 test: keep the invariant tests for the jonpol01 perf trio
Each pick keeps the two tests that pin its invariant; the rest restated
the same behaviour from other angles.

- #117399: keep test_a_repeated_bullet_is_kept_once_in_first_position and
  test_a_bullet_with_continuation_lines_is_never_touched (9 removed).
- #117789: keep test_archiving_a_skill_does_not_move_the_epoch and
  test_installing_a_real_skill_still_moves_the_epoch (4 removed).
- #117951 (rewritten): one test, test_recording_a_reply_does_not_open_a_
  second_connection — counts sqlite3.connect per record_obligation, red on
  main by construction (main opens 2).
2026-09-22 15:55:09 +05:30
John Paul Soliva
6df4571177 perf(bot-mode): the capability epoch counts the skills the model can invoke, not archived ones
`capability_fingerprint`'s skills component globbed `**/SKILL.md` raw, while every other reader of
that tree — `skills_list`/`skill_view`, the prompt's own `<available_skills>` index, `skill_count`,
`prompt_size` — walks it through `iter_skill_index_files`, which prunes EXCLUDED_SKILL_DIRS
(`.archive`, `.curator_backups`, `node_modules`, `.venv`, …) and each package's support dirs.

So the fingerprint was the only place counting an archived skill as a skill. Since a changed digest
rebuilds the stored Bot Chat system prompt, archiving a skill — or the curator writing a backup, or
a skill vendoring node_modules — re-prefilled the whole prompt cache for a prompt whose skills
section had not changed.

Measured on a 120-skill profile: 11.42ms -> 6.26ms, and it runs twice per turn (turn start and the
prompt-restore path). The surface drops 120 -> 89 entries, and all 31 dropped sit under a pruned
dir — none of them appear in the prompt index, skills_list or skill_count.

One-time effect: the digest changes once on upgrade, so each Bot Chat rebuilds and re-prefills once.
After that the epoch stops moving for churn it should never have tracked.

Regressions: installing a real skill still moves the epoch; archiving one settles and further
archived packages never move it again; no EXCLUDED_SKILL_DIRS entry can move it (parametrised over
the whole set); a curator backup and a skill's `references/` dir are both inert; and the epoch
tracks exactly what the skills walker reports. Restoring the raw glob fails 20 of the 21.

Fixes #117788

(cherry picked from commit a6b50a40e117f8afca28208541be714578bbc568)
2026-09-22 15:55:09 +05:30
Siddharth Balyan
70f5dc5f46 feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
2026-09-22 13:57:51 +05:30
alt-glitch
62aa4b7d0d refactor(voice): WSL detection through hermes_platform.host.runtime.is_wsl
One runtime predicate keeps WSL detection from diverging between call sites.
The shared result is cached per process.
2026-09-22 01:13:23 -07:00
alt-glitch
0bdf33b7d9 test: repoint WSL/container probe patches to hermes_platform.host.runtime (group B)
The probe cache now belongs to hermes_platform.host.runtime.
Call-site name patches stay where production resolves them.
2026-09-22 01:13:23 -07:00
kshitijk4poor
78dfd6e6e7 docs(skills): explain why an undecodable ledger warns, without review jargon
The docstring of test_list_entries_treats_an_undecodable_ledger_as_empty_and_warns
cited "(re-gate M8/S1)" (tests/tools/test_skill_ledger_delta.py:88; re-gate
G2 S1) — a review-process artifact meaningless to future readers. State the
behavioural reason instead: silence would make a corrupt ledger look like a
fresh install with no history.
2026-09-22 13:40:40 +05:30
kshitijk4poor
d3684a7d79 test(skills): assert the missing-ledger case emits no warning records
test_list_entries_is_silent_when_the_ledger_is_merely_missing checked
`"listing empty" not in caplog.text` (tests/tools/test_skill_ledger_delta.py:40-44;
simplify quality #1) — a negative substring match that goes vacuously
green if the message is ever reworded, and stays green if a *different*
warning fires. Assert on caplog.records at WARNING or above instead, which
checks the actual invariant (no warning for a merely absent ledger)
independent of wording.
2026-09-22 13:40:40 +05:30
kshitijk4poor
dd309d046f fix(skills): warn when list_entries() hits a corrupt ledger, and pin the UnicodeError guard
list_entries() (tools/skill_ledger.py:414) caught (OSError, UnicodeError) but
returned [] silently, so a ledger that EXISTS but cannot be decoded made
`hermes curator ledger` print "ledger is empty" and `hermes curator rollback`
report "no ledger entry with id" with no hint that the file is damaged — while
gc_blobs()/compact_ledger() already warn on the same condition (:309, :345).
Now warn for any read failure except FileNotFoundError (a missing ledger is the
normal empty state). `%s`-lazy logger call, same "ledger unreadable" prefix.

Also adds the regression test the UnicodeError widening (2213062284) lacked:
re-gate mutation M8 reverted the except clause to `except OSError:` and the
whole suite stayed green. New tests write b"\xff" to ledger_path() and assert
list_entries() == [], get_entry() is None and "listing empty" in caplog; a
second test pins that a merely missing ledger stays silent.

Red/green: `except OSError:` -> UnicodeDecodeError (1 red); warning dropped ->
caplog assert red; warning made unconditional -> missing-ledger test red.
2026-09-22 13:40:40 +05:30
kshitijk4poor
142401166f test(skills): make the unreadable-ledger fixture transport-agnostic
The `unreadable` param of test_gc_keeps_rollback_blobs_when_the_ledger_cannot_be_read
(tests/tools/test_skill_ledger_delta.py:122-130) monkeypatched Path.read_text class-wide
to raise PermissionError for the ledger path. That couples the test to one read primitive:
switching gc_blobs() to read_bytes()/open() would make the param pass vacuously without
testing anything. Replace it with a real filesystem condition — a directory at the ledger
path — which fails read_text() with an OSError on every platform (IsADirectoryError on
POSIX, PermissionError on Windows) regardless of how the file is read. The monkeypatch
fixture and the with-block indentation go away; the directory is removed before the
ledger bytes are restored for the rollback check.

RED with the guard defeated (`return 0, 0` -> `lines = []`):
  [unreadable] assert (2, 14) == (0, 0)  (blobs were deleted); GREEN at head.
2026-09-22 13:40:40 +05:30
kshitijk4poor
0c5beaad77 test(skills): pin compact_ledger() no-op on an undecodable ledger
Gate mutation M2 (moving `raw.decode("utf-8")` outside the try in compact_ledger(),
tools/skill_ledger.py:306-307) survived the existing suite: nothing exercised compaction
over ledger bytes that are not UTF-8. Add the 4-line regression: with the ledger set to
b"\xff", compact_ledger() returns (0, 0, 0), the bytes are left exactly as they were,
and the "compaction skipped" warning from the previous commit is emitted.

RED with the decode moved outside the try (UnicodeDecodeError escapes), GREEN at head.
2026-09-22 13:40:40 +05:30