Commit Graph

2366 Commits

Author SHA1 Message Date
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
Siddharth Balyan
70f5dc5f46 feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
2026-09-22 13:57:51 +05:30
teknium1
43504da0e5 test(config): regression guard so config.yaml comment loss cannot come back
- tests/hermes_cli/test_config_yaml_comment_preservation.py: a hand-commented config goes
  through config set, config unset, save_config (plugin enable + memory.provider), a
  _config_version migration bump and a direct atomic_config_write; every comment, the key
  order and the quoted "off" must survive, and the boilerplate is appended only on create.
  6/8 red on base.
- scripts/check_config_yaml_writers.py (wired into the lint workflow): AST scan that fails
  on any atomic_yaml_write / yaml.dump / yaml.safe_dump of a config path, or any PyYAML dump
  inside the config system, outside the writer module. Flags all nine base-tree writers.
2026-09-22 01:07:59 -07:00
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
ethernet
4d737cb461 Merge branch 'fix/r2-mac-shell' into ethie/pm-clean 2026-09-22 00:28:37 -04:00
ethernet
c9bd7459c2 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	scripts/run_tests_parallel.py
#	tui_gateway/model_switch.py
2026-09-22 00:28:30 -04:00
ethernet
75a13386fb fix: make macOS shell and mode contracts portable
The macOS runner uses Bash 3.2, which rejects parameter case conversion.
Normalize the GHCR owner with portable tr and verify mixed-case input.

The macOS runner can clear setgid from a directory when chmod applies 2770.
Compare scratch permissions with the native chmod result while preserving all
bits the host accepts.

Tests: scripts/run_tests.sh tests/scripts/test_termux_build_driver.py tests/test_scratch_dir.py
Shell: bash -n scripts/termux/build_builder_image.sh
2026-09-22 00:09:09 -04:00
teknium1
439eb0395e fix(tests): relocated pytest basetemps stop piling up in $HOME; runner sweeps roots killed runs left
Two leaks from the test temp plumbing:

tests/conftest.py relocates pytest's basetemp out of the native Hermes
home (#111101) with mkdtemp(dir=native.parent), which is the operator's
$HOME, and nothing removed it: 123 hermes-pytest-basetemp-* dirs (552 MB)
appeared there in a day, one per bare pytest process. The relocated
basetemp now goes into one prunable root (/var/tmp/hermes-pytest on
POSIX, a non-dotted sibling of the native home elsewhere), is removed at
pytest_unconfigure, and idle siblings from killed runs are swept on entry.

scripts/run_tests_parallel.py deletes each per-file temp root in finally,
but a SIGKILLed runner (tool timeout, stray pkill) never gets there and
leaks one root per in-flight worker: 983 r-* roots (3.4 GB) in three
days. The runner now sweeps 24h-idle roots at start and forces read-only
permission fixtures writable before rmtree instead of skipping them.
2026-09-21 20:05:27 -07:00
ethernet
5f1d3294b9 Merge branch 'rev/tests-infra' into ethie/pm-clean 2026-09-21 19:52:58 -04:00
ethernet
1afa348e1c Merge branch 'rev/pm-core' into ethie/pm-clean 2026-09-21 19:52:58 -04:00
ethernet
843a0095ec Merge branch 'rev/termux' into ethie/pm-clean 2026-09-21 19:52:58 -04:00
ethernet
2250b501b4 refactor(pm): rename cache_lock.py to uv_cache_prune.py
The module prunes uv caches to the lock; it takes no lock. Two importers.
2026-09-21 19:09:56 -04:00
ethernet
4fda664b0f tests: resolve platforms() specs in the lane selector and runner note
scripts/ci/list_os_marked_tests.py matched the literal lane word inside a
platforms() string, so 80 files gated platforms("posix") never reached the
macOS lane ("posix = linux or macOS" was false for CI), and "any" files
reached none. The selector now resolves specs the way the conftest gate
does (posix ⊇ linux+macos, any ⊇ all, "not X" admits the rest) and only
looks inside mark.platforms(...) calls, dropping a false positive whose
only "macos" was inside a generated-file string.

scripts/run_tests_parallel.py still grepped the retired linux_only/
macos_only/windows_only names, so its "N files SKIPPED on this host" note
had been silent since the migration. It now shares the selector's
resolver and names the spec and the lane(s) it runs on.

Restore the platforms("windows") mark that
test_suppress_platform_ver_console_stubs_syscmd_ver lost in the
windows_only migration (its docstring still declared it); it passed
vacuously on Linux and was deselected on the Windows lane.
2026-09-21 18:53:24 -04:00
ethernet
3c53fbdc8b install.ps1: emit the -Json failure frame when Fail ends a stage
Fail ends the script with `exit 1`, which unwinds past the stage
dispatcher's try/catch, so the catch that frames failures as JSON never
ran for the installer's own fatal errors: a `-Stage repository -Json`
run whose clone failed printed the reason via Write-Host only and put
zero frames on stdout (verified on Windows: 0 frames before, 1 after).
Only thrown exceptions were framed.

Fail now emits the failure frame itself when running under -Stage -Json,
so every fatal path yields exactly one frame carrying the original
reason, matching install.sh's EXIT-trap framing.
2026-09-21 18:44:30 -04:00
ethernet
d7466aa3ba release: one strict stable-tag grammar shared by every selector
scripts/releases/semver.STABLE_TAG accepted any-width majors, so
is_valid_version('2026.9.21') was True and docker.require_stable_tag /
stable.py / release.py admitted the legacy CalVer tags that
hermes_cli.source_releases and get_last_tag() already refused. A
workflow_call carrying GitHub's current 'latest' (v2026.9.21) would have
passed the docker publish gate.

hermes_cli.update_channel already owns the canary tag shape; it now owns
STABLE_TAG_RE too (three-digit major cap, no leading zeros, no suffix)
and every stable selector imports it. release.py drops its private
_SEMVER_TAG_RE + CalVer exclusion pair, which the capped major makes
redundant.
2026-09-21 18:42:01 -04:00
ethernet
5e8fbd4919 install.sh: refuse Termux hosts and point at the APT package
check_platform accepted Termux as plain Linux, so `curl | bash` on a
phone walked the source-install ladder: a glibc uv, a lock whose CPython
is the bundled bionic build with no Android wheels, and on-device sdist
builds. The signed APT package is the only supported Termux shape, so
detect Termux the way the runtime does (TERMUX_VERSION or the com.termux
PREFIX) and stop before any stage with `pkg install hermes-agent` and
the setup docs URL. --json surfaces the reason in the stage frame.
2026-09-21 18:41:42 -04:00
ethernet
04bfa58d47 install.sh: refuse to clone over an unrelated destination
The clone publication step did `mv <staged>/tree "$INSTALL_DIR"`. When
INSTALL_DIR already existed as a non-git directory, mv moved the
checkout INSIDE it as INSTALL_DIR/tree and the stage reported success,
leaving later stages to read pm/lock.json from a directory that holds
the user's files instead of a checkout. An existing file made mv fail
with a generic "cannot publish" error.

Before staging the clone, refuse a destination that exists and is not a
Hermes checkout (non-empty dir, file, or symlink) and say what to do; an
empty directory is taken over so the checkout lands AT the path.
2026-09-21 18:40:30 -04:00
ethernet
6a6771e5e7 termux: stamp the deb as a self-contained runtime, not a bundled payload
build_deb.sh wrote HERMES_DESKTOP_VARIANT=bundled, so the Termux stamp
carried payload=bundled and every 'bundled' reader treated the tree as
the repo/ of an Electron payload. hermes uninstall --data then called
resolve_bundle_layout on it and rejected the plan: the deb has no
enclosing app to protect, so data-only removal was impossible on Termux.

Add a 'runtime' variant to write_install_stamp.py for a sealed CLI
runtime with no desktop app around it. It stays sealed for the update
gate and channel identity (source_check, update_channel accept it next
to bundled/light) while is_bundled_payload keeps answering False, so
cleanup planning protects the APT-owned package tree the ordinary way
and never asks where the app is.
2026-09-21 18:37:21 -04:00
ethernet
f999b5ecd2 wip fixin stuff 2026-09-21 18:19:03 -04:00
ethernet
14b3232cbd fix(test-runner): key the scratch root by user, not a shared literal
The runner's per-run temp roots live under a fixed `/var/tmp/hermes-pytest`, chosen for
real reasons (disk-backed, not hidden, short enough for AF_UNIX sun_path). But a fixed
literal in a world-writable sticky dir belongs to whoever creates it first: a root-owned
root — a container or system-service run — makes every later `makedirs`/`mkdtemp` there
fail with EPERM for every other user on the host, with no way back that does not need
root. That is exactly what happened on luna: /var/tmp/hermes-pytest is root:root 755, so
every local suite run by the login user died at `runner crashed: PermissionError(13)`
before collecting a single test.

Key the name by uid (with the same non-/var/tmp fallback), so no run can be blocked by
another user's leftovers. Invariant test: two uids never share a scratch root, and the
root is created under the expected parent.
2026-09-21 17:24:29 -04:00
ethernet
4efb36f81e merge: integrate upstream desktop features through PM preparation
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.

Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.
2026-09-21 15:17:44 -04:00
ethernet
034c3fdb02 ci: update ty to avoid full-tree allocation failure 2026-09-21 14:43:35 -04:00
ethernet
d2c60dbdf8 fix(test-runner): preserve PATHEXT for native PowerShell children 2026-09-21 14:03:12 -04:00
ethernet
3a6e61b189 ci: drop the PM toolchain smoke matrix; key the tools cache on the PM code
The 6-target cold->warm smoke existed to prove setup-pm still cold-boots
when its code changes without a lock bump, and to reach linux-arm64,
darwin-x64 and win32-arm64. Both are now covered without a dedicated
workflow:

* The tools cache key hashes pm/**, the action itself and
  scripts/ci/setup_toolchain.py, not only pm/lock.json. A provisioning
  change misses the cache on every lane that uses the action, so the
  cold path runs where the tests already are.
* tests-os gains a windows-11-arm leg running the same windows-marked
  files (tests/pm carries ten of them). It needs the ARM64 build deps
  because several extras build from sdist, and fewer workers on the
  4-core runner.

The Windows SDK adapter test is the one thing left that no other lane
ran natively; it keeps its two Windows runners under
windows-bundle-sdk.yml, path-triggered on the signing scripts. The
run-scoped cache cleanup workflow and its script only served the smoke
and go with it.
2026-09-21 12:07:53 -04:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
ethernet
9ee4397408 fix: round-9 Windows lane — retry argv keeps --force; mint fixture carries venv_sync; stamp probe reports the child
- desktop-update/windows.ps1: the legacy-install retry re-sends the identical request
  (--force included); the contract test compares both attempts.
- test_mint_launchers: the bootstrap imports hermes_cli.venv_sync/steward before
  prepare_launch can return early for a fixture repo; copy them into the tree.
- test_source_build_env: when the pwsh child writes no stamp, fail with the child's
  stdout/stderr instead of a bare FileNotFoundError (the Windows lane hides the cause).
2026-09-21 06:48:59 -04:00
ethernet
f7561667c7 fix: round-8 CI backlog (Windows-only lane)
- tests/conftest.py: `real_bash` fixture — the Windows runners resolve `bash` to System32's
  WSL launcher (UTF-16 "no installed distributions", exit 1); prefer Git for Windows'. Used
  by the setup-pin, install stage-frame and source-launcher shell tests.
- source-build-env.ps1: Test-Path before Remove-Item — under $ErrorActionPreference='Stop'
  a missing identity variable aborted the try block before the child ran (Windows PS 5.1
  raised where pwsh on Unix did not).
- desktop-update/windows.ps1: `--force` precedes the target arguments (the hand-off contract
  test reads argv in that order).
- test_install_ps1_desktop_stage: assert the current contract — the shared completion tail
  (source_completion.py --desktop) builds the products and -IncludeDesktop selects the desktop
  product inside `products` rather than adding a stage. The test predated the completion-tail
  refactor and had been red on the Windows lane since.
- test_windows_native_support: the restart watcher argv is `runtime_command` shaped
  ([python, -I, -c, bootstrap, pid, delay, ...]).
- test_mint_launchers: create the fixture repo's pm/ dir before copying pm/environments.py.
- test_browser_use_pm: console-script launchers report sys.argv[0] without `.exe`.
- test_update_stale_gateway_yield (from main): `_verify_fleet_after_update` has no
  `node_failures` here (PM owns node).
2026-09-21 06:03:05 -04:00
ethernet
9ac6af3809 merge origin/main (28 commits) into ethie/pm-clean
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
2026-09-21 05:07:21 -04:00
ethernet
834e20ee42 fix: round-7 CI backlog (win32 cold lanes, Windows-only lane, 8 python)
- run_tests.sh forwards `ProgramFiles(x86)` (setuptools finds vswhere under it; the name
  cannot be read with ${!var}) instead of DISTUTILS_USE_SDK/MSSdk, which made setuptools take
  `link.exe` from PATH — Git for Windows' coreutils `link` under a bash step. The PATH
  reordering in windows-build-deps.ps1 is dropped too: it broke the initializer contract
  test ("PATH lost or reordered") on win32-x64.
- tests/hermes_cli/test_stale_pid_guard.py: the POSIX kill-path test lived inside the
  platforms("windows") class and asserted `_kill_pids_posix` on a Windows host; it is a
  module-level platforms("posix") test now.
- test_update_concurrent_quarantine: pin hermes_constants.project_venv_dir to the fixture's
  venv layout (a CI checkout has no venv/ dir, so the launcher prefix resolved elsewhere).
- test_update_autostash: commit hermes_bootstrap.py into the fixture repo, or
  `stash push --include-untracked` takes it and the baseline import probe dies before its
  health marker (CI has no editable finder to fall back on).
- test_update_finish: the two npm-graph completion tests skip when the checkout has no
  node_modules (the Python lane never runs `npm ci`; install E2E exercises the real path).
2026-09-21 04:54:26 -04:00
teknium1
2463550c97 feat(website): plugin pages render the README from the pinned commit by default
The README section was gated on `readme: true` in the catalog YAML and no entry set it, so
all 222 plugin pages shipped without one. READMEs now render for every GitHub/GitLab entry
(fetched at the reviewed sha, subdir first then repo root, common casings and docs/README.md
as fallbacks); `readme: false` opts an entry out. Build proof: 222/222 READMEs rendered.
2026-09-21 01:15:33 -07:00
ethernet
1d602a04e3 fix: round-6 CI backlog (32 python, docker arm64, win32-arm64, windows lane, desktop lint)
- pm.environment: the `--no-install-package <name>` root read parses [project].name without
  tomllib on the pre-3.11 bootstrap python (Docker arm64 stage_runtime). Both parsers proven
  to agree on the real pyproject.
- windows-build-deps.ps1 / run_tests.sh: with DISTUTILS_USE_SDK, setuptools takes link.exe
  from PATH; under a bash-hosted step Git for Windows' coreutils `link` shadowed MSVC's.
  The MSVC linker directory now leads PATH (cl.exe was already found — this was the next
  failure in the ruamel-yaml-clib build).
- tests/install/e2e-assets/source-build-env.ps1: clear the identity variables through the
  env: drive — [Environment]::SetEnvironmentVariable(..., 'Process') on .NET/Unix does not
  reach spawned children, so the stamp child still saw GITHUB_SHA. Red→green under nix pwsh.
- old-updater surface: warm_agent_browser_npx_cache is a permanent def in both facade and
  sibling (the frozen surface names both); its compat pointer is retired, and the shims test's
  __module__ check holds.
- tests re-seamed / de-faked: import guard tests use the probe_root fixture (CI has no
  editable finder), the takeover child tree gets a hermes_constants stub, posix.sh hand-off
  test pins HERMES_HOME (our script honours the ambient one), the two Windows-layout
  PYTHONPATH tests are platforms("windows") (they fake `Lib/site-packages` on Linux; pm's
  site_packages() is host-correct), setup-pin test picks Git bash over System32's WSL stub,
  expose_cli's Windows test asserts the installer-convention convergence (the branch retired
  "windows-installer-owned"), plugin-manifest satisfied-dep fixture uses a core dep (pyyaml is
  gone), housekeeping test yaml imports go through hermes_yaml.
- apps/desktop main.ts: three imports restored in round 4 whose users main removed.
2026-09-21 03:44:00 -04:00
ethernet
f3b1399211 fix: round-5 CI backlog after the 779-commit main merge
Merge fallout (my resolution errors, all caught by CI):
- hermes_cli/backup.py + gateway.py: `theirs` on those hunks re-imported clusters HEAD had
  already moved to backup_restore.py / kept in the facade. backup.py loses the 349-line
  duplicate (main's #110179 fix is ported into backup_restore._import_db_member); the
  systemd service-unit cluster returns to gateway.py (PM's _prepare_service_launcher /
  _pm_managed_node_dirs / _systemd_command have no home in main's extraction) with main's
  utf-8-sig read. gateway_service_unit.py is dropped.
- gateway/run.py: main's plugin-update chore is not profile-scoped (the housekeeping
  ordering test pins the scope/drain sequence).
- pyproject + 30 test files: `import yaml` -> `import hermes_yaml as yaml` (pm-clean has no
  pyyaml); gateway/config._bundled_platform_manifest_name reads through hermes_yaml.
- tests re-seamed onto pm-clean's shape: residency admission (installed_engine),
  supervisor child env (binary is a constructor argument), update import guard
  (update_cmd_deps is gone; our probe already scrubs PYTHONPATH — both #115032 invariants
  pass), shallow-count git responses (stash path asks `status --porcelain -z`); dropped
  tests for retired code (_run_node_bootstrap/_ensure_tui_node, Windows resume demotion).
- tests/tools/test_local_env_blocklist.py: restore the two helpers the suite-reduction
  commit dropped and the blocklist import.

Real fixes:
- pm: classify_uv_failure/ResolutionConflict move beside the uv runner (pm.environment,
  stdlib-only). pm.workspace imports tomllib at module level and cannot load on the 3.10
  bootstrap python that streams uv output in the Docker arm64 image.
- tools/browser_tool.warm_agent_browser_npx_cache: back as a permanent definition — it is on
  the frozen old-updater surface, and the revert-scheduled compat pointer does not count.
- hermes_cli/memory_setup: the dashboard's pip row uses pm.environments.
  running_from_selected_environment for installed vs restart_required.
- scripts/windows-build-deps.ps1: export DISTUTILS_USE_SDK/MSSdk so setuptools trusts the
  primed MSVC environment instead of asking vswhere (`env -i` test runner on win32-arm64
  compiling ruamel-yaml-clib); run_tests.sh forwards them.
- tests/pm/test_windows_build_deps.py: start the protocol test from a parent env without the
  toolchain variables the runner job already exports.
- tests/conftest.py scrubs HERMES_BUNDLED_PLUGINS (Nix-wrapped hermes on the dev host);
  tests/home_io_guard.py treats sys.path site-packages under the real home as the
  interpreter's installation (PM-activated developer shell).
- tests-js: four `curly` lint errors from main's new scripts.
2026-09-21 02:47:48 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
ethernet
3b181dd848 merge origin/main into ethie/pm-clean 2026-09-21 00:38:40 -04:00
teknium1
10c273813d feat(plugin-catalog): screenshots and readme entry fields for the plugin pages
Two optional, submitter-controlled fields on a catalog entry feed the
entry's own page at /docs/plugins/<name>:

- `screenshots:` — up to 6 https URLs on GitHub hosts (same host rule as
  `image`, so the site never fetches from third-party hosts and a raw URL
  pinned to the sha is as immutable as the code).
- `readme: true` — the docs build renders the README from the PINNED
  commit (raw.githubusercontent.com / gitlab.com raw at <sha>), never live
  content, so what a user reads is what the reviewer read.

Validator rejects malformed values (admission), the loader parses and
drops off-host screenshots with a warning (client), and the extractor
emits `screenshots`, `readme`, `readmeUrl` and a `maintainerSlug` for the
author pages. Tests on all three.
2026-09-20 20:43:13 -07:00
liuhao1024
04845f5f3e fix(desktop): skip the local gateway restart on update when the Desktop is remote-served
A Desktop whose active connection is remote (SSH/remote/cloud, including the
registry primary) owns no local messaging gateway, yet the update hand-off
always ran `hermes update --gateway`. On hosts where launchd/service recovery
fails, the updater falls back to a detached local `gateway run --replace`;
with the same Telegram bot token as the remote VPS gateway, the two processes
compete for getUpdates and Telegram rejects one consumer, taking the
production bot offline (#117529).

Pass the ownership down the hand-off: globalRemoteActive() now adds
--no-gateway (posix) / -NoGateway (windows) when the Desktop is remote-served,
and both orchestrators drop --gateway from every update invocation (initial +
retry). The local-ownership default keeps --gateway exactly as before.
2026-09-20 19:30:16 -07:00
fangliquan
8c5a5deb79 fix(install): probe the command link directory for dependencies 2026-09-20 15:22:22 -07:00
joaomarcos
f58605b5c6 fix(desktop): keep Windows update hand-off hidden 2026-09-20 11:14:36 -07:00
liuhao1024
8828e356f7 fix(tests): let run_tests_parallel take the explicit file list from a file
--files carries the whole list as one argv element, and Linux caps a
single argument at MAX_ARG_STRLEN (128 KiB) - a much smaller limit than
ARG_MAX. The whole-suite list (~210 KB) dies with E2BIG in execve before
the runner's first line runs, so 'run the whole suite except one file'
cannot be expressed through --files at all.

Add --files-from PATH (or '-' for stdin), one path per line, mutually
exclusive with --files. A bare '-' after --files-from is normalized to
the '='-joined form because argparse treats '-' as a positional.
2026-09-20 10:51:40 -07:00
teknium1
6159bf4d87 runner: drop the RLIMIT_DATA worker cap
On the CI runner the pillow-heif HEIF encode in tests/tools/test_image_source.py hangs under
the cap (2/2 runs, 36/40 then 300 s timeout; 40/40 in 26 s on main). The wheel's encoder spins
on a failed allocation instead of erroring, so a heap cap on C code trades an OOM for a hang.
Keep the leak fix and the no-relaunch-on-kill rule.
2026-09-20 09:00:16 -07:00
teknium1
439ebe0ae9 fix(tests): stop the code_kernel reader-thread leak that OOM-killed test workers; cap worker heap
tests/tools/test_local_env_blocklist.py::TestPythonpathSelectiveStrip::
test_execute_code_composition_strips_inherited_hermes_entries hands code_kernel a MagicMock
process whose stdout/stderr only fake read(); the kernel drains with read1(), which on a bare
MagicMock never returns EOF. _stdout_reader died on `buf += chunk`, but _stderr_reader's
`while chunk := stderr.read1(4096)` spun forever appending mocks (each call growing
mock_calls) in a daemon thread that outlived the test: ~1 GB/min until the kernel killed the
worker. Five OOM incidents on this file (08-30, 09-13, 09-14, 09-16, 09-19), always blamed on
whichever test ran next. The fake now returns EOF from read1() as well.

Runner guardrails so the next runaway is a traceback, not a swap storm:
- each pytest worker runs under RLIMIT_DATA (8 GiB, Linux; HERMES_TEST_WORKER_MEM_GB, 0 = off).
  RLIMIT_AS is avoided on purpose: browsers spawned by tests reserve huge address space.
- a worker killed by signal or the file timeout is never --file-retries relaunched; a runaway
  relaunched while the first tree is still being reaped doubled the damage on 09-16.

Live: the file went from 20 min / 20 GB to 5.6 s / 140 MB. An allocate-forever probe dies with
MemoryError in 5 s; a SIGKILL'd worker launches once on this runner, twice on base.
2026-09-20 09:00:16 -07:00
ethernet
925c08ceca fix: CI python-tests backlog — no import-time dependency syncs, CI-shaped test fixtures
Production:
- agent/bedrock_adapter.py, agent/vertex_adapter.py: pm.ensure_import ran at
  module import. In any process that imports these modules without a committed
  PM selection (CI's build_environment test venv, a fresh checkout) that sync
  rebuilt the dependency environment mid-process and replaced sys.path with a
  generation missing the caller's own packages (anthropic, aiohttp vanished).
  The extra is now ensured at first client build / credential request.
- plugins/platforms/matrix/adapter.py: a complete install needs no
  ensure_and_bind round trip; only a partial one syncs.
- tools/browser_tool.py: drop the facade's duplicate warm_agent_browser_npx_cache
  shim; the compat pointer already resolves to browser_tool_install.

Test harness:
- tests/home_io_guard.py: PATH-entry probes (shutil.which) and the running
  interpreter's own installation (stdlib reads, realpath ancestry, fixture
  symlinks into it) are not Hermes state; a patched Path.expanduser must not
  crash the guard. run_tests.sh no longer filters PATH — the guard owns it.
- tests/tui_gateway/conftest.py: import hermes_bootstrap before any file opens
  a MagicMock hermes_constants window (6 files exited the process at boot).
- tests/hermes_cli/conftest.py probe_root: scratch checkouts the import guard
  probes need hermes_bootstrap.py (the launcher imports it).
- tests/pm/_fixtures.py stage_host_python: a copied relocatable python needs
  its stdlib beside it (No module named 'encodings' on CI).
- tests/install/e2e-assets/smoke-env.mjs: dependency-free env shaping so the
  source-build-env probe runs under bare node (main deleted the Playwright
  entry it was imported through).
- adapt main's new tests to branch seams (model_metadata_http, launch
  completion tail, CI toolchain exports uv after python, source_launch
  hermes_cli stub, systemd_notify single marker).
2026-09-20 11:47:06 -04:00
ethernet
a41aabff6c fix: platform legs — docker arm64 bootstrap, win32-arm64 native builds, nix desktop-backend cleanup
- pm.environment owns _RESOLVER_MARKERS: the streaming uv runner imported
  pm.workspace, whose tomllib import fails on the 3.10 system python that
  bootstraps the Docker arm64 image (No module named 'tomllib'). Invariant test
  proves runtime staging needs neither tomllib nor pm.workspace.
- run_tests.sh forwards the MSVC/SDK/Rust/OpenSSL toolchain variables through
  its env -i scrub so PM tests that compile ruamel-yaml-clib on Windows arm64
  find cl.exe (previously 'Visual C++ 14.0 or greater is required').
- nix desktop-backend check: the spawned backend outlives cage's process group
  and kept writing under the temp HERMES_HOME during rmtree; stop every
  process bound to the throwaway HOME before cleanup.
- windows: test_launcher_runtime_selection imports runtime_state from
  hermes_cli (moved in bbec973514); the ' spaced ' suffix row loses its
  trailing space on win32 (the filesystem strips it).
- macOS: test_sealed_worker_command copies the interpreter into the payload
  (the escape guard resolves symlinks) and links the host lib tree.
2026-09-20 10:42:30 -04:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
teknium1
40bfcbda34 chore(lint): P33 — module-level constant frozen from a bridged HERMES_* env var
The advisory profile-scope lint had no pattern for the shape behind #115635 (a module
CONSTANT = _float_env(...)/os.environ.get("HERMES_...") of a var gateway/run.py bridges from
config.yaml). Fires on origin/main's `_RECONNECT_ATTENTION_AFTER_SECONDS`; 6 advisory hits on
head, all env-only knobs or already keyed per home (agent/redact.py).
2026-09-19 22:34:52 -07:00
ethernet
9214174e07 fix(ci): lint legs after the main merge
- check_no_tmp_literals: resolve scratch via tempfile/os.tmpdir; the termux
  container mount point is one marked variable per script
- ruff TID251: desktop E2E fixtures may reach PM internals like tests do;
  the pm.runtime_stage ban message no longer names a module that never existed
- auth_codex: build the capped httpx stream subclass on first use so importing
  hermes_cli.auth_codex no longer forces httpx (the lazy proxy in auth_constants
  was defeated by a module-scope base class; broke lanes without httpx)
- desktop-smoke: launchApp is a parameter; the bundle-env test substitutes a
  refusing launcher instead of letting Playwright spawn a dying binary
  (3 unhandled rejections failed the tests-js lane)
2026-09-19 23:25:18 -04:00
ethernet
119bc61c3b fix: read text files with encoding=utf-8-sig in modules merged from main 2026-09-19 22:57:45 -04:00
ethernet
e1576d06a6 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
2026-09-19 22:57:07 -04:00
teknium1
ece388eca8 fix: test-runner scratch root is /var/tmp/hermes-pytest; two tests follow TMPDIR
~/.cache/hermes-pytest failed seven files: a dot-dir ancestor made the hidden-dir search tests
see every fixture as hidden, and the longer root pushed the PulseAudio/voice AF_UNIX test
sockets past sun_path. /var/tmp is the FHS disk-backed temp root (never tmpfs), non-hidden,
and shorter than the old /tmp root. test_tool_result_storage asserts STORAGE_DIR instead of a
literal, and the zh-Hans bot-mode mirror follows the English code block it must copy.
2026-09-19 10:44:26 -07:00