Commit Graph

2105 Commits

Author SHA1 Message Date
teknium1
0988a99743 test(runner): forward HERMES_E2E_REQUIRE_TUI, CI and GITHUB_ACTIONS through the hermetic env
run_tests.sh starts every run from env -i with an allowlist, so the e2e
job's HERMES_E2E_REQUIRE_TUI=1 never reached the terminal suite (a missing
ui-tui build skipped instead of failing) and the upgrade suite's CI branch
in sandbox_required_reason() was dead. With ui-tui/dist removed and
HERMES_E2E_REQUIRE_TUI=1: before, 2 skipped; after, 2 failed.
2026-09-23 17:55:23 -07:00
teknium1
88e45a48b6 fix: run_tests.sh forwards HERMES_RUN_E2E so the documented iron-proxy E2E runs
The egress developer guide says `HERMES_RUN_E2E=1 scripts/run_tests.sh
tests/agent/test_iron_proxy_e2e.py`, but the runner's `env -i` dropped the
variable, so all three cases skipped silently. Forward it next to the other
opt-in knobs (HERMES_RUN_SLOW_PET_TESTS, HERMES_E2E_BROWSER); before: 3
skipped, after: 3 passed. The test docstring no longer names a --run-e2e
option that never existed.
2026-09-23 10:34:54 -07:00
teknium1
0f1544f903 fix(desktop): report a failed gateway restart as a manual step, not a failed update
The salvaged restart (#119815) exited 9 when `hermes gateway start --all`
failed, which makes the hand-off's finally block write ok:false, show the
error finale and log "detached update FAILED" in the Desktop even though
the code, venv and Desktop build all verified. Move the restart after the
outcome is settled and surface a restart failure through Write-Result's
existing manual flag instead: the Desktop's boot dialog shows the
`hermes gateway start --all` follow-up, and the update stays a success.
Publish a UI stage so the marquee names the step.

Extend the salvaged test with the exit-code invariant (red on the PR's
own head) and correct the update_cmd_windows docstring that said the
Desktop never restarts the messaging gateway.

Closes #119809
2026-09-23 06:05:29 -07:00
KoNit-K
d10f1735bc fix(desktop): restart gateways after Windows update 2026-09-23 06:05:29 -07:00
teknium1
7ed6534c7e Merge origin/main: browser fence composed with the dispatch/retry split (#115184); server registration, i18n, contracts 2026-09-23 03:25:06 -07:00
teknium1
4b17b284ce test: keep cross-cutting architecture lints; drop stale OS-fake baseline entries
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)

Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
2026-09-23 03:15:26 -07:00
teknium1
38137d2b6e test: purge low-value tests, lane js06 (375 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
85cd82f1dd test: purge low-value tests, lane py15 (375 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
teknium1
4f1dcc2dad test: purge low-value tests, lane py11 (303 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
alt-glitch
5c215ffa21 feat: catalog marks onboarding plugins; catalog rows carry the app's presence
The onboarding card needs to list catalog plugins beside the hosted
connectors (NS-960 D1, D4) and grey a plugin whose app is absent (D5).
The catalog had no curated flag, and the manage_catalog row left
app_state empty.

- `onboarding: true` and `title` on catalog entries (loader, validator,
  docs); set on blender, nvidia-app and nvidia-broadcast.
- hermes_cli/plugin_catalog_presence.py reads the plugin.json at the
  catalog's pinned commit once per pin and judges its app declaration with
  the hermes_platform resolver the installer and the Plugins-tab pill use.
  No declaration or an unreadable one is `unknown`, never `present`.
- `plugins.manage action=onboarding` lists the curated entries this OS
  runs (platform mismatch is the only exclusion) with app_state and the
  sentence the card greys the row with.
- manage_catalog plugin rows now carry app_state and the catalog title.
- A live entry that differs from the in-tree entry at the same pin (new
  metadata) now follows the same newer-catalog rule as a new pin, so a
  checkout that adds `onboarding` is not masked by a published doc that
  predates it.
2026-09-23 09:26:07 +05:30
Siddharth Balyan
70f5dc5f46 feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
2026-09-22 13:57:51 +05:30
teknium1
43504da0e5 test(config): regression guard so config.yaml comment loss cannot come back
- tests/hermes_cli/test_config_yaml_comment_preservation.py: a hand-commented config goes
  through config set, config unset, save_config (plugin enable + memory.provider), a
  _config_version migration bump and a direct atomic_config_write; every comment, the key
  order and the quoted "off" must survive, and the boilerplate is appended only on create.
  6/8 red on base.
- scripts/check_config_yaml_writers.py (wired into the lint workflow): AST scan that fails
  on any atomic_yaml_write / yaml.dump / yaml.safe_dump of a config path, or any PyYAML dump
  inside the config system, outside the writer module. Flags all nine base-tree writers.
2026-09-22 01:07:59 -07:00
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
teknium1
439eb0395e fix(tests): relocated pytest basetemps stop piling up in $HOME; runner sweeps roots killed runs left
Two leaks from the test temp plumbing:

tests/conftest.py relocates pytest's basetemp out of the native Hermes
home (#111101) with mkdtemp(dir=native.parent), which is the operator's
$HOME, and nothing removed it: 123 hermes-pytest-basetemp-* dirs (552 MB)
appeared there in a day, one per bare pytest process. The relocated
basetemp now goes into one prunable root (/var/tmp/hermes-pytest on
POSIX, a non-dotted sibling of the native home elsewhere), is removed at
pytest_unconfigure, and idle siblings from killed runs are swept on entry.

scripts/run_tests_parallel.py deletes each per-file temp root in finally,
but a SIGKILLed runner (tool timeout, stray pkill) never gets there and
leaks one root per in-flight worker: 983 r-* roots (3.4 GB) in three
days. The runner now sweeps 24h-idle roots at start and forces read-only
permission fixtures writable before rmtree instead of skipping them.
2026-09-21 20:05:27 -07:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
2463550c97 feat(website): plugin pages render the README from the pinned commit by default
The README section was gated on `readme: true` in the catalog YAML and no entry set it, so
all 222 plugin pages shipped without one. READMEs now render for every GitHub/GitLab entry
(fetched at the reviewed sha, subdir first then repo root, common casings and docs/README.md
as fallbacks); `readme: false` opts an entry out. Build proof: 222/222 READMEs rendered.
2026-09-21 01:15:33 -07:00
teknium1
10c273813d feat(plugin-catalog): screenshots and readme entry fields for the plugin pages
Two optional, submitter-controlled fields on a catalog entry feed the
entry's own page at /docs/plugins/<name>:

- `screenshots:` — up to 6 https URLs on GitHub hosts (same host rule as
  `image`, so the site never fetches from third-party hosts and a raw URL
  pinned to the sha is as immutable as the code).
- `readme: true` — the docs build renders the README from the PINNED
  commit (raw.githubusercontent.com / gitlab.com raw at <sha>), never live
  content, so what a user reads is what the reviewer read.

Validator rejects malformed values (admission), the loader parses and
drops off-host screenshots with a warning (client), and the extractor
emits `screenshots`, `readme`, `readmeUrl` and a `maintainerSlug` for the
author pages. Tests on all three.
2026-09-20 20:43:13 -07:00
liuhao1024
04845f5f3e fix(desktop): skip the local gateway restart on update when the Desktop is remote-served
A Desktop whose active connection is remote (SSH/remote/cloud, including the
registry primary) owns no local messaging gateway, yet the update hand-off
always ran `hermes update --gateway`. On hosts where launchd/service recovery
fails, the updater falls back to a detached local `gateway run --replace`;
with the same Telegram bot token as the remote VPS gateway, the two processes
compete for getUpdates and Telegram rejects one consumer, taking the
production bot offline (#117529).

Pass the ownership down the hand-off: globalRemoteActive() now adds
--no-gateway (posix) / -NoGateway (windows) when the Desktop is remote-served,
and both orchestrators drop --gateway from every update invocation (initial +
retry). The local-ownership default keeps --gateway exactly as before.
2026-09-20 19:30:16 -07:00
fangliquan
8c5a5deb79 fix(install): probe the command link directory for dependencies 2026-09-20 15:22:22 -07:00
joaomarcos
f58605b5c6 fix(desktop): keep Windows update hand-off hidden 2026-09-20 11:14:36 -07:00
liuhao1024
8828e356f7 fix(tests): let run_tests_parallel take the explicit file list from a file
--files carries the whole list as one argv element, and Linux caps a
single argument at MAX_ARG_STRLEN (128 KiB) - a much smaller limit than
ARG_MAX. The whole-suite list (~210 KB) dies with E2BIG in execve before
the runner's first line runs, so 'run the whole suite except one file'
cannot be expressed through --files at all.

Add --files-from PATH (or '-' for stdin), one path per line, mutually
exclusive with --files. A bare '-' after --files-from is normalized to
the '='-joined form because argparse treats '-' as a positional.
2026-09-20 10:51:40 -07:00
teknium1
6159bf4d87 runner: drop the RLIMIT_DATA worker cap
On the CI runner the pillow-heif HEIF encode in tests/tools/test_image_source.py hangs under
the cap (2/2 runs, 36/40 then 300 s timeout; 40/40 in 26 s on main). The wheel's encoder spins
on a failed allocation instead of erroring, so a heap cap on C code trades an OOM for a hang.
Keep the leak fix and the no-relaunch-on-kill rule.
2026-09-20 09:00:16 -07:00
teknium1
439ebe0ae9 fix(tests): stop the code_kernel reader-thread leak that OOM-killed test workers; cap worker heap
tests/tools/test_local_env_blocklist.py::TestPythonpathSelectiveStrip::
test_execute_code_composition_strips_inherited_hermes_entries hands code_kernel a MagicMock
process whose stdout/stderr only fake read(); the kernel drains with read1(), which on a bare
MagicMock never returns EOF. _stdout_reader died on `buf += chunk`, but _stderr_reader's
`while chunk := stderr.read1(4096)` spun forever appending mocks (each call growing
mock_calls) in a daemon thread that outlived the test: ~1 GB/min until the kernel killed the
worker. Five OOM incidents on this file (08-30, 09-13, 09-14, 09-16, 09-19), always blamed on
whichever test ran next. The fake now returns EOF from read1() as well.

Runner guardrails so the next runaway is a traceback, not a swap storm:
- each pytest worker runs under RLIMIT_DATA (8 GiB, Linux; HERMES_TEST_WORKER_MEM_GB, 0 = off).
  RLIMIT_AS is avoided on purpose: browsers spawned by tests reserve huge address space.
- a worker killed by signal or the file timeout is never --file-retries relaunched; a runaway
  relaunched while the first tree is still being reaped doubled the damage on 09-16.

Live: the file went from 20 min / 20 GB to 5.6 s / 140 MB. An allocate-forever probe dies with
MemoryError in 5 s; a SIGKILL'd worker launches once on this runner, twice on base.
2026-09-20 09:00:16 -07:00
teknium1
a5ea71c409 Merge origin/main: browser env socket-safe TMPDIR composed with the Bot Desktop display 2026-09-19 22:51:13 -07:00
teknium1
40bfcbda34 chore(lint): P33 — module-level constant frozen from a bridged HERMES_* env var
The advisory profile-scope lint had no pattern for the shape behind #115635 (a module
CONSTANT = _float_env(...)/os.environ.get("HERMES_...") of a var gateway/run.py bridges from
config.yaml). Fires on origin/main's `_RECONNECT_ATTENTION_AFTER_SECONDS`; 6 advisory hits on
head, all env-only knobs or already keyed per home (agent/redact.py).
2026-09-19 22:34:52 -07:00
teknium1
ece388eca8 fix: test-runner scratch root is /var/tmp/hermes-pytest; two tests follow TMPDIR
~/.cache/hermes-pytest failed seven files: a dot-dir ancestor made the hidden-dir search tests
see every fixture as hidden, and the longer root pushed the PulseAudio/voice AF_UNIX test
sockets past sun_path. /var/tmp is the FHS disk-backed temp root (never tmpfs), non-hidden,
and shorter than the old /tmp root. test_tool_result_storage asserts STORAGE_DIR instead of a
literal, and the zh-Hans bot-mode mirror follows the English code block it must copy.
2026-09-19 10:44:26 -07:00
teknium1
0746903679 fix: test-runner scratch lives in ~/.cache/hermes-pytest; env-less home fallback cannot raise
A runner root under HERMES_HOME/cache made conftest relocate every basetemp out of the live
Hermes home into ~/hermes-pytest-basetemp-*: deeper paths pushed AF_UNIX test sockets past
sun_path, and a profile home under $HOME renders as ~/… (which one test compared verbatim).
The root now sits in $XDG_CACHE_HOME/hermes-pytest (disk-backed, outside the Hermes home,
as short as the old /tmp root) and the test asserts the displayed form.

get_real_home()'s tempfile fallback raised RuntimeError on Windows when a child env carried
no HOME/USERPROFILE; it falls back to the old literal instead of crashing env construction.
2026-09-19 10:44:26 -07:00
teknium1
07df62d604 fix: sockets keep a short temp root; lint skips git-ignored artifacts; test runner scratch leaves /tmp
Chrome puts its SingletonSocket under $TMPDIR and AF_UNIX paths cap at 104/108 bytes, so a
deep HERMES_HOME (profile homes, test homes) made the new scratch TMPDIR kill Chrome at
startup ("Socket path too long") — two browser test files went red on the branch and green
on main. hermes_constants.socket_safe_tmpdir() keeps the scratch root when it fits and falls
back to the OS root for sockets only; the browser env and the code kernel RPC socket use it.

check_no_tmp_literals walked git-ignored runner artifacts (test_durations.json) and failed on
whatever the last test run wrote; it now skips `git ls-files --others --ignored` paths.

run_tests_parallel created its per-file temp roots in the system temp dir and exported no
TMPDIR, so a full-suite run wrote gigabytes of fixtures to tmpfs (3,225 leftover roots, 9.9 GB,
were sitting in /tmp on the dev box). Roots now live under HERMES_HOME/cache/scratch/pytest and
the test process inherits TMPDIR=<root>, so the existing cleanup removes every temp file.
2026-09-19 10:44:26 -07:00
teknium1
4586cde64d chore: mark the deliberate /tmp literals and shrink the lint baseline to one code block
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
2026-09-19 10:44:26 -07:00
teknium1
e16fee4db1 refactor: resolve runtime temp paths via tempfile/TMPDIR instead of literal /tmp
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.

Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
2026-09-19 10:44:26 -07:00
teknium1
3999096d18 ci: forbid literal /tmp paths outside a burn-down baseline
scripts/check_no_tmp_literals.py flags /tmp path tokens in production code, skills,
docs and prompt strings (tests, CI workflows, Dockerfiles, lockfiles, i18n mirror,
code comments and docstrings exempt; ${TMPDIR:-/tmp} idiom exempt). Opt out one line
with 'no-tmp: ok — <why>' on the line or the line above. _BASELINE lists pre-existing
hits per file: growth fails, burn-down is advisory (--strict-baseline / --print-baseline
to refresh). Wired into lint.yml next to check_compat_pointers.
2026-09-19 10:44:26 -07:00
teknium1
d0dbf2cbb6 fix(install): resolve uv shims before salvage and validate the copy in place
Install-Uv accepted any file at $HermesHome\bin\uv.exe, and copied whatever
`Get-Command uv` returned into that location. Chocolatey's bin\uv.exe is a
ShimGen launcher that locates ..\lib\uv\tools\uv.exe RELATIVE to itself, so
the copy is dead on arrival; `& exe --version` does not throw on a nonzero
exit, so the launcher passed the try/catch and the Python stage then failed
with "Python 3.11 not available" (#110350). The re-run path trusted the same
broken copy again.

Building on KoNit-K's Test-ManagedUvBinary and its three call sites:

- Test-ManagedUvBinary merges stderr, relaxes the error preference, and
  returns the `uv <version>` line only on exit 0 -- a launcher's error text
  can no longer surface as "Managed uv found (Cannot find file ...)".
- Resolve-UvShimTarget maps a candidate to the standalone binary before the
  copy: `<name>.shim` sidecar (Scoop), the Chocolatey bin\ -> lib\<pkg>\tools\
  layout, symlinks (winget Links\); other reparse points (WindowsApps
  app-execution aliases) have no copyable file and skip the salvage.
- The salvage rung validates the candidate where it lives, copies, then
  validates the COPY at its new location and removes it on failure, so the
  stage fails honestly instead of reporting success over a dead launcher.
- scripts/tests/test-install-ps1-uv-shim-validation.ps1 drives the real
  Install-Uv with compiled fake uv binaries (a working uv and a
  location-relative launcher) under stubbed installer rungs; wired into
  installer-tests.yml for pwsh 7 and Windows PowerShell 5.1.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-19 02:44:02 -07:00
KoNit-K
3662a1926d fix(install): validate every managed uv acceptance path
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-19 02:44:02 -07:00
KoNit-K
23298685f0 fix(install): reject broken copied uv shims 2026-09-19 02:44:02 -07:00
teknium1
7009611013 fix(desktop): document the afterExtract ordering, trim tests, refresh stale hook references
Follow-up to the salvaged #106846 commit (@JoaoMarcos44):

- after-extract.mjs: spell out WHY the stamp moved (electron-builder's
  beforeCopyExtraFiles rebuilds the PE with resedit for the ELECTRONASAR
  resource; rcedit then cannot commit to that exe, deterministically —
  #105629), and why disableAsarIntegrity was not taken.
- after-extract.test.mjs: two invariants — the hook wiring (afterExtract set,
  afterPack unset, ASAR integrity still on) and the stamp target
  (electron.exe on win32, nothing on other platforms). Red on origin/main.
- set-exe-identity.mjs / scripts/install.ps1: comments still named the
  afterPack hook / after-pack.mjs.
2026-09-19 02:31:27 -07:00
teknium1
c07708671d fix(gateway): every adapter session key goes through one seam (+ lint)
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.

Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.

Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.

Phase 2 of #88715.
2026-09-18 22:04:43 -07:00
teknium1
0ddba07ad7 fix(ci): report an interpreter crash as CRASHED, not "no tests ran"
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.

The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
2026-09-18 19:41:51 -07:00
kshitijk4poor
8925c70a1c docs(whatsapp): group access section says what the gateway admits; env reference rows; trim bridge tests
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
2026-09-19 03:15:40 +05:30
Waldo
81fd9dc773 fix(whatsapp): honor group ingress policy in bridge
Use the configured group policy and group-JID allowlist at Node bridge intake instead of applying the DM sender allowlist to group participants.

Co-authored-by: Martin Gontovnikas <m@gon.to>
2026-09-19 03:15:40 +05:30
jinlingzi-cmd
3926c4209c fix: WhatsApp group messages dropped when LID sender has no lid-mapping (#72529) 2026-09-19 03:15:40 +05:30
Amit C
8a55373dbf fix(whatsapp): authorize first-contact LID senders 2026-09-19 03:15:40 +05:30
teknium1
b1be493b41 Merge origin/main: SDK export list (WORKSPACE_PAGE_HEADER_AREA beside the profile-group exports) 2026-09-18 14:09:20 -07:00
Andrey
b34ebc084a feat(send): add WhatsApp native mentions
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
2026-09-19 00:24:50 +05:30
teknium1
49397cf2b4 fix(tests): parallel runner reports a known flag's missing value as usage, not per file
The bare-flag check asked pytest's parser which tokens it does not know,
but wrapped parse_known_args in the same blanket except that guards
parser construction. A known flag with a missing value (`--tb` alone)
raises pytest.UsageError there, which the except turned into "nothing
unknown", so discovery ran and every per-file pytest died with
"argument --tb: expected one argument".

Keep the fallback around building the parser only; let parse_known_args
run outside it and surface UsageError (and the unknown-token list) as
this runner's own usage error before discovery. One invariant test.
2026-09-18 10:23:21 -07:00
teknium1
d7bedcee1e fix(tests): parallel runner rejects unknown bare flags with usage instead of sweeping
A bare token this runner does not own used to be forwarded to every per-file
pytest, so a typo (`--jbs`, or `--help` before #114065's fix) discovered the
whole suite and each file died with "unrecognized arguments" — hours to learn
about a typo. Validate the bare passthrough tokens against pytest's own
argparse parser (installed plugins loaded), and fail once with this runner's
usage (exit 2) before discovery. argparse handles the attached-short-value
(`-rA`), combined-flag (`-xvs`) and `-k expr` forms, so real pytest flags keep
passing through; tokens after a literal `--` are the caller's explicit choice
and are never validated. If pytest's parser cannot be built, the check is
skipped and behaviour is unchanged.

Follow-up to KoNit-K's `-h`/`--help` interception (#114065). Fixes #114059.
2026-09-18 10:23:21 -07:00
KoNit-K
41332e7851 fix(tests): handle parallel runner help flags 2026-09-18 10:23:21 -07:00
teknium1
3dcf0d49ad fix(install): let Rolldown name the missing binding; repair on every OS
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
2026-09-17 00:25:57 -07:00
Gille
ae9f42accf fix(install): repair missing Rolldown bindings 2026-09-17 00:25:57 -07:00
teknium1
1a4d876176 Merge origin/main: profile-scoped browser env, sudo prompt copy, clone special files
Conflicts (all keep-both): tools/browser_tool.py takes main's scope-bound
passthrough + loopback NO_PROXY and still routes the env through the Bot
Desktop's desktop_env; i18n gains main's sudoDesc/sudoCommandUnavailable
beside our sudoInstallDesc; test_profiles keeps both sides' new tests.
2026-09-16 17:58:54 -07:00
Teknium
73521a8e37 fix(update): one bad workspaces glob no longer aborts the lockfile-churn cleanup
Path.glob raises NotImplementedError for a non-relative pattern, which a string `workspaces` (iterated char by char, so "/") or an absolute entry produces. The (OSError, ValueError, TypeError) catch missed it, so the error escaped to the caller's suppress(Exception) and no lock was reverted at all -- back to autostash every run. Non-list values are now ignored and each pattern is tried on its own so a bad one just owns nothing.

install.sh: read the workspace globs with `while read` instead of an unquoted $(...) so they are never pathname-expanded against the caller's CWD before `case` sees the pattern.
2026-09-16 17:44:36 -07:00