d4e5ecce59cd1886163e976d9539418dd210b8bd
4528 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ed568bd947 | fix(runtime): route dependency hints through pm | ||
|
|
80273b4507 |
refactor(pm): route Python tool installs through PM
Use PM-selected interpreters and tool entrypoints for Browser Use, Hindsight and Python language servers. Sync the declared Google Chat extras instead of changing the active environment with pip. Verified the affected 11-file Nix test subset: 359 passed, 7 skipped. The full suite and real third-party package installation were not run. |
||
|
|
410ac37a0b |
fix(wake): select a supported engine by default
The fixed openWakeWord default selects an unavailable engine on native Windows ARM64 and Intel macOS. Use auto and the existing PM platform gates to prefer openWakeWord, then sherpa, then Porcupine. Keep explicit provider choices unchanged. Porcupine still requires its access key, and wake detection remains disabled until the user enables it. Expose auto in the config UI and document the backend-platform selection. Verified config loading, platform selection, explicit-provider preservation, key requirements, and the config schema. No microphone detection was run. |
||
|
|
284dbaf537 |
fix(pm): isolate bootstrap dependencies and unify YAML on ruamel
Activation reaches plugin discovery before the application dependencies exist. Give PM its own locked Python project and runtime so it can install or repair the application without importing that dependency tree. Keep PM outside the application workspace. A shared uv workspace resolves the application graph and cannot provide this isolation. Route mutations through an isolated worker and preserve transaction callbacks, cancellation, custom package registrations, and correlated receipts. Use the same runtime builder for source installs and packaged payloads. Keep offline wheelhouse support in that builder. Nix builds the independent PM lock as a separate derivation. Refuse lazy-disabled bootstrap before installing tools or dependencies. Move first-party YAML readers and writers to ruamel. Keep the application lock's transitive PyYAML requirements for third-party packages. Verification: - Focused canonical Python suite: 177 passed, 1 host-gated skip. - Electron backend probes: 12 passed. Electron typecheck passed. - Both uv locks, scoped lint, Bash syntax, and whitespace checks passed. - Cold activation, corrupt-app repair, offline staging, and relocation ran. - Built and exercised the Nix PM runtime and standalone YAML merge script. Six broader caller test files retain the same 24 failing test IDs as an archive of HEAD. The existing real-home guard blocks those tests before they can exercise the affected paths. No full-suite pass is claimed. Native Windows signing and full Bionic package execution remain unverified. |
||
|
|
8f6d98e4c3 | fix activation of devenv, use /usr/bin/env bash everywhere | ||
|
|
b3bfc3afe5 |
Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts: # apps/desktop/electron/backend-connection-state.test.ts # apps/desktop/electron/backend-connection-state.ts # apps/desktop/electron/backend-exit.test.ts # apps/desktop/electron/main.ts # apps/desktop/electron/pool-spawn-coordinator.test.ts # apps/desktop/electron/pool-stop.ts # apps/desktop/electron/preload.ts # apps/desktop/src/app/settings/about-settings.tsx # apps/desktop/src/app/updates-overlay.tsx # apps/desktop/src/global.d.ts # apps/desktop/src/store/notifications.ts # apps/desktop/src/store/updates.ts # gateway/config_loader.py # hermes_cli/banner.py # plugins/platforms/dingtalk/adapter.py # tests/hermes_cli/test_plugins_cmd.py # tests/test_live_system_guard.py # tui_gateway/server.py # website/docs/user-guide/desktop.md |
||
|
|
a6d65cdd09 |
fix(state): single durable-shape authority for async_delegations
Review follow-up on #94701: the delegation tool's _initialize_schema still carried its own CREATE TABLE + ALTER column list for async_delegations, leaving a second durable-shape authority even with the column declared in SCHEMA_SQL. Its legacy ALTER added origin_session_id as bare TEXT (nullable, no default); reconciliation repairs missing column names only, so a database first opened through the tool kept a non-canonical shape forever (#94691). Remove the private DDL entirely. The tool's initializer now calls a new reconcile_state_schema() in hermes_state_schema, which replays the canonical SCHEMA_SQL (idempotent CREATE IF NOT EXISTS for every table, canonical indexes included) and reuses SessionDB's declarative _reconcile_columns for missing-column backfill — one reconciliation implementation, one authority. Because _parse_schema_columns reconstructs each column's full constraint expression (type, NOT NULL, DEFAULT), the tool-first legacy path now adds origin_session_id as TEXT NOT NULL DEFAULT '' — the canonical shape — and SQLite backfills existing rows with the '' default. Opening-order regressions compare FULL PRAGMA table_info metadata (type, notnull, dflt_value, pk) plus the canonical index set across fresh SessionDB→tool, legacy→SessionDB, and legacy→tool→SessionDB, each preserving a pre-existing legacy delegation row. |
||
|
|
df0eed4f6b |
fix(session_search): a bare session id never reads another profile's state.db
Reading a session by id that missed the caller's store fell through to _locate_session_db(), which opened every profile's state.db read-only and returned the first owner's full transcript — no opt-in, no profile named, and the miss path even fired after an explicit non-matching profile= read. Any caller holding an id (ids appear in logs and tool output) could read a foreign profile's conversation. Profiles are isolated islands by design. A miss now stays a miss, with a hint to name the owning profile (profile=<name> / @session:<profile>/<id>), which remains the sanctioned, explicit cross-profile read. The schema eval runner no longer needs to fake the scan. Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779. |
||
|
|
0e927c914d |
Guided first launch behind HERMES_GUEST_ONBOARDING: intro, guided chat, first task in default (NS-848, PR B1) (#107958)
* feat(desktop): port guided onboarding substrate Add seeded session creation, transcript directives, profile routing, and the shared window and pane primitives needed by the guided flow. Keep later-step mounts deferred and exclude provider selection and retry machinery. * refactor(desktop): anti-slop cleanup for substrate Assemble seed parameters in the existing create helper and use the owning transcript attribute type. Read the guaranteed gateway and connection contracts directly to remove runtime type probes and unchecked assertions. * test(desktop): create-overrides invariants Verify that reasoning and title overrides do not select a provider or model. Empty overrides and seeds add no parameters. * feat(desktop): port first-run cinematic window Play the cinematic behind the guest onboarding launch flag using bundled Collapse and JetBrains Mono. Give the native window its own controller and restore the app on skip, renderer deadman or native watchdog. Drop the perf scenario because it depends on the removed replay hook. Guided chat kickoff and app-shell gate wiring remain with their later steps. * refactor(desktop): anti-slop cleanup for cinematic Preserve audio and canvas behavior through named types and inferred results. Split the viewport node and frame drawing to keep control flow bounded. Cut comments that only repeat the code. * feat(desktop): add onboarding gate and answers stores Track cinematic, guided chat, handoff and completion in one phase record. Queue the guide after the intro and share pending kickoff work between callers. Keep existing saved answers while dropping retired preferences. Leave intro seen-state ownership with the cinematic store. * feat(desktop): port guided onboarding chat Add guided setup cards, runbooks, machine context, and onboarding presence. Connect transcript rendering and first-build progress to the desktop behind the onboarding flag. Leave session kickoff and handoff execution for the next step. * refactor(desktop): anti-slop cleanup for guided chat Keep directive and layout lookups typed. Remove unsafe test casts and isolate onboarding transcript calculations without changing the flow. * feat(desktop): connect guided onboarding to durable first-build handoff Start the guide only after its profile backend confirms bootstrap readiness. Seed or adopt the welcome chat, then transfer the first build to default with a durable receipt and explicit retry. Wire cinematic completion, screen stand-down, layout growth and progress check-ins. Save agreed preferences before creating the build and release prompt slots after storage refusal. * refactor(desktop): anti-slop cleanup for onboarding handoff Reuse the gateway request and error contracts. Isolate guide adoption and snapshot validation while preserving receipt recovery and reasoning overrides. Validate persisted receipt fields at the JSON boundary without coercion. Keep corrupt identities rejected and retain only the permitted test mocks. * fix(desktop): guided chat review fixes Wire the native machine probe so guided setup can suggest a name and offer the right first task. Restore the comments that explain the flow boundaries. The directive registration uses the launch flag to preserve ordinary chat. Ruling 6 folds active.ts into assembly to keep activity ownership together and removes the second greeting source so the seeded and visible greetings agree. * fix(desktop): handoff review fixes Probe the guide backend before switching profiles so a readiness refusal keeps classic onboarding on the current backend. Restore list-valued personalization coverage and routing rationale. Remove the obsolete setup status fixture. * chore(desktop): onboarding script cull and rehearsal recipe Document a temporary-state rehearsal using the existing onboarding flag and optional portal stand-in. Keep the main scripts unchanged and retain window growth for the guided chat. * fix(connectors): reject incomplete catalog responses * feat(gateway): scope connector controls to the owning session * feat(desktop): connect apps through native session-owned controls * feat(desktop): gate connector cards and enable free-tier access Use the launch flag before mounting connector controls so classic transcripts add no status requests. Allow existing free-tier identities through the read-only tool gateway gate and test the owning-profile RPC path with A’s launch gate. Keep authorization links out of previews. * style(desktop): format connector translations Apply Prettier to the connector copy blocks while preserving upstream translations and free-tier wording. * refactor(desktop): anti-slop cleanup for connector card Use the transcript JSON contract and concrete RPC parameters. Preserve malformed-value filtering at one string boundary and make the fixture and row types explicit. Keep connector execution and cancellation behavior unchanged. * feat(desktop): detect initial language from the OS Use the native machine locale when no supported language is saved. Preserve explicit choices and leave inferred languages out of config. * refactor(desktop): anti-slop cleanup for initial locale detection Keep unvalidated config values at the existing validation boundary. Pass no saved choice after that boundary has ruled it out, preserving locale precedence. * test(desktop): onboarding port test set Make native window tests reject duplicate IPC handlers and isolate disabled onboarding. Assert the active gate mock when onboarding re-enables. Keep the test set limited to behavior carried by the port. * fix(desktop): recover failed guide kickoff and reveal once The review found that a failed guide create stranded the solo shell and draft profile, and solo boot faded an already visible window a second time. Restore the prior route and layout, release onboarding through its existing phase record, and surface create failures. Let the film own the reveal while solo boot animates the visible resize. * fix(desktop): preserve transcript ownership across cards and handoff The review reproduced answers submitted to the focused chat, repeated questions disabled across sessions, handoff recovery using foreground identity, and mount-dependent progress history. Target each card’s own composer, scope settlement to its message and session, carry the issuing guide through handoff, and derive progress from its transcript with streaming activity. Reuse the existing owner ladder for exact and profile-only routes. * fix(gateway): preserve connector ownership with profile routing The review found that shared-primary profile metadata was rejected before connector dispatch, while desktop controls treated a missing registry id as missing ownership. Accept profile only as routing metadata and keep the live transport as authorization. Resolve card ownership through the existing exact/profile ladder, retaining ambient routing only for the single-backend case. * fix(desktop): resolve plugin roots and gate the Basic layout The review found that the first plugin build was seeded with a different installation’s fixed path, and the director ruled that flag-off layouts must match main. Resolve the running desktop’s plugin root before seeding a plugin build and register Basic only when onboarding is enabled. Keep the runbook wording and the ordinary four layout presets intact. * fix(desktop): clear review-fix slop findings The slop gate flagged an undocumented layout-data assertion and unknown-return types in the new test selectors. Record the layout registry invariant and preserve each selector’s return type. The only remaining production finding is the accepted connector-tools baseline. * fix(desktop): detect the OS language on a fresh install The review found that the merged English config default prevented the desktop from probing the OS language on a fresh install. Add an opt-in saved-values read so an absent choice remains distinct from saved English. Preserve default-valued English only for explicit language saves; unrelated settings saves must not turn a merged default into a language choice. Older backends ignore the new query options and keep returning merged English, preserving their existing desktop behavior. * test(desktop): make the flag-off layout registry test deterministic The flag-off test awaited the full controller import, pulling in the UI graph and installing application watchers just to read layout presets. That import took 9.5 seconds locally and timed out in the director's run. Move the existing trees and registration into a small layout-presets module. Production and the synchronous test use the same flag-gated registration, without starting the controller in the test. Keep the real registry invariant and dispose the test's contributions after completion. * fix(desktop): keep the transcript parser and ::ask behind the onboarding flag Register the guided chat's question card only with onboarding enabled. Restore main's whole-paragraph parser and contribution rendering when the flag is off, including its streaming prose behavior. Keep segmentation for the guided flow until B4 decides the parser's wider use. Restore main's two parser test files so its existing product and plugin contracts remain the flag-off check. * test: drop the onboarding and connector tests pending a later ticket Apply the director's ruling to remove B1's added test files and restore main's existing suites. Keep only the gateway route-reader mock contract that main's profile tests need against the shipped activation behavior; their cases and assertions stay intact. The flow's shape is not settled and B3/B4 rewrite it. The connector layer will also be reworked. The live CDP run is the flow check until a follow-up ticket brings tests back. --------- Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com> |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
4fc64b6730 |
refactor(terminal): probe NOPASSWD only when a prompt would fire; drop the host-only probe
With BaseEnvironment supplying a backend-scoped probe to every production caller, the module-level host-only `_sudo_nopasswd_works` (and its TERMINAL_ENV gate) had no callers left; the `or` fallback and the outer try/except around a callback that already fails closed were dead too. Move the probe under `should_prompt_for_sudo`: headless callers (gateway, cron, delegated children) reach `(command, None)` whether or not the probe runs, so the extra backend round trip — an ssh exec on the SSH backend — was pure waste on every headless sudo command. test_subagent_sudo_prompt no longer needs to patch the host probe out: bare `_transform_sudo_command(cmd)` calls have no probe by construction. |
||
|
|
39abca492d |
fix(terminal): gate sudo probes by backend cancellation safety
Opt in only Local, Docker, SSH and Singularity: a timed-out `sudo -n true` probe on those backends kills one process, while SDK adapters (Modal, Daytona, Vercel) cancel by terminating the whole sandbox. Net of the original PR's commits f7db10ef + 9965b1b9 (the intermediate _ThreadedProcessHandle special-case was superseded by this gate). |
||
|
|
d38064e3c1 | fix(terminal): probe remote passwordless sudo | ||
|
|
45a6101f36 |
fix(gateway): secondary-profile send_message, notices and /loop wakeups go out via their own bot
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved its live adapter by bare platform from runner.adapters — the DEFAULT profile's map — so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy, every plugin platform), its "Gateway shutting down/restarted" and /update notices, its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left through the default bot (Telegram DMs landed in the user's chat with the other bot). Every such door now resolves through the profile-aware, fail-closed resolver already used by the inbound reply path (authz_mixin: _adapters_for_profile / _authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear error, never the default bot. - tools/send_message_senders.py::_live_adapter — resolve via runner._authorization_adapter(platform, get_active_profile_name()); shared by _send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and the WeCom standalone sender. - gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay- aware resolve_delivery_transport callers); _authorization_adapter reuses it. - gateway/run_shutdown.py — shutdown/restart notice for a running session uses the session's source transport / agent:<profile>: key lane, never self.adapters. - gateway/slash_commands.py + run_notifications.py — /restart and /update markers persist `profile`; the restart notice, update result and update prompt resolve the requester's own adapter (legacy markers fall back to the session_key lane). /goal, /heartbeat, /approve, /deny confirmations use the source's own transport. - gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its route; the wakeup watcher scans every served profile's store under its own scope (same shape as _handoff_watcher) and fires through that profile's adapter map. - plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in the owning profile (its adapters and its home channels). - plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter. - docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity). Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py, tests/gateway/test_multiplex_notice_egress_profile_adapter.py. Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com> |
||
|
|
580322ef1e |
fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for every name tagged in the process-global `_SECRET_SOURCES` map. That map is filled by EVERY served profile's secret-source hydration, while `os.environ` only ever holds the LAUNCH (default) profile's values — so once any profile's 1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio MCP server was started with the default profile's token. Resolve those names through the active profile's secret scope (`get_secret`) instead: the routed profile's value, or omitted when that profile has none. Under multiplex `get_secret` never falls through to environ; single-profile runs keep the .env overlay + environ behaviour, so the existing "vault vars reach MCP subprocesses" contract still holds there. `secret_source_names()` exposes the tagged NAMES only — values are never read from the shared map. Docs: the multi-profile guide's "MCP subprocesses only see their own profile's secrets" claim is now true for source-injected names too; say so explicitly. |
||
|
|
4bdd64b334 |
The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cbcf7b72f7 |
feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop * feat(gateway): /signin signs the free tier into a Nous account from a DM * feat(cli): chat surfaces name /signin as the sign-in verb * fix(auth): review follow-ups for the shared sign-in flow and /signin * fix(i18n): carry the /status free-tier line in every locale catalog * refactor(cli): the chat sign-in command is /login * fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules |
||
|
|
a2db110ccc |
feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous provider) instead of forcing the setup wizard. The identity is persisted through the same path a real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off. * test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin * docs(user-guide): free tier and signing in New page explaining what a fresh install gets before any key or sign-in (free inference on nous/welcome plus connectors), how the free tier coexists with a user's own API key, how to sign in with hermes auth upgrade and keep connectors, how to turn the free tier off with nous.guest, what hermes logout does in each state, a troubleshooting table, and a plain privacy note. Wired into the Using Hermes sidebar. * fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user is told they were never signed in. Logging out of a real Nous account now also clears the cross-profile store, so a profile logout is not silently re-adopted on the next boot. * fix(model): switching off the free tier points at signing in, never hops providers * Name the free tier in the gateway startup notice and tell explicit-provider installs about it once * Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token * fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free tier before declaring nothing configured. On a fresh install the first command lands in chat on nous/welcome; a failed setup still falls through to the existing guidance. * Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors The device-code flow runs as usual, with a promotion intent registered on the portal between the code request and the token poll so the account that approves the code inherits the free tier's connectors. The promotion status decides the outcome: only a completed one is followed by the token grant, which is persisted over the free-tier singleton and the shared store. Declined, superseded, retired and busy outcomes each print their own plain copy, and a retired identity is cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals. * Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off * fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the claim code, not the generic device page. A failed mint is attempted once per process so several bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check. * fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure A credential-pool entry can select a paid Nous key while the profile singleton is still the free tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and the pin in model normalization is removed since it had no route to look at. A failed background identity setup releases its latch so a later attempt in the same process can try again. * fix(auth): decide the Nous model together with the route on every credential-pool swap The credential pool can move a Nous agent between the welcome host and the portal host after init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model. * fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start * fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the shared account. Locks are taken in the documented order (profile, then shared). A minted credential is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and trigger a second mint. Retiring a dead credential removes only that credential from both stores. Guest exchange uses the resolver's canonical portal URL. * fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential The welcome host serves one model, so a rotation onto it is refused for any conversation on another model instead of silently switching that conversation to nous/welcome (the model pin applies only when a route is first chosen). The connector token path now treats the free tier as absent when nous.guest is false, including cached tokens, and shares the one dead-credential rule with inference: a retired identity is replaced once rather than returning its stale token. * fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only A free-tier identity in the shared store is not an OAuth credential to offer for import; a real sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted state (no token refresh at boot), so an expired free-tier token cannot stall the online message. |
||
|
|
6756d11b5f |
fix(pm): ship full chromium without headless shell
Full Chromium serves both headed and headless sessions. The separate shell duplicates the browser payload and is not needed for either mode. Remove the shell from PM and Docker. Select the managed Chromium executable for agent-browser and the full Chromium channel for direct Playwright callers. Route setup through PM and remove retired packages from cached bundle stores without changing the user's tool store. Update signing, architecture checks, launch probes and install guidance. Leave llama packages and Docker archive cleanup unchanged. Verification: - Real agent-browser navigation, clicks, DOM reads and screenshots pass in headed and headless modes with the same Chromium executable. - The direct Playwright doctor probe passes. - Focused Python and desktop packaging tests pass, as do six Docker checks and both real-browser task-scroll tests. - The built linux/amd64 image is 1.393 GB compressed, 223.6 MB smaller. - The broader PM suite and two unrelated setup tests still fail. Those failures reproduce on unchanged HEAD. - Five updated eval scripts parse; their full scenarios were not run. |
||
|
|
d9ca9c974d |
feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped the agent cold: the login classifier excludes one-time-code fields on purpose (a password must never land in an OTP box) and there was no tool for the second step, so the only move was to ask in chat. browser_vault_enter_code Fills the one-time code the current page asks for. Two sources, same invariant as passwords (the code goes to the page over the supervisor socket and never enters model context): - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib, verified against the RFC test vectors), 1Password `op item get --otp`, Bitwarden `bw get totp`. Nobody is asked. - no seed: the surface prompts "Verification code for {site}"; the user types what their phone/email/app shows. Enter on empty / Skip declines and the tool returns code_declined ("do not ask again this turn"). no_code_field tells the model the site wants a passkey / hardware key / app approval: hand it to the user's device and wait for navigation. Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order. Surfaces CLI: sudo-style panel, code shown as typed (not a secret worth masking, typos must be visible), Enter submits, ESC/empty skips. Desktop: "Verification code for {site}" card via vault.code.request / vault.code.respond (gateway), owner-routed like the other vault prompts. Settings → Passwords & Logins: optional "Authenticator key" field on the add form (base32 or otpauth:// link); items with one show a "2FA auto" badge. `hermes vault add` asks for the same optional key. browser_vault_fill's result now says what to do next ("if the site asks for a verification code, call browser_vault_enter_code with this handle"). Six locales. Verified live (real model, local 2FA site that checks the TOTP; CLI PTY): A. login saved with authenticator key → signed in through 2FA, zero prompts, code/password absent from the transcript B. login without key → code panel → user types code → signed in C. panel dismissed → agent stops and explains, never asks in chat Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip). |
||
|
|
96fbc47f14 |
fix(vault): dogfood fixes — offer save-login on every backend, keep the model off passwords, bind to the login tab
Found by using the feature as a user (natural prompts, real sites, CLI PTY + native Electron), not by naming tools:
- browser_vault_save_login was registered but never offered: toolsets.py is a hand-maintained list. Added, with an
invariant test that every registered browser_vault_* tool is in the browser toolset.
- Vault tools were absent on the DEFAULT backend (Browser Use): the gate deferred to check_browser_requirements(),
which is False by design there. Gate = is_browser_use_cli_mode() or check_browser_requirements().
- The model typed a page-shown demo password with browser_type and offered to take one in chat: the vault rules
lived only on the vault tools. browser_type/browser_exec now carry a vault note when the vault tools are
present ("call browser_vault_list first … never type a password with this tool, never accept one in chat, even
if the page shows it"); the browser_exec login-wall line points at the vault instead of "ask the user".
- On Browser Use the saved item was bound to chrome://new-tab-page: the supervisor's default page session is the
daemon's blank tab. browser_vault_save_login now focuses the tab holding a password field before reading its
origin (focus_page("", accept=probe); about:/chrome: pages are never candidates). Live E2E leg added.
- Settings row: "identifier · Added <date>", origin omitted when it duplicates the label.
Live (real model): CLI on Browser Use — first visit prompts, signs in, saves; second visit fills silently; GitHub
decline (Enter or ESC) stops the agent, which refuses chat passwords. CLI on the built-in stack — same three
scenarios pass. Desktop native Electron — same three scenarios plus Settings list/remove pass. Password never in
a transcript, UI, or a file outside vault/.
|
||
|
|
98d11c95f4 |
feat(vault): zero-setup UX — save a login on the page that needs it, managers auto-detected, one "Passwords & Logins" surface
Nobody should have to learn `hermes vault add` or find a toggle before "log into GitHub" works. - browser_vault_save_login: when the agent reaches a sign-in page with no saved login it asks the user on THEIR surface (CLI two-step panel on the sudo modal: identifier shown, password masked; Desktop card with labelled Email/username + Password fields). The answer goes to the encrypted vault bound to the page origin and is filled at once; the model gets back only the handle and identifier. Declining returns save_declined; headless sessions get prompt_unavailable. Never a password in chat. - Vault tools ride with the browser toolset (check_browser_requirements) instead of appearing only once the vault has items — an empty vault is exactly when save_login is needed. browser_vault_list hints at it when empty. - 1Password / Bitwarden are login sources as soon as their CLI is installed; `vault.<name>.enabled` is opt-OUT only. Settings shows Detected/Locked/Unlocked/Off/Not detected with a switch only for installed managers; `hermes vault sources` reports detection, `--disable`/`--enable` flip the opt-out. - Desktop nav/page renamed "Passwords & Logins"; empty state tells the user they do not need to add anything; all five locales updated. Docs rewritten from "how it works" to "say log into X". - New per-thread SaveLoginPrompt callback (agent/vault_backends/unlock.py) installed beside the unlock prompt on every CLI site and the gateway bridge (vault.save_login.request/respond/expire), propagated to worker threads via tools.thread_context. Live: CLI PTY (real model, packaged Chromium, local login server) — panel shown, identifier + masked password typed, server received the correct password, password absent from terminal transcript and from every file under HERMES_HOME outside vault/. Native Electron (headless, isolated HOME/HERMES_HOME, own Vite + CDP port) — card shown, "Save & sign in", server received the password, Settings lists the saved item, password absent from the rendered UI. |
||
|
|
f5e28f8119 |
fix(vault): declare tool parameters the registry understands; attach a supervisor to local built-in sessions
Found by a model-driven live run (hermes chat -q against a real login page): the agent found the vault tools, typed the identifier, then failed twice for reasons the direct-call E2E could not see. - The three schemas spelled their arguments `input_schema` (Anthropic shape). The registry and every provider adapter read `parameters`, so the model was shown browser_vault_fill with NO arguments and called it with an empty handle. Renamed; a registry-wide invariant test now fails on any schema without `parameters`. - A local built-in session (agent-browser --session) carries no cdp_url, so nothing ever started a supervisor for it and the fill refused with supervisor_required. _ensure_supervisor asks the daemon for the packaged Chromium's endpoint (`get cdp-url`, carries no secret) and attaches on demand; the fail-closed test now pins that the daemon is only ever asked `get`, never handed an eval with the password. Live: same run after the fix -> browser_navigate, browser_vault_list, browser_type, browser_vault_fill, browser_click; the test server received the correct password; landing title "Welcome"; the password string is absent from the whole transcript. |
||
|
|
c8998c3957 |
feat(vault): fill payment cards and addresses at checkout, cards behind a confirm prompt
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could fill them: a dead surface holding real card numbers. browser_vault_fill now handles all three kinds through the same origin-bound, supervisor-only, redacted path: - classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens (cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code, country-name) with label/name heuristics as backup; a combined "MM/YY" control gets exp_month+exp_year and suppresses the split fills; inspection now covers <select> (country, state, expiry month) and the fill script picks an option by value or text. - Every payment fill goes through request_elicitation_consent (gateway button round-trip or CLI panel) before a byte is written; declined → payment_declined, headless sessions are refused. A prompt injection that reaches a checkout can ask, not spend. Card values join the redaction registry like passwords; the result lists targeted field tokens only. - Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the field) because a card without a bound origin is unfillable. - The tool descriptions, docs and CLI copy drop "Phase 1 / login only". Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the country <select> untouched, and neither the card number nor the CVC appears in any result. browser_vault_tool also: focuses the tab on the bound origin holding the right form before the origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's input tool" (rewritten per session by model_tools); _check_vault_available is registered uncached because its answer is per profile (vault dir + config) and the probe is a file stat. |
||
|
|
1bd4e33d36 |
fix(vault): make browser_vault_fill work on the default Browser Use backend
On the default backend (browser.backend unset → browser_exec) the vault tools were advertised but could never fill: the CDP supervisor that carries the secret-bearing eval is started only by the built-in browser_* session path, so _eval_js_secret failed closed with supervisor_required and the origin pre-check fell back to an agent-browser CLI eval against a browser browser_exec never touched. - browser_exec now attaches SUPERVISOR_REGISTRY to the CDP endpoint it just routed the harness to (BU_CDP_WS/BU_CDP_URL), so the fill talks to the SAME browser over the same secret-capable WebSocket. BU direct-cloud (BU_AUTOSPAWN) exposes no endpoint and keeps the supervisor_required refusal. - CDPSupervisor.focus_page(origin, accept=<js>) (used by browser_vault_fill, next commits) re-attaches the page session to the open tab on the item's origin whose DOM holds the form being filled (browser_exec opens its own tabs; the supervisor's initial attach picks the first page target, which is chrome://new-tab-page). browser_vault_fill uses it before the origin pre-check with a per-kind probe (password input / card fields / address fields). Live: evals/vault_fill_live_e2e.py drives the real browser_exec tool against Hermes' packaged Chromium with the login page in the third tab; A/B with the attach line disabled fails at "did not attach a supervisor", enabled fills the password into the /login tab and card fields into the /checkout tab with every model-facing read scrubbed. Also: browser_vault_list/fill described the workflow as "type the identifier with fill_input", a helper that exists only inside browser_exec code (toolset browser-use) and is a ghost on the built-in stack. model_tools._rewrite_browser_vault substitutes the concrete name from the session's actual tool set (`fill_input` inside browser_exec, or browser_type), the same dynamic cross-reference pattern browser_navigate uses for web_search. |
||
|
|
4140901d15 |
fix(browser): stop refusing credential-named query params on cloud browser/extract backends
browser_navigate / browser_exec / web_extract refused any URL whose query carried a
credential-NAMED parameter (token, signature, access_token, ...) when the backend was
a cloud provider. That is exactly the shape of magic links, OAuth callbacks and signed
CDN assets, so on Browserbase/Browser Use the agent could not finish a sign-in flow or
open an X video asset ("Blocked: URL contains a credential-like query parameter").
The floor protected nothing: the cloud browser already sees every cookie and typed
password of the session, and with the credential vault it receives the real password at
fill time. Hermes' own secrets leaking into a URL stay blocked by the value-shaped
_PREFIX_RE check (_secret_url_error), which is backend-independent. IMDS and
private-address floors are unchanged.
|
||
|
|
dedc99ec6c |
fix(vault): ownership and race findings from the second independent review
Manager tokens: lock generation fence (a Lock acknowledged while `bw unlock` / `op signin` is still running discards the late token); tokens record the unlocking gateway session and are released when THAT session ends, not when any sibling session in the profile is torn down. 1Password: OP_CONNECT_HOST/TOKEN come from the profile's scoped secret store like the service token (Connect outranks a service token inside op), never from the launch environment. Vault RPCs bind params.profile (home + secret scope) so a shared remote backend serving several profiles locks/lists/unlocks the requested one; unknown profile → RPC error, not a crash. Fill target: inspection stamps are `<nonce>:<index>`; a fill resolves only its own inspection's stamps, so an interleaved second inspection can no longer redirect A's password into a newly mounted field (real Chrome: 0 filled, both fields empty). Desktop Settings: every RPC goes through the owner profile's socket (requestGatewayForProfile), query keys carry (connection, profile), an owner change closes dialogs and wipes drafts (a master password typed for A is never submitted to B; a late list from A never paints under B), and vault.add secrets travel in a ref consumed by the mutationFn instead of mutation variables. Three owner-routing invariant tests on the real component. Docs/PR body: session-scoped release, lock-race semantics, bw --passwordenv. |
||
|
|
5aa9e7f033 |
fix(vault): independent-review findings — vendor contracts, profile scope, transport, target binding
Bitwarden unlock now uses the CLI's documented non-interactive channel: `bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child environment only (bw 2026.x rejects a piped password with "Master password is required"). Verified against the real published binary. Manager session tokens are keyed by (profile home, backend): a Desktop gateway hosting several profiles can no longer reuse or lock another profile's session. Status probes (`vault.sources`, is_unlocked) no longer refresh the idle TTL; only real manager calls do. Gateway session teardown locks the profile's managers (a per-session unlock ends with the session). 1Password service-account token comes from the profile-scoped secret store (get_secret), not ambient os.environ. `vault.source.set` no longer references a module constant (bind_module rebinding dropped it → NameError on every Settings toggle). Fill target binding: inspection stamps each input with a per-inspection slot attribute; the fill resolves by stamp and requires type=password, then strips every stamp. A DOM reflow between inspect and fill can no longer redirect the password into a text field (reproduced in real Chrome before, 0 filled after). Redaction boundary: no 4-char floor, CR/LF-normalized form registered (what a text input actually stores), JSON object KEYS scrubbed in both browser redactors; longest value first. Docs now state the real trust model: accidental-disclosure protection, not an execution sandbox. Desktop: the mid-turn card sends the master password through the owning session's socket (requestForOwnedSession), never the ambient foreground gateway; `vault.unlock.expire` clears a stale card; Settings keeps the master password out of react-query mutation variables (ref consumed by the mutationFn). One renderer invariant test for the routing. |
||
|
|
b51da65258 |
fix: adapt execute_code cell authority to the widened prompt-callback table
_callback_api() now yields (getter, setter) pairs for every per-thread prompt (approval, sudo, vault unlock); the kernel cell captured and restored the old fixed 4-tuple. Iterate the table so a cell carries every callback and a future addition needs no change here. Test recorder unpacks the new shape. Also: perfectionist import order in ui-tui interfaces.ts (CI lint). |
||
|
|
e9f7d44f3f |
fix(vault): carry the unlock prompt onto tool worker threads; honour binary_path in install checks
Live Desktop repro: the fixture model called browser_vault_unlock and got unlock_unavailable although the renderer was interactive. tool_executor runs handlers on a propagated worker thread; thread_context only copied the approval and sudo thread-local callbacks, so the unlock prompt registered by _wire_callbacks was invisible there and can_prompt_here() said nobody could answer. The callback table in thread_context now lists every per-thread prompt (approval, sudo, vault unlock) so a new one cannot silently drop off worker threads again. After the fix the same turn shows the masked card and completes. vault.sources / `hermes vault sources` report a manager as installed when its configured binary_path exists, not only when it is on PATH. |
||
|
|
8e7e9be217 |
feat(vault): sign in with 1Password or Bitwarden logins, unlocked per session
The browser vault now draws from three login sources behind one handle
shape: the local encrypted vault (vault_…), 1Password Login items (op:…)
and Bitwarden Password Manager logins (bw:…). browser_vault_list aggregates
metadata across them; browser_vault_fill routes by prefix and resolves the
password at fill time only, through the manager CLI.
External managers are locked until the user unlocks them for the current
session. The new browser_vault_unlock tool (and the fill path, implicitly)
asks the surface to show a masked master-password prompt — CLI panel
(reuses the sudo panel state), TUI/Desktop via a vault.unlock.request
blocking card. The password goes to `op signin --raw` / `bw unlock --raw`
on stdin, never argv or env; only the session token is kept, in memory,
with a 30-minute idle TTL, cleared on session close or `vault.lock`.
Headless contexts (cron, webhook, api_server, -q) can never prompt: the
manager is reported as locked with unlock=unavailable_in_this_session and
fill refuses — the same posture approvals take where nobody can answer.
Config: vault.onepassword / vault.bitwarden {enabled, binary_path, …};
a 1Password service-account token skips the prompt for headless use.
RPC: vault.sources, vault.source.set, vault.unlock, vault.lock for Settings.
Tests (2, real subprocess against a fake bw; each proven red by sabotage):
headless never prompts or spawns; unlock feeds stdin only, token never
enters os.environ, fill routes by prefix and the password only reaches the
fill script.
|
||
|
|
cde0ad0edd |
feat: agent signs into sites from an encrypted local vault (CLI, browser fill, Desktop Settings)
Consolidated re-apply of #96988 onto current main. Ported from Merit-Systems/OpenInstinct (MIT) opaque-handle autofill design: the model sees vault handles + login metadata, the password is resolved and filled server-side over the supervised CDP socket, and filled values are scrubbed from every browser tool result by an unconditional redaction registry. Rebase adaptations to the Sep-2026 facade/sibling layout: - toolsets: one _HERMES_CORE_TOOLS entry (the browser toolset derives from it) - hermes_cli/main.py: vault parser registered via the subcommand owner table - file_safety: vault/ joins the _READ_DENIED_DIRS credential-dir table - redact: registry scrub runs before the redact_secrets early-return - browser_vault_tool: _run_browser_command now lives in browser_tool_session |
||
|
|
4f12985cd2 |
fix(gateway): refuse relay sender fields from a logged-in client and say what the author trusts
bot_relay.deliver accepted from_profile, from_handle and from_connection from any admitted JSON-RPC client. The handler now refuses them with error 4095 when the calling transport carries a browser login identity, since a logged-in browser never relays for another connection. The DeliveryAuthor docstring now says the author is trusted because an admitted client relays it, not because the sender is verified. |
||
|
|
091ac34ba9 |
fix(bot-mode): qualify the peer-dm author id with the sender's hostname
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id. |
||
|
|
3db4defcc1 |
fix(bot-mode): qualify a relayed author from the local connection too
delivery_turn_author kept the bare bot:<profile> id when the sender's connection was the Desktop's own "local", so a DM relayed from that machine collided with the recipient's profile of the same name. A relayed DM always crosses gateways, so the connection id is now part of the id whenever the Desktop sends one, and only the direct message_agent path in tools/bot_mode_dm.py stays bare. The relay.ts and session_auto_continue.py comments added earlier are cut to one line each. |
||
|
|
d3c8bacbe3 |
fix(bot-mode): qualify relayed authors with the sender's connection id
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
|
||
|
|
6881e4d3fc |
fix(bot-mode): carry the relay sender into a live Bot Chat turn
When the target Bot Chat is already open on this gateway, the relay handler delivers through `prompt.submit` with `queued: true`, and that branch dropped the envelope's sender. The model still saw the text prefix, but the turn reached the agent unattributed, the exact case the subprocess branch fixes. The relay handler now stamps the author on the submit as a `DeliveryAuthor`, an in-process object a JSON client cannot build, so `prompt.submit` accepts it the way it accepts a hosted-room callback and refuses a dict with error 4124. The busy queue keeps an authored envelope in its own slot, the drain hands the author to the turn runner, and the runner passes it to an agent that declares the keyword. A plain prompt after an authored dm carries no author. Local deliveries to a desktop-owned Bot Chat take the live-owner mailbox instead. The admission intent and the mailbox record now carry the author, a retry under the same id with a different author is refused, and the owner gateway hands the author to the turn it runs. Isolated compute turns still run unattributed, because the compute-host frame has no author field. |
||
|
|
16d15869ba |
feat(bot-mode): carry the sender through local and relay deliveries
a bot dm arrived as an ordinary user message. the only trace of the sender
was the "Message from" text prefix, which the model reads and nothing else
does. the recipient's memory provider saw its own configured user.
message_agent now passes the sender as {"id": "bot:<profile>", "name":
<handle>, "is_bot": true} to the delivery runner (--author <json>), which sets
HERMES_TURN_AUTHOR on the recipient one-shot only. the -Q turn reads it and
passes turn_author into run_conversation. the desktop relay forwards the
envelope's from_profile/from_handle to bot_relay.deliver, which sets the same
variable on its delivery turn. the runner drops any inherited author first so
a delivery without one stays unattributed. the text prefix is unchanged.
|
||
|
|
ac07e20407 |
fix: resolve subagent control authority from the live session slot
Subagent list/tail/steer/interrupt authorized against a per-record copy of the owning session's transport (`owner_transport`). That copy had to be re-synced at every reattach site; `_rebind_live_transport` did it for session.resume/activate but prompt.submit and the queued-prompt drain still attached bare, so a client that reconnected through a prompt (the common path on a remote gateway / Bot Mode switch) streamed fine while `subagent.list` returned [] and controls rejected. Read `owner_session_record["transport"]` at check time instead: the slot is already mutated by every attach/detach/viewer-failover path, so no site can forget the sync. `owner_transport` stays as the capture-time "commissioned by a gateway session" marker (None = no RPC authority ever); non-dict owners keep the exact-object rule. Drops the registration-time re-read and the attach-time registry loop. Diagnosis credit: nftpoetrist (#106663) — their prompt.submit / drain regression tests pass against this change with no call-site edits. |
||
|
|
b1915cb02d |
fix(bot-mode): keep the relay waiter watching past the Desktop deliver deadline
The sender-side waiter gave up at 900s while the Desktop held bot_relay.deliver open for 1500s, so a turn finishing between minute 15 and minute 25 wrote a reply nobody read. REPLY_WAIT_SECONDS now rebuilds the Desktop budget from the same numbers and waits 60s past it. The two turn constants move into tools/bot_relay.py so the gateway handler and the waiter share one definition. |
||
|
|
49ef015ca3 | fix: join delegated work before finite chat exits | ||
|
|
9cd1453f26 |
fix(sessions): a stale explicit-close stamp no longer makes a session permanently uncompressible
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq. publish_compression_child() fails closed on any non-automatic end stamp and end_session() is first-stamp-wins, so a stale tui_close on a session the TUI still routes turned every rotation into "compute the summary, then discard it" (#106459). The host that still routes the session clears the stale explicit close via SessionDB.reopen_if_explicitly_closed() before the turn starts; publication never heals explicit closes. Review probes by @ehz0ah. |
||
|
|
22356075f1 |
fix(cron): make user-written bare-platform home targets continuable
A managed cron with `deliver: "slack"` (no captured origin) delivers to the Slack home channel, but the brief was never mirrored into that session and the in_channel seed never fired: bare-platform targets resolved with no `_resolved_from` provenance, so `_target_mirror_eligible` treated them like `all` broadcast expansions. - `_resolve_delivery_targets` derives `from_broadcast` from the raw token and passes it to `_resolve_single_delivery_target`; a user-written bare platform token is tagged `_resolved_from: "home"`, `all` expansions stay untagged (fan-out is never continuable). - `_target_mirror_eligible`: `home` eligible under the same flags as `origin_fallback` (per-job attach_to_session wins, else cron.mirror_delivery). - `_MIRROR_PROVENANCE_RANK`: `home` == `origin_fallback` so token order through dedup cannot strip eligibility. - Docs, tool schema text, cron/gateway AGENTS.md updated; existing exact-dict pins carry the new key. Squashed from PR #101819 (commits 5fe612c9, 65e36926, 10e0af1d, f2cc4639): they predate the cron/scheduler.py -> scheduler_delivery.py split and do not cherry-pick individually onto current main. |
||
|
|
948c9dbfc0 |
fix(files): scope backslash compensation to local Windows
The shared quoting helper doubled regex backslashes for remote shells because the controller ran Windows. Apply compensation only when the backend executes locally; serialized POSIX command text stays literal. Path translation and file-write payloads remain unchanged. Verified with the real LocalEnvironment argv path and JSON command text delivered to a real Bash stdin. Tests compare received bytes for quotes, metacharacters, newlines, empty values and backslash runs. The serialized case fails before the fix while the local case already passes. Integrated: 114 tests passed, 6 host skips. Ruff and whitespace checks pass. No Docker/SSH service or POSIX host run is claimed. The rejected grep multiline rewrite is not included. |
||
|
|
f2a19b0e06 |
refactor: remove unused extraction copies
Use the existing session-export, transcription and DingTalk owners. Remove unused setup/watchdog helpers and the no-op package migration hook. No package overrides it; user-state migrations keep their own existing owners. Keep version-change installation coverage and the scheduled external plugin compatibility blocks. Verified with the real adapter, transcription, session-snapshot and PM core suites. No live messaging service or user-state operation. |
||
|
|
fd605dcacf |
fix(wake): use one engine family and honor selected capture
Move the pyopen adapter to the shared engine module and remove shadowing definitions. Request audio-io only for local capture. Pass the caller's resolved capture mode into engine construction, so auto-selected client audio does not install microphone packages. Verified through the listener entry and existing engine/detector tests with disposable state. No live microphone or native SDK acceptance. |
||
|
|
b4d04eb8fd |
Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding
* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg
The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.
On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.
* docs(tool-search): connectors section — remote tools through the bridge
The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.
* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots
dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.
The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.
The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).
Live, 311 local tools + gateway, before -> after:
"send gmail email": 5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
"read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
"linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.
* refactor(tool-search): connector leg into tools/connector_search.py
tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.
No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.
* fix(tool-search): at most 7 queries per call, the gateway's search limit
One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.
The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.
* fix(tool-search): the model is told that connectors__ names are manage_connections accounts
tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.
The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.
Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.
Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.
* fix(connectors): /stop halts a connector batch before the next remote call
dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.
The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.
Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.
* test(connections): schema assertions become dispatch contracts
test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.
Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.
Test count in the file goes from 26 to 25.
* docs(tool-search): connector batches are one gateway request per entry
The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.
Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.
Docs only, no test.
* fix(tools): the between-turns refresh never rewrites the bridge tools
The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.
The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.
Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.
* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation
The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.
* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote
bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.
The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.
Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.
* fix(connectors): search keeps the twin a colliding name reaches, and says so
format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.
Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
|
||
|
|
cf4b78e91f |
fix(tool-search): a query no tool answers returns nothing, not five tools sharing one word (#106676)
search_catalog admitted every document with BM25 score > 0 and then padded
to `limit`. BM25 sums over the tokens a document shares with the query, so
on a 300-tool catalog "send gmail email" returned five incident tools that
shared only "email", and the discriminating word ("gmail", in no document)
had no say. The model read those as the answer.
Admission is now the query's rarest token: a document is a result only if it
contains the query token with the highest IDF, the one that names the intent.
Common verbs ("send", "read", "create") sit in dozens of documents and never
gate; vendor and object words ("gmail", "github", "incident") do. A token no
document carries admits nothing, and the existing empty-group hint tells the
model to retry without it. The name-substring fallback is deleted: it admitted
tools that matched no query token at all.
Result descriptions are clipped at 500 characters instead of 400. Over 353
vendor tool descriptions, 500 keeps 91% whole and every first sentence
(first-sentence max 329); 400 kept 82%.
Measured on the live 311-tool catalog with 25 hand-labelled queries:
precision@5 0.18 -> 0.43, wrong names returned 102 -> 66, false positives on
absent intents 17 -> 13. Live before/after: "send gmail email" went from five
betterstack tools to an empty group with the retry hint; "linear create issue"
and "betterstack incident" are unchanged.
|
||
|
|
e4cc7f09d9 |
merge: integrate upstream catalog with PM publication
Keep upstream's reviewed catalog as the only plugin name index. Catalog pins and custom update sources share staged PM validation. Publish code and dependencies with recovery after process death. Reject a concurrent enablement change before publishing disabled code. Use the manifest loader's supported version in the installer. Keep probe cooldowns for timeouts, not TLS failures that a CA change fixes. Preserve the backup, uninstall, browser and memory-provider repairs. Verified with the canonical runner on native Windows ARM64, real Git repositories, local TLS endpoints and UV dependency generations. Desktop catalog tests and both TypeScript checks pass. The full suite and native release builds were not run. No remote push. |
||
|
|
b7bef04861 |
fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver every surface funnels through, `kanban_db.create_task`: board-project inheritance now runs only when the caller left `workspace_kind` open (`None`), and `workspace_kind` defaults to scratch after that check. The tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and `board=` scoping but drops its handler-local sentinel logic, since the resolver now owns the rule; the `self_task` project inheritance for dispatcher-owned workers is unchanged. Sibling surfaces had the same bug through the same line and are fixed by the same change: - CLI `hermes kanban create --workspace scratch` on a project-scoped board produced a project worktree; `--workspace` no longer defaults in argparse so the resolver can tell "omitted" from "scratch". - Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same; `CreateTaskBody.workspace_kind` defaults to `None` for the same reason. - `kanban_swarm.create_swarm` threads `None` through for consistency. Tests: one resolver invariant in test_kanban_board_project.py (explicit scratch stays scratch, omitted still inherits) and the salvaged tool test folded into a single parametrized matrix over scoped/unscoped target boards. |