Setup needs declared extras instead of unbounded package installs. Pin ddgs,
langfuse and Piper without changing any existing resolved package versions.
Use upload-time cutoffs for the reviewed pins.
NeuTTS excludes Python 3.14. KittenTTS requires misaki, which excludes Python
3.13 and newer. Keep matching dependency markers and PM gates so bundles omit
these engines and explicit requests fail instead of reporting empty success.
Piper also excludes Intel macOS and Windows ARM64 because its native closure
lacks wheels there.
Verify the parent sync gate against the native Python version, including an
installed-anchor override. The test fails without that gate. Check declaration
and runtime gate agreement across bundle targets. Isolate an existing test's
home lookup, whose failure also reproduces on the base commit.
Verification: public PM lock/check and fresh 85-package environment passed in
an isolated manager runtime. DDGS, Langfuse and Piper imports passed on Linux.
No full suite or cross-host execution was performed.
Sherpa 1.13.8 supplies native Windows ARM64 Python and core wheels.
Remove its platform exclusion so auto selects the keyless engine there.
Declare pypinyin for both wake install paths. The tokenizer imports it
unconditionally, including for English phrases. Include sentencepiece
in the lazy engine extra too. Keep the matching core in the lockfile
and admit the reviewed release with bounded package cutoffs.
Verified native ARM64 Python 3.14 with the unchanged Hermes engine:
three phrases detected twice each, with no triggers in 180 seconds of
silence. Engine construction failed without pypinyin and passed with it.
The tested wheels match the lockfile hashes. Focused tests: 78 passed,
2 skipped. uv lock --check passed. No MSIX rebuild or microphone test.
The dependency audit reports the same existing httpx2/httpcore2
advisories as the parent commit. No unrelated packages changed.
Build TUI, web, desktop UI and runnable agent products from explicit
prepared inputs. Keep dependency preparation separate from distribution
packaging, with PM and native builds sharing uv environment construction.
Docker copies compiled frontend products instead of build dependencies.
Nix retains uv2nix environments and consumes shared assembly through store
references. Native desktop and Termux use the same launcher and frontend
contracts. Preserve the independent PM runtime and source imports from
arbitrary working directories.
Keep failed frontend builds from replacing the previous product, reject
source/output overlap, and bound dependency-process output draining.
Include hermes_wisdom in the Nix wheel: real CLI smoke tests exposed its
missing package declaration on the base revision too.
Verified focused Python and JavaScript suites, Docker build/runtime checks,
Nix desktop and CLI/ACP checks, standalone TUI and packaged Electron PTY,
and real full-Chromium interaction. Native signed installers, Android device
installation and the full repository suite remain CI verification.
Activation reaches plugin discovery before the application dependencies
exist. Give PM its own locked Python project and runtime so it can install
or repair the application without importing that dependency tree.
Keep PM outside the application workspace. A shared uv workspace resolves
the application graph and cannot provide this isolation. Route mutations
through an isolated worker and preserve transaction callbacks, cancellation,
custom package registrations, and correlated receipts.
Use the same runtime builder for source installs and packaged payloads.
Keep offline wheelhouse support in that builder. Nix builds the independent
PM lock as a separate derivation. Refuse lazy-disabled bootstrap before
installing tools or dependencies.
Move first-party YAML readers and writers to ruamel. Keep the application
lock's transitive PyYAML requirements for third-party packages.
Verification:
- Focused canonical Python suite: 177 passed, 1 host-gated skip.
- Electron backend probes: 12 passed. Electron typecheck passed.
- Both uv locks, scoped lint, Bash syntax, and whitespace checks passed.
- Cold activation, corrupt-app repair, offline staging, and relocation ran.
- Built and exercised the Nix PM runtime and standalone YAML merge script.
Six broader caller test files retain the same 24 failing test IDs as an
archive of HEAD. The existing real-home guard blocks those tests before
they can exercise the affected paths. No full-suite pass is claimed.
Native Windows signing and full Bionic package execution remain unverified.
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
Termux removes old package files, so a pinned URL and hash do not keep
build inputs available. Preserve the exact bytes without changing pins.
Archive every PM HTTP artifact and the Termux runtime inputs by SHA256.
CI reads R2 first. Only a missing object permits an upstream download,
hash verification, immutable upload, and verified readback. Seed the
actual toolchain and payload stores before their consumers run.
Use the public archive as a pinned fallback in PM, bootstrap installers,
and Nix fetchers. Keep network retries bounded and report attempted URLs.
Keep publication credentials in protected CI jobs, not installed clients.
Verification:
- 283 targeted tests passed; five POSIX tests skipped on Windows.
- All 87 preserved Termux packages passed local archive miss/hit checks.
- Native ARM64 ripgrep installed through the mirror and ran successfully.
- Wheel import, workflow lint, Python lint, shell syntax, and pins passed.
Live R2 publication, POSIX tests, and Nix builds remain for native CI.
The real-byte archive checks used loopback HTTP, not the live bucket.
Use a verified writable copy of pinned Python for uv builds. Keep the
signed launchers and bundled Python as the MSIX runtime so plugin
rebuilds do not remove package access.
Identify Microsoft Store builds from the artifact stamp. Use StoreContext
for update checks, downloads, and installation requests inside Hermes.
Keep relaunch and recovery ownership across the update. Do not report
unknown checks or no-op requests as successful installations.
Open the build commit link in the system browser.
Verification: 168 Python tests, 107 Electron tests, and 121 renderer tests
passed. Typechecks, lint, lock validation, and desktop builds passed.
A real packaged process completed plugin admission, import, and repair.
Store-acquired download, installation, and relaunch remain unverified.
This host has only the sideloaded MSIX. Its real Store API call returned
an error, which the app reports as unknown. macOS updates are unchanged.
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.
Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.
Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
Keep downloads bound to their remote representation and publish through
atomic destination-local staging. Serialize shared partial ownership.
Keep explicit CA trust scoped to provider probes. Preserve checkpoint
history and edited files, validate all profile inputs before dependency
publication, and separate data removal from installed runtime ownership.
Exclude machine-specific PM state from portable transfers. Keep plugin
files and nested skill tools intact. Preserve native test isolation.
Focused native Windows receipts cover the individual repairs and their
integration. This commit does not claim a full-suite or release build.
Merge upstream b1f003e186 while preserving PM runtime ownership and
Python 3.14 worker startup, Windows signing, and macOS wait recovery.
Keep retired runtime modules deleted. Port upstream updater preflight
checks into the checkout strategy and preserve live build logging.
Carry checkpoint filename handling and process recovery into the current
module layout. Regenerate locks and adapt incoming platform test markers.
Focused Python and JavaScript tests, desktop and root-test typechecks,
conflict-path lint checks, lock validation, and retired-import checks pass.
The full test suite and packaged release builds were not run.
The pyopen-wakeword universal2 wheel contains an ARM64-only TFLite library.
Exclude that dependency on Intel macOS in both wake extras and the PM gate.
Keep sherpa and Porcupine available without changing the selected provider.
Wake status now refuses unsupported engines even when lazy installs are
allowed. It recommends supported alternatives instead of a failing install.
The documentation identifies the platform limit and the configuration command.
Marker and status regressions failed before the change. Focused tests pass:
52 passed and 2 skipped on Python 3.11, 50 passed and 4 skipped on native
Python 3.14. Frozen exports retain both Intel alternatives. No locked package
versions changed. Native release acceptance remains pending.
Finish bootstrap uv before PM replaces its store entry. Keep failure
receipts stdlib-only and align the cryptography requirement and override
with the locked version.
Let bundle builders declare launch paths and update ownership. Remove
payload discovery, Store probing, and the unused develop command.
Derive Nix Python from the PM lock and share its provenance stamp.
Document setup, activation, optional dependencies, and distribution
ownership. Targeted Windows tests, relocated runtime launches, Electron
bundling, and bilingual docs builds pass. Native Nix and signed-package
acceptance remain CI gates.
Pin Android psutil to upstream commit
380bd2b59c67b0e1b04bbf3a90b11744f4f96644. It contains the unreleased
platform recognition and disk_partitions fixes. Other platforms keep 7.2.2.
Remove the local psutil source patch.
Retain the Git source through requirement normalization and wheel builds.
Use its locked version for offline installation. Provision Git in the
builder and host parsing dependencies through an isolated uv environment.
Wire the source-pin regression tests into Termux verification.
Targeted tests, lock checks, Ruff and shell/workflow checks passed. A real
wheel from the pinned source built and ran on Windows ARM64. Cold-cache
host normalization also passed without packaging installed on the host.
Full bionic compilation remains unverified.
The pywinpty 2.x cap selects a source build whose PyO3 version rejects
Python 3.14. Require >=3.0.5,<4 and remove the ARM64 exclusion. Both
architectures have compatible wheels with the required native binaries.
Only pywinpty changes version in the lockfile. Existing PTY bridge and
process-registry tests passed under Python 3.14.7 on x64 emulation and
native ARM64: each run had 88 passed and 38 skipped.
The ARM64 PTY test environment excluded cryptography because its local
source build needs OpenSSL. Full signed bundle acceptance was not run.
requires-python >=3.14,<3.15 (was <3.14 ceiling): the old cap existed because
Rust-backed transitives lacked cp314 wheels; tflite-runtime (openwakeword's
hard linux dep) still caps at cp311, so the wake engine moves to
pyopen-wakeword 1.1.0 — py3 wheels + bundled tensorflowlite_c lib, whose
bundled melspectrogram/embedding models are byte-identical to the openWakeWord
v0.5.1 files (verified by sha256), and it loads the shipped hey_hermes.tflite.
Drops openwakeword/onnxruntime/ai-edge-litert/wake-tflite machinery entirely.
win32-arm64 gets a marker gate (no pyopen-wakeword wheel there; porcupine
covers wake). uv.lock regenerated for 3.14 (254 pkgs, all platforms).
Real-lib smoke verified: engine builds against the actual wheel, silence
scores 0, noise scores ~0.003 (threshold 0.6), scores flow 1:1 per frame.
Icon generation selected the application venv, where resvg-py was
missing. Adding it to the dev extra also selected it for production
payloads built with --all-extras.
Use a locked icon-build dependency group in an isolated uv environment.
Keep its wheel cache separate from the PM cache copied into payloads.
Route generation and structural checks through the same runner.
Verified clean-source generation without resvg in the runtime venv,
runtime dependency export exclusion, targeted Python and JS tests,
and the full web workspace build. Signed desktop packaging was not run.
pytest-xdist is gone: run_tests.sh no longer dispatches on host, the
Windows xdist (--dist loadfile) arm is deleted, and pytest-xdist is
dropped from the dev extra (uv lock removes it + execnet). The
per-file subprocess runner (run_tests_parallel.py) is THE runner on
every host — the shape the CI linux lane already used.
- tests-os.yml: the macOS/Windows lanes now run scripts/run_tests.sh
like every other lane, passing the marker-narrowed file list via
--files and -m after --. The runner natively tolerates per-file
empty collections (exit 5, platform-gated file) and fails the run
when NOTHING collected — replacing the hand-rolled exit-5 branch.
- tests/gateway/conftest.py: the hasattr(config, 'workerinput')
controller-guard was xdist-only dead code; the file-locked cache
already handled the per-file model. Removed.
- tests/tools/test_browser_supervisor.py: the port came from xdist's
worker_id (fixed 9225 under per-file isolation — concurrent files
collide). Binds an ephemeral port instead (bind 0, read back).
- ~35 comment sites named xdist as the isolation mechanism; they now
describe the shared-process hazard they actually guard against
(bare pytest runs, same-file ordering) without naming a runner that
no longer exists.
check-appinstaller-update.py imports winrt.windows.applicationmodel
(Package/PackageUpdateAvailability) AND winrt.windows.management.
deployment (PackageManager) — separate winrt-windows-* distributions.
'[all]' on winrt-windows-management-deployment made uv lock reference
the whole family WITH an 'all' extra none of them ship ('does not have
an extra named all'), desyncing uv.lock so every '--locked' CI sync
failed. Declare both namespaces plain (>=3.2.1,<4; win32-gated);
uv lock regenerated (0 bogus extras, 6 winrt entries); lock --check
passes; CI's exact extra set resolves.
- desktop.md: dist:win is MSIX-only, not NSIS+MSI
- BUILDING.md: sign-nested-chromium is LIVE (after-pack.mjs wires it), not dead
- pyproject: lazy_deps.py comment -> pm
- photon docs: drop dead PHOTON_NODE_BIN rows (adapter is pm-store-first);
restore PHOTON_MENTION_PATTERNS row my earlier edit wrongly removed
- urllib_security/models docstrings: SSL_CERT_FILE/certifi fallback ->
platform trust store (post truststore port)
Cherry-pick of 03b45db777 from ethie/bundles-local-models: TLS trust was
five hand-rolled ladders (agent/ssl_verify, hermes_cli/auth,
agent/model_metadata, hermes_cli/urllib_security, gateway/run's SSL_CERT_FILE
mutation) all ultimately pointing at certifi's frozen list. Trust now comes
from the platform verifier via truststore: CryptoAPI on Windows,
Security.framework on macOS, OpenSSL's store on Linux; install_truststore()
patches ssl.SSLContext process-wide. agent/ssl_verify.py is the one
authority; agent/ssl_guard.py deleted.
Also closes the session-flagged coverage gap: hermes_cli/main.py (CLI
entrypoint) and tui_gateway/entry.py now call install_truststore() so
subcommands/help that never construct an AIAgent still get OS-store trust.
check-appinstaller-update.py imports the winrt Management.Deployment API
to ask the OS whether a newer MSIX package is registered. The dep was
never declared, so every real App Installer update-check fell back to
{available: null, error: 'winrt import failed'} — silent 'unknown'
verdicts on all out-of-store Windows installs.
Declare it win32-gated (the checker only runs on Windows) with the
>=3.2.1,<4 bound per dep policy; uv lock adds the winrt-windows-* family
(80 packages, all winrt-*, no unrelated re-resolve).
The pm package and its lockfile were missing from
[tool.setuptools.packages.find] include and package-data, so a sealed
wheel built with pm but no package manager inside it. test_wheel_ships_pm_package_and_lock_json caught it red; add pm/pm.* to the include
list and pm/lock.json to package-data.
The static py-modules list in pyproject.toml drifted from the source
tree each time the root layout changed. An installed wheel then
raised ModuleNotFoundError on import: hermes_state_common, and every
gateway or CLI start failed. The list missed hermes_state_holders,
mini_swe_runner, and the 15 new hermes_state_* modules from this
branch.
setup.py now derives py_modules from the source tree at build time.
setuptools package discovery sees only directories with an
__init__.py, so root single-file modules need py_modules in every
wheel build. setup() kwargs merge with pyproject.toml, and setup.py
is the only legitimate wheel or sdist builder, so the derived list is
the single source of truth.
The nix build is the only wheel consumer. Its source filter keeps
every root .py file, so the build sandbox derives the same set as the
checkout. Editable installs do not read py_modules: build_editable
never runs bdist_wheel.
Verified: wheel and sdist built with HERMES_NIX_BUILD=1 carry all 38
root modules. The guard still blocks builds without the nix env var.
The sealed uv2nix venv from nix build .#default imports
hermes_state_holders, hermes_state_sessions, mini_swe_runner, and
hermes_startup_watchdog.
(cherry picked from commit e23d467b52)
Invariant test updated to pin the derived list (no static py-modules; every root module packaged code imports ships).
The hermes_state split added 15 root-level sibling modules but
[tool.setuptools] py-modules still listed only the old set, so the built
wheel/sdist (and the uv2nix sealed venv) failed at 'import hermes_state'
with ModuleNotFoundError. hermes_state_holders and hermes_state_registry
were already missing on main (registry is imported by gateway/run.py and
hermes_state.py).
Add an invariant test that parses pyproject and walks every import in the
packaged root modules and packages, requiring any import that resolves to
a root-level *.py to be declared - so the class of bug is caught by CI
rather than by a broken install.
Platform gating moves from scattered per-package markers to a
readable per-extra authority (settled 2026-09-02 plan, task 9):
- pyproject [tool.hermes.extras-platforms]: extra -> PEP 508 marker.
matrix = linux-only (python-olm has no win/darwin build);
google-chat / mem0 = all but win32-arm64 (grpcio has no win_arm64
wheel). The per-package markers stay — uv's resolver needs them for
--all-extras (probe: a fully marker-gated extra syncs clean on
Windows); this table is the readable single authority for WHICH
extra is supported WHERE.
- pm.extras.extra_supported(): gate consult with an installed-override
rule (an extra whose anchors import on this machine is supported
regardless — dev machines, hand-synced venvs) and a fail-open-on-
malformed-marker posture. Cached table read.
- available() reads gated-off extras as unavailable; ensure_import()
raises InstallError naming the gate instead of a resolver no-op —
the adapter degrades with a readable reason, never a mystery.
Live-verified on this win32 host: matrix -> not supported (gate:
sys_platform == 'linux'), web -> supported, table parses.
tests/pm: 169 passed, 0 failed (4 new gate tests).
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.
Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
No shims, one marker system. The full mechanical sweep:
- 165 test files converted: every @pytest.mark.<os>_only decorator and
pytestmark assignment is now platforms("<os>"). Docstring/comment
prose mentioning the trio rewritten to the platforms() vocabulary.
- conftest: _OS_MARKS, the legacy skip loop, the double-mark reject for
the trio, and the selector-tagging shim are deleted. The single
platforms() gate does all host gating; _reject_contradictory_platform_
marks replaces the old reject (one platforms() marker per test — a
module-level pytestmark stacked on a per-test marker is the historic
skipped-everywhere-green-everywhere failure and stays a hard error).
- pyproject: only the platforms marker is registered.
- list_os_marked_tests.py: rewritten for the single vocabulary — takes
a platform name (linux/macos/windows), matches quoted platforms(...)
specs including negated and any-of forms, exits nonzero on empty
selection. Its test file rewritten to match (bare identifiers must
not match; the spec must appear inside the string literal).
- tests.yml macOS lane: marker: macos feeds the lister; the pytest
selection is -m "platforms and not integration" — the marker name is
the selector, the conftest's per-test host skips are the gate.
- test_os_marker_gating.py rewritten for the new reject (the old file
tested deleted machinery).
- AGENTS.md / CONTRIBUTING.md updated to the single vocabulary.
Verified: full-tree compile; marker unit tests (36); lister tests;
macOS-lane simulation (21 files, 180 collected / 429 deselected on a
windows host); broad xdist slices through scripts/run_tests.sh
(367 + 1674 passed, zero refactor-attributable failures — the 6 red
tests in the second slice fail identically with the changes stashed).
The per-file model pays a spawn+import wall of ~0.5-1.5s per file x ~3400
files on the Windows runner — a ~6-minute floor before any test runs
(measured: bare spawn 140ms serial / 531ms under 32-way contention; the
full pytest boot 1.4-1.8s per file; one-process collection of the whole
tree takes only 43s, so the architecture itself is the ~8x multiplier).
xdist with --dist loadfile pays that wall once per worker and keeps each
file's tests on ONE worker. Cross-file state pollution is bounded to files
co-scheduled on a worker; failures from that are stateful-test bugs to fix
(the same class PR #29016 fixed when it moved the other way).
Linux stays on the per-file runner: its spawn floor is ~15ms, so isolation
is nearly free there. Add pytest-xdist==3.8.0 to the dev extras and
scripts/run_xdist.sh. Windows-only experiment; the local 6h36m xdist run
on a polluted dev box (MSIX PYTHONPATH + SAC) is not representative — CI
renders the real verdict.
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.
P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.
P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.
P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.
P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.
P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.
Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.
Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
Each release exact-pins at least one dependency to a version published
days before the release (v0.20.6: snowballstemmer==3.1.1, psutil==7.2.2).
For two weeks after release the relative exclude-newer cutoff filters
those versions out, so any venv that predates the release cannot resolve
the new pins at all ('no version of snowballstemmer==3.1.1' — observed
2026-08-29 updating three production installs v0.20.0 -> v0.20.6, one
Termux and two Linux servers; the Termux host additionally bricked on
psutil==7.2.2 sdist resolution, and cryptography's isolated build
environment resolved maturin/setuptools-rust under the same cutoff).
Same zero-float-protection logic as the setuptools/pillow/mcp/
firecrawl-anydoc exemptions: the pin bump WAS the review, so the cutoff
adds nothing for an exact pin and can only brick. Extend
exclude-newer-package to every exact-pinned package in
[project].dependencies / optional-dependencies (table moved to
one-key-per-line — 97 entries), plus maturin and setuptools-rust for
wheel-less sdist builds of the exempted cryptography pin.
test_exact_pinned_deps_exempt_from_exclude_newer enforces the invariant
going forward: adding a name==version pin without a matching
exclude-newer-package entry fails CI.
Windows installer editable builds fail in uv's isolated sandbox with
ModuleNotFoundError: wheel.cli because build-system.requires only listed
setuptools. setuptools.build_meta and our setup.py bdist_wheel guard both
import wheel during the build.
Also whitelist wheel in tool.uv.exclude-newer-package so the existing
build-system exclude-newer brick guard stays green.
Fixes#96488
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
Co-authored-by: Olympusbuildz <Olympus.roots@outlook.com>
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>