Two test functions: the lifecycle contract (clone/rename open their own DB) and the
save_config contract (this profile's own default spelling becomes the placeholder; a path
elsewhere is kept as given, so a deliberately shared store still works).
The setup schema offered db_path as f"{display_hermes_home()}/memory_store.db",
so `hermes memory setup` (Enter on the field) and the dashboard form wrote the
active profile's concrete path into plugins.hermes-memory-store. initialize()
only expands a literal $HERMES_HOME, and neither `profile create --clone` nor
`profile rename` rewrites config.yaml, so:
- a cloned profile opened the source profile's memory_store.db: facts stored
in one profile were recalled into the other's prompts, both ways;
- a renamed profile opened profiles/<old>/memory_store.db, which MemoryStore
re-created as an empty DB in a ghost directory under the old name, while the
real facts sat orphaned in the renamed directory.
The schema default is now "$HERMES_HOME/memory_store.db", the value the docs
already give, which initialize() resolves against whichever profile opens it.
save_config also stores a db_path equal to this profile's own DB as the
placeholder, so a config written by an older setup is repaired on its next save
from the CLI or the dashboard. A path anywhere else is kept as given.
- Desktop drawer -> Linear-style modal (smaller than Settings): main column
holds diagnostics, description, result/summary, dependencies, comments,
activity, runs, and the worker log tail; a right property sidebar holds the
inline editors (assignee, model override) plus priority/tenant/workspace/
created rows, estimate, and attachments. Backdrop click or Esc closes.
- GET /tasks/:id gains 'link_tasks' ({id,title,status} per linked task) so
Blocks/Blocked By chips render titles instead of raw ids; older backends
fall back to short ids. Additive; 'links' shape unchanged.
- New backend tests (test_kanban_link_tasks.py) + drawer tests for title
chips and the id fallback.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Follow-up on the two salvaged commits:
- One _sidecar_payload_text() helper decides the /send text for both the adapter
path and _standalone_send (cron / send_message), which had the same raw-markdown
leak on URL-bearing messages and never stripped with PHOTON_MARKDOWN=false.
- BlueBubbles (same iMessage surface) keeps [label](url) targets as bare URLs.
- Tests trimmed to two invariants.
The gateway's tool-progress path emits terminal commands as fenced code
blocks on any adapter whose supports_code_blocks is True. The Photon
adapter set that flag from PHOTON_MARKDOWN, but the sidecar's /send
router (send-format.mjs) silently routes every URL-bearing message
through the plain-text builder — where fences survive as literal backtick
characters — and even the markdown path renders a fence as inline
monospace text, not a block. Net effect: raw fenced terminal commands
(and raw markdown around them) surfaced in iMessage bubbles no matter
what the user's prompt-level style rules said.
Fix the whole class: the adapter never claims code-block support, so the
gateway emits its compact one-line tool preview instead. Prose markdown
passthrough (bold/italic/headings) is unchanged, and no config change is
required — tool_progress can stay at its user's preferred mode.
Tests: new capability + E2E tests pin that the gateway cannot emit a
fence for Photon with real adapter + real display resolution; the old
supports_code_blocks-mirrors-env expectation (which pinned the buggy
behavior) now asserts the flag is always False.
chooseSendFormat() routes markdown containing a raw http(s) URL through
spectrum-ts' text() builder, because the markdown builder's iMessage data
detection 500s on those messages (#73615). text() ships the payload
verbatim, and nothing strips the markers on the way down, so any reply
that mentions a link arrives in iMessage as literal markdown source:
**Release 1.2.0** is out
- **EUR 5** off this month
https://example.com/releases/1.2.0
Every ** is visible in the bubble. Remove the URL from that same reply
and it renders correctly, which is what makes the URL the trigger rather
than the content.
spectrum-ts documents markdown() as degrading to readable plain text on
platforms without native support "instead of surfacing raw ** markers".
Selecting text() ourselves opts out of that guarantee, so the adapter has
to honour it instead.
Strip in _sidecar_send() when the payload is markdown and the sidecar will
downgrade it. The format key is deliberately preserved: the sidecar owns
the builder choice (test_rich_links.py pins that contract), and an older
sidecar without chooseSendFormat must keep rendering natively.
_send_plain_fallback() selects the text builder explicitly via
markdown=False and had the same leak, so it strips too.
Reuses the shared strip_markdown() helper rather than adding a second
implementation, so the PHOTON_MARKDOWN=false path and this path produce
identical output and both inherit any future fix to that helper.
Diagnosis previously reported in #85733, which was closed unmerged.
Stripping must not take the URL with it, though. The shared helper
collapsed [label](url) to label alone, which on this path is worse than
raw markdown: iMessage auto-links bare URLs and nothing else, so the
reply arrives with a description and no way to reach the link.
Open the itinerary on Google Flights <- URL gone entirely
So strip_markdown() takes keep_link_targets, which rewrites
[label](https://url) as "label\nurl" (own line, because these URLs are
often long) and leaves non-http targets such as mailto: or relative
paths label-only, since iMessage won't linkify those either. The default
is unchanged, so the SMS, IRC, Feishu and QQ callers keep dropping the
target as before.
plugins/platforms/line/adapter.py already carries a private
strip_markdown_preserving_urls() for exactly this reason ("LINE
auto-links bare URLs only"). This moves the behaviour behind the shared
helper instead, so Photon's three plain-text paths -- the downgrade,
_send_plain_fallback(), and PHOTON_MARKDOWN=false -- all agree.
gateway/platforms/bluebubbles.py is the same iMessage surface with the
same loss; left alone here to keep this change to one platform.
Deleted: the seven files whose subject was plugins/memory/hindsight itself
(provider, config schema, env perms, local-runtime hint, templates, health
grace timeout, root guard) and the hindsight-only cases inside the multiplex
identity-scope, provider-thread, BOM-tolerance and session-switch suites.
Retargeted: the dashboard memory-provider config-surface tests used hindsight
as the only bundled provider with a flat `<home>/<name>/config.json` declared
schema; they now install a synthetic user plugin (`flatprov`) into the isolated
HERMES_HOME so the generic declared/live PUT + GET paths stay covered. The
lazy_deps plugin-owned-range test keeps mem0ai as its subject.
Parse memory-provider manifests through the declared hermes_yaml layer so manifest-name deny-list entries cannot fail open when PyYAML is absent. Exercise enabled and disabled profiles A→B→A to lock down config and module cache isolation.
Salvage of #66777 (@chrisyoung2005): the dashboard toggle toast now says a restart is needed
(#71595). Moved the ``restart_required`` stamp from a dashboard router wrapper into
dashboard_set_agent_plugin_enabled so the ``plugins.manage`` RPC (TUI/Desktop) carries it too;
contract + generated TS regenerated. The salvaged endpoint tests stubbed the toggle they were
checking — replaced by invariant tests that run the real commands against a temp HERMES_HOME
(all red on origin/main): platform disable gates the loader, alias toggle writes the canonical
key, status honours bundled defaults + memory.provider, remove forgets config / resets
memory.provider / unlinks a symlink only, disabled memory provider is not loaded.
test_deferred_platform_client_tools: bundled a2a is keyed ``platforms/a2a`` now.
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995
The OS lanes are marker-driven: list_os_marked_tests.py picks the files
a lane imports from their platforms() specs and the lane selects with
-m platforms. A test gated with skipif(sys.platform != "win32") is
therefore never imported on the Windows lane and skipped everywhere else
— it runs on no host. skipif(sys.platform == "win32") tests were merely
invisible to the lane bookkeeping, but the rule the tree now follows is
one host marker, never a bare skipif.
Mechanical mapping, semantics preserved: skip-on-Windows → "posix",
skip-off-Windows → "windows", skip-off-Linux → "linux", skip-on-macOS →
"not macos". The former skip reasons stay as trailing comments. A
non-host condition (os.geteuid() == 0) stays a separate skipif beside
the marker, spelled getattr(os, "geteuid", ...) so the decorator still
imports on Windows.
Where the conversion would stack two platforms() marks on one test (the
conftest rejects that at collection) the narrower mark wins:
- test_update_wedged_gateway: the class is already platforms("linux");
its per-test "needs UNIX sockets" marks were redundant and are gone.
- test_process_registry.TestSystemdCgroupIsolation: the class-level
skip-on-Windows moves onto the 11 methods that had no host mark; the
11 platforms("linux") methods keep theirs.
- test_file_ops_single_roundtrip: the two fifo tests drop their
platforms("linux") in favour of the module's "posix" (mkfifo exists on
macOS; both tests already skip when it does not).
- test_linux_desktop_entry / test_gateway_job_teardown_live: duplicate
or wider marks removed.
_daemon_subprocess_env copied dict(os.environ), which under multiplex is
the LAUNCH profile's: profile X's daemon started with the default
profile's HERMES_HOME and credentials. It now builds the child env with
served_profile_child_env(inherit_credentials=False) — the manager reads
the LLM keys from the profile's own 0600 env file, so the child needs no
credentials from us.
check_local_runtime spawned 'python -c import ... sentence_transformers'
(a torch cold start) from is_available(), unavailable_reason() and
initialize() on every session. The verdict is now kept per interpreter
path for the process; a reinstall publishes a new PM generation, so a
new path re-probes.
Merge fallout (my resolution errors, all caught by CI):
- hermes_cli/backup.py + gateway.py: `theirs` on those hunks re-imported clusters HEAD had
already moved to backup_restore.py / kept in the facade. backup.py loses the 349-line
duplicate (main's #110179 fix is ported into backup_restore._import_db_member); the
systemd service-unit cluster returns to gateway.py (PM's _prepare_service_launcher /
_pm_managed_node_dirs / _systemd_command have no home in main's extraction) with main's
utf-8-sig read. gateway_service_unit.py is dropped.
- gateway/run.py: main's plugin-update chore is not profile-scoped (the housekeeping
ordering test pins the scope/drain sequence).
- pyproject + 30 test files: `import yaml` -> `import hermes_yaml as yaml` (pm-clean has no
pyyaml); gateway/config._bundled_platform_manifest_name reads through hermes_yaml.
- tests re-seamed onto pm-clean's shape: residency admission (installed_engine),
supervisor child env (binary is a constructor argument), update import guard
(update_cmd_deps is gone; our probe already scrubs PYTHONPATH — both #115032 invariants
pass), shallow-count git responses (stash path asks `status --porcelain -z`); dropped
tests for retired code (_run_node_bootstrap/_ensure_tui_node, Windows resume demotion).
- tests/tools/test_local_env_blocklist.py: restore the two helpers the suite-reduction
commit dropped and the blocklist import.
Real fixes:
- pm: classify_uv_failure/ResolutionConflict move beside the uv runner (pm.environment,
stdlib-only). pm.workspace imports tomllib at module level and cannot load on the 3.10
bootstrap python that streams uv output in the Docker arm64 image.
- tools/browser_tool.warm_agent_browser_npx_cache: back as a permanent definition — it is on
the frozen old-updater surface, and the revert-scheduled compat pointer does not count.
- hermes_cli/memory_setup: the dashboard's pip row uses pm.environments.
running_from_selected_environment for installed vs restart_required.
- scripts/windows-build-deps.ps1: export DISTUTILS_USE_SDK/MSSdk so setuptools trusts the
primed MSVC environment instead of asking vswhere (`env -i` test runner on win32-arm64
compiling ruamel-yaml-clib); run_tests.sh forwards them.
- tests/pm/test_windows_build_deps.py: start the protocol test from a parent env without the
toolchain variables the runner job already exports.
- tests/conftest.py scrubs HERMES_BUNDLED_PLUGINS (Nix-wrapped hermes on the dev host);
tests/home_io_guard.py treats sys.path site-packages under the real home as the
interpreter's installation (PM-activated developer shell).
- tests-js: four `curly` lint errors from main's new scripts.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).
_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".
(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
`test_diag_severity_tokens_route_through_host_theme_tokens` matched the CSS
text byte-for-byte including the alignment padding (`warning: var(`), so a
reformat that keeps the computed value failed it (probe: single-space the
three declarations -> red) while it still could not distinguish which
token/fallback each rung used beyond that exact substring. Parse the
`--hermes-diag-*: var(--color-*, #hex);` declarations with a regex and assert
on the captured token/fallback pairs; a wrong fallback colour still fails.
The three `--hermes-diag-*` tokens were literals declared on the consuming
elements, so the theme engine's `<html>`-level custom properties could never
reach them and light presets rendered the warning badge at 1.8:1 contrast.
Chain them through the host tokens themes already set (`--color-warning`,
`--color-destructive`) with the shipped literals as fallbacks; same selector
list, so nodes rendered outside `.hermes-kanban` keep a value. Error and
critical share `--color-destructive` (critical keeps its bold weight).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
`_resolve_max_depth` is a process-wide `lru_cache`, so its once-per-value
warning survives across tests; any earlier in-process resolution of the same
invalid string turned `test_invalid_max_depth_warns_and_preserves_default_capture`'s
"exactly one warning" into 0 (probe: resolve 'nope' first, run the test ->
`assert 0 == 1`). An autouse fixture clears the cache around every test so the
assertions describe the plugin, not the order the tests ran in.
#116713 parsed HERMES_LANGFUSE_MAX_DEPTH inside `_safe_value`, i.e. once per
captured prompt, response, tool input and tool output. With an invalid value
(`abc`) every captured field logged the same "Invalid ... Falling back to 4"
WARNING for the life of the process — one multi-tool turn fills agent.log.
Resolve the depth in `_resolve_max_depth`, an lru_cache keyed on the raw env
string: the warning fires once per distinct bad value, a changed env var is
still picked up by a long-lived process (mirrors `_capture_mode`, which also
reads per call and warns once), and valid values skip the int() parse after
the first call.
Follow-up to #116713 (independent review finding).
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.
Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:
- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
`_discord_free_response_channels` and `_missed_message_backfill_channels` now share
the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).
Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
The cherry-picked ownership check walked every ancestor up to `/`, so a
HERMES_HOME kept inside a dotfiles checkout (`~/.git`) turned every
root-level test_* scratch file into a protected "git-owned" file and
silently disabled the plugin's core contract. Cut the walk at
HERMES_HOME for in-home paths; out-of-home (/tmp/hermes-*) trees keep
the full walk since is_safe_path already bounds them.
Tests trimmed to the two invariants: quick() drops a stale tracked entry
for a committed test inside a linked worktree (.git pointer FILE) instead
of deleting it, and root-level scratch is still deleted even with a .git
above HERMES_HOME. Dropped the contributor's literal /tmp test (the repo
never writes /tmp) and the guess_category-only case the quick() test
already drives.
guess_category() matched test_*/tmp_* by basename alone, so a committed
regression test inside a git worktree under $HERMES_HOME/worktrees/ or a
/tmp/hermes-* checkout was tracked and auto-deleted by quick() at session
end (#115295; the protected-top-level-dir half landed in #114770).
Classify such files as non-disposable whenever a .git entry (directory or
linked-worktree pointer file) exists on the directory chain. quick() and
dry_run() already re-validate stored "test" entries through
guess_category(), so stale pre-fix tracked.json entries are dropped from
tracking instead of deleted — no separate migration needed. Scratch
test_* files outside git-owned trees keep aging out as before.
Fixes#115295
The three tests asserted literal phrases in a module docstring and two
markdown files; they guard prose, not behaviour, and would break on any
rewording. The docs change itself is the deliverable.
The prior commit only documented the <=2-member DM auto-classification
in the adapter.py module docstring — source an operator configuring
MATRIX_REQUIRE_MENTION etc. would never read. Issue #114733 explicitly
asks for the callout wherever MATRIX_ALLOWED_ROOMS,
MATRIX_FREE_RESPONSE_ROOMS, MATRIX_REQUIRE_MENTION, and MATRIX_AUTO_THREAD
are documented for operators, i.e. the env-var reference table and the
Matrix user guide. Add the same rule + bypassed vars + escape hatch to
both, matching the precedent set by 08aa67e473 for this kind of env-var
clarification, and extend the doc-content test to cover them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.
Fixes#114733
attachTouchDrag() armed a drag on ANY touch pointerdown and immediately
called preventDefault(), which suppresses the synthesized click
TaskCard.handleClick relies on to call props.onOpen(). There was no
movement threshold, so a finger drifting even ~2-3px on a normal tap --
which is universal on real touch hardware -- was enough to arm the
drag and swallow the open.
Fix: defer starting the drag proxy and calling preventDefault() until
the pointer has actually moved past an 8px threshold (matches the
common native drag-affordance convention). A stationary tap never
crosses the threshold, dragging is never armed, and the click fires
normally. A real drag still claims the gesture identically to before,
just after the same few pixels of travel every touch drag implementation
already tolerates.
The bundle (plugins/kanban/dashboard/dist/index.js) has no build step --
it is hand-maintained directly, as established by prior kanban dashboard
PRs (#114882, #108694) -- so the fix is applied there.
Closes#115568.
Testing: no jsdom/vitest harness exists for this bundle (confirmed by
PR #114882's review follow-up, which explicitly rejected turning a
"live-repro jsdom harness" into a pytest because jsdom/react aren't
declared in the root package.json and the Python CI job has no
node_modules -- such a test would be vacuous in CI). Per that
precedent and the "never read source code in tests" rule (no
regex/substring pin on the bundle text), this PR instead extracts
attachTouchDrag() verbatim at test time via Node (already present:
tests-js/ + vitest are in the repo) and drives it through real
pointerdown/pointermove/pointerup sequences against a minimal DOM
stub -- a behavioral test, not a source-shape test. Proven red on the
unfixed bundle (asserts preventDefault is called on a stationary tap)
and green on the fix; skips cleanly via shutil.which("node") if Node
is unavailable in a given lane.
Verification:
- node tests/plugins/fixtures/kanban_touch_drag_probe.js against the
ORIGINAL (unfixed) bundle: fails with "FAIL: a stationary tap called
preventDefault (suppresses the click)", exit 1 -- confirms the probe
reproduces the reported bug
- Same probe against the fixed bundle: "PASS", exit 0
- scripts/run_tests.sh tests/plugins/test_kanban_dashboard_plugin.py --
42/42 passed (1 new, 41 unchanged)
- node --check plugins/kanban/dashboard/dist/index.js -- syntax OK
get_board() buckets one list_tasks() fetch, so the done column
inherited the shared priority DESC, created_at ASC order — creation
order, which says nothing about when work finished. Sort the done
bucket newest-completed-first (completed_at DESC NULLS LAST, id DESC)
and expose that as a completed-desc list_tasks sort key; queue lanes
keep the FIFO dispatch default.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
_fetch_discovery followed redirects but only pinned the document's
self-asserted issuer field, so one cleartext or attacker-hosted hop
could serve a forged document claiming the configured issuer with
attacker jwks_uri and token_endpoint. Verify then accepted
attacker-signed ID tokens and the code exchange POSTed the client
secret to the attacker's token endpoint.
The resolved response.url must now share the configured issuer's
origin (scheme, host, port with default-port normalisation) before
the body is parsed. Same-origin canonicalisation redirects still
pass, and the issuer-field pin remains as the misconfig check it is.