RunClock ticked from tasks.started_at (the task's first-ever start), so after
a review timeout + retry a healthy current run showed cumulative card age
(e.g. working · 2h for a minutes-old run).
_task_dict() now emits current_run_started_at from the run row
tasks.current_run_id points at (batched single query, same as latest
summaries), and RunClock prefers it, falling back to started_at for older
backends. Fixes#99819.
The control test only added one thing over the first new test: a quick()
deletion check for untracked scratch in a git-init'ed HERMES_HOME. It also
passed on base, so it pinned no new behaviour. Save both entries in the
first test's quick() call and assert the tracked file survives while the
scratch beside it is deleted, then drop the control. The stack now adds
one test, and that test is still red on base.
Update test_scratch_outside_git_trees_still_cleaned's docstring. It said
only a .git below HERMES_HOME marks a file git-owned, which stopped being
true once an enclosing repo that tracks the file counts too.
Keep the stack at two invariant tests: the tracked-file test now also
covers a file committed after first classification in the same process
and a HERMES_HOME nested in an enclosing repo; the separate stale-cache
test is redundant now that there is no cache.
Co-authored-by: David Crandall <david@convergentdesign.dev>
`_git_tracked_index()` cached one `git ls-files` set per process, so a
long-lived process could keep classifying a file as untracked after git had
taken ownership of it. `guess_category()` runs on every post-tool-call, so the
gateway can cache the index while `test_scratch.py` is still untracked; an
auto-snapshot then stages and commits it, `quick()` re-validates the stored
"test" entry through `guess_category()`, the stale cache still reports the file
as untracked, and `_delete_item()` unlinks a file git now owns (`git status`
shows `D test_scratch.py`).
Key the cached set on the git index's stat signature as well as the home, so a
staged, committed or unstaged change is a new cache key and no invalidation
hook is needed anywhere. The index path comes from `git rev-parse --git-path
index`, cached per home — so the per-call cost stays a single `os.stat`, which
also covers a linked worktree (where `.git` is a pointer file) and an explicit
`GIT_INDEX_FILE`. Fail-open behaviour is unchanged: no index and no git mean the
empty set, and the guard never raises.
Regression test: classify an untracked `test_scratch.py`, then `git add` and
`git commit` it in the same process with no `cache_clear()`, and assert
`quick()` leaves it on disk while an untracked control beside it is still
cleaned.
(cherry picked from commit a3271750608ca4e02e5205311a09b906e16d33d5)
The `_inside_git_worktree()` guard sliced its parent chain at HERMES_HOME, so
`.git` entries strictly below the home were the only ones consulted. That misses
the case where HERMES_HOME IS a git checkout: on this install `~/.hermes` is the
userfiles repo, so `~/.hermes/scripts/` and any top-level `tmp_*` file sit inside
a worktree yet resolve to no `.git` below the home.
Observed live: the bundled disk-cleanup plugin classified two COMMITTED
regression tests in `~/.hermes/scripts/` as disposable session scratch and
unlinked them — `scripts/test_analyze_upstream_opportunities.py` and
`scripts/test_customization_protocol_v2.py`; the deletion was then committed by
the next auto-snapshot (ce8dc99), so the suite that would have caught a scanner
regression was silently gone. 13 tracked `test_*` files under `scripts/` plus 17
tracked top-level `tmp_*` files were in the same blast radius.
Fix: when HERMES_HOME itself carries `.git`, ask git whether it TRACKS the exact
path (`git ls-files`, one cached subprocess per process keyed on the home) rather
than treating the whole home as protected — a scratch file merely living beside
tracked ones must still be cleanable, or cleanup would be disabled entirely.
Tests: two regression tests — a git-tracked `test_*` file is never classified
disposable and a stale pre-fix tracked.json entry is dropped by quick()'s
re-validation instead of deleted; plus the control that an UNTRACKED `test_*`
file in the same repo is still cleaned. Existing suite 30 -> 32 passed.
(cherry picked from commit c5555940108286c4c1070e84d6d694ad9a35bb76)
(cherry picked from commit c7c39e60d4dc467bf9681c3dcb5781ac96557724)
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Desktop opened /events with no since, and a missing cursor was read as
0, so every open replayed task_events history. Seed the socket from the
snapshot or this connection's last frame, and start a cursorless stream
at MAX(id). An explicit since still replays from there.
Fixes#81537
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
directly when the pool yields no token (no second uncached pool load,
no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
helper's error fallback reads the profile-scoped override, not the raw
process env.
- model setup flow: the confirm guards get the resolved Codex base, not
the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):
- the OAuth context-length probe (agent/model_metadata.py) and the
/model picker's live discovery (hermes_cli/codex_models.py) both GET
https://chatgpt.com/backend-api/codex/models with
Authorization: Bearer <gateway key> whenever model.context_length is
not pinned — the key is sent to a service it does not belong to and
cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.
Fix, mirroring the quota probe's existing gate in auth_codex:
- both catalog sites now decline to probe non-JWT credentials (real
Codex access tokens are JWTs; a gateway key is not one) and fall
back to the static table / offline sources — same outcome as the
doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
/models instead of chatgpt.com (catalog URLs are built from the
resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
same way the text client does.
Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.
(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
These came back in #120220 but exercise nothing in CI, so they are not
restored coverage:
- 12 Playwright specs under apps/desktop/e2e: the Desktop E2E job is
`if: false` in ci.yaml, and the new core lane (#120287) only runs
apps/desktop/e2e/core/*.spec.ts. Dropped rather than left for that lane.
- tests/plugins/platforms/test_discord_voice_receive.py: skips on the
nacl/discord import (discord.py[voice] is only in the messaging extra CI
does not install) and copied the production opus loader.
- electron/command-screenshot.test.ts: darwin-only; the JS lane is Linux.
- electron/wsl-path-bridge-gate.test.ts: faked process.platform = win32.
- test_kanban_dashboard_plugin markdown sanitiser case + probe fixture:
sliced function text out of dist/index.js by brace counting and eval'd it.
The plugin exposes no seam to reach MarkdownBlock behaviourally.
- test_kanban_dashboard_plugin::test_dashboard_markdown_html_is_sanitized_before_render:
runs the real bundle's sanitizer and MarkdownBlock under node; script/event
handler/javascript: HTML never reaches dangerouslySetInnerHTML (security).
- test_meta_ai_profile::test_images_ride_user_turns_not_tool_results: vision on,
tool-message vision off, so screenshots never 400 in tool results (#101668).
- session-info.test.ts "rebuilt-runtime re-bind" (7 cases, replaces the regex
test tests/desktop/test_bots_chat_stream_rekey.py): the pane adopts a rebuilt
runtime id only when lineage matches and the old runtime is idle (#93942, #94417).
Review on #120071 (yoniebans) found tests deleted as 'dead' that were only
mis-gated, and regression tests with no remaining equivalent:
- Discord voice receive (real NaCl/RTP: wrong-key drop, DAVE decrypt
failure, malformed padding, non-allowlisted speaker) moves out of the
never-selected tests/integration/ to tests/plugins/platforms/, where no
conftest stubs discord.py; 34 pass with the messaging extra.
- Five *_windows_live.py files had a bare skipif(win32), so
list_os_marked_tests.py never selected them; now windows_only.
- iron-proxy token-swap E2E (opt-in gate) is back.
- Gateway regressions: cached agent iteration cap (#48127), clarify JSON
never renders as progress (#52374), handoff watcher DB work off the event
loop, delivery ledger single connection, checkpoint prune and memory trim
housekeeping ticks.
- Desktop: command-screenshot IPC rejection, inactive WSL bridge never
spawns wsl.exe, bootstrap runner git binary.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
Under multiplexing `get_provider_env` resolved a scoped miss (`get_env_value` -> None) by falling
through to a bare `os.getenv`, which is the LAUNCH profile's `.env`: a routed profile with no
EXA_API_KEY/PARALLEL_API_KEY of its own searched on another profile's key, and the (key, client)
cache in `cached_sdk_client` faithfully built it a client on that borrowed key. Independent-review
major on #120092 (pre-existing on main, but this PR's claim covered only the both-have-keys case).
The bare `os.getenv` rung now runs only when no profile secret scope is bound (stripped installs,
plain CLI, systemd-injected env keep working); with a scope bound a miss is "" and the provider
raises its usual missing-key error. Isolation is between profiles; no environ fallthrough on a
scoped miss.
Tests: the A->B->A row gains a no-key profile (refused, launch key never used) plus a
single-profile control; the openrouter video sibling gets the same absence row (already correct
on this PR's head, red on main).
Sibling sites of #119986's class: the openai and meta-ai image backends
resolved their API key through get_secret but the base URL through
os.environ, so on a multiplexed gateway a routed profile's key was sent to
the launch profile's endpoint. Both fields now come from the same scope
(get_secret_str, like the DeepInfra video backend after #119986).
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:
- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
the credential pool, not the environment. Chat and image_gen/openrouter find
it through resolve_runtime_provider. video_gen reported OpenRouter
unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
routed profile's video jobs were submitted, polled and downloaded with the
launch profile's key and billed to that account. A profile whose key lived
only in its own .env could not use the backend at all.
The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.
OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.
Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.
Refs #68502
Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>
Two test functions: the lifecycle contract (clone/rename open their own DB) and the
save_config contract (this profile's own default spelling becomes the placeholder; a path
elsewhere is kept as given, so a deliberately shared store still works).
The setup schema offered db_path as f"{display_hermes_home()}/memory_store.db",
so `hermes memory setup` (Enter on the field) and the dashboard form wrote the
active profile's concrete path into plugins.hermes-memory-store. initialize()
only expands a literal $HERMES_HOME, and neither `profile create --clone` nor
`profile rename` rewrites config.yaml, so:
- a cloned profile opened the source profile's memory_store.db: facts stored
in one profile were recalled into the other's prompts, both ways;
- a renamed profile opened profiles/<old>/memory_store.db, which MemoryStore
re-created as an empty DB in a ghost directory under the old name, while the
real facts sat orphaned in the renamed directory.
The schema default is now "$HERMES_HOME/memory_store.db", the value the docs
already give, which initialize() resolves against whichever profile opens it.
save_config also stores a db_path equal to this profile's own DB as the
placeholder, so a config written by an older setup is repaired on its next save
from the CLI or the dashboard. A path anywhere else is kept as given.
- Desktop drawer -> Linear-style modal (smaller than Settings): main column
holds diagnostics, description, result/summary, dependencies, comments,
activity, runs, and the worker log tail; a right property sidebar holds the
inline editors (assignee, model override) plus priority/tenant/workspace/
created rows, estimate, and attachments. Backdrop click or Esc closes.
- GET /tasks/:id gains 'link_tasks' ({id,title,status} per linked task) so
Blocks/Blocked By chips render titles instead of raw ids; older backends
fall back to short ids. Additive; 'links' shape unchanged.
- New backend tests (test_kanban_link_tasks.py) + drawer tests for title
chips and the id fallback.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Follow-up on the two salvaged commits:
- One _sidecar_payload_text() helper decides the /send text for both the adapter
path and _standalone_send (cron / send_message), which had the same raw-markdown
leak on URL-bearing messages and never stripped with PHOTON_MARKDOWN=false.
- BlueBubbles (same iMessage surface) keeps [label](url) targets as bare URLs.
- Tests trimmed to two invariants.
The gateway's tool-progress path emits terminal commands as fenced code
blocks on any adapter whose supports_code_blocks is True. The Photon
adapter set that flag from PHOTON_MARKDOWN, but the sidecar's /send
router (send-format.mjs) silently routes every URL-bearing message
through the plain-text builder — where fences survive as literal backtick
characters — and even the markdown path renders a fence as inline
monospace text, not a block. Net effect: raw fenced terminal commands
(and raw markdown around them) surfaced in iMessage bubbles no matter
what the user's prompt-level style rules said.
Fix the whole class: the adapter never claims code-block support, so the
gateway emits its compact one-line tool preview instead. Prose markdown
passthrough (bold/italic/headings) is unchanged, and no config change is
required — tool_progress can stay at its user's preferred mode.
Tests: new capability + E2E tests pin that the gateway cannot emit a
fence for Photon with real adapter + real display resolution; the old
supports_code_blocks-mirrors-env expectation (which pinned the buggy
behavior) now asserts the flag is always False.
chooseSendFormat() routes markdown containing a raw http(s) URL through
spectrum-ts' text() builder, because the markdown builder's iMessage data
detection 500s on those messages (#73615). text() ships the payload
verbatim, and nothing strips the markers on the way down, so any reply
that mentions a link arrives in iMessage as literal markdown source:
**Release 1.2.0** is out
- **EUR 5** off this month
https://example.com/releases/1.2.0
Every ** is visible in the bubble. Remove the URL from that same reply
and it renders correctly, which is what makes the URL the trigger rather
than the content.
spectrum-ts documents markdown() as degrading to readable plain text on
platforms without native support "instead of surfacing raw ** markers".
Selecting text() ourselves opts out of that guarantee, so the adapter has
to honour it instead.
Strip in _sidecar_send() when the payload is markdown and the sidecar will
downgrade it. The format key is deliberately preserved: the sidecar owns
the builder choice (test_rich_links.py pins that contract), and an older
sidecar without chooseSendFormat must keep rendering natively.
_send_plain_fallback() selects the text builder explicitly via
markdown=False and had the same leak, so it strips too.
Reuses the shared strip_markdown() helper rather than adding a second
implementation, so the PHOTON_MARKDOWN=false path and this path produce
identical output and both inherit any future fix to that helper.
Diagnosis previously reported in #85733, which was closed unmerged.
Stripping must not take the URL with it, though. The shared helper
collapsed [label](url) to label alone, which on this path is worse than
raw markdown: iMessage auto-links bare URLs and nothing else, so the
reply arrives with a description and no way to reach the link.
Open the itinerary on Google Flights <- URL gone entirely
So strip_markdown() takes keep_link_targets, which rewrites
[label](https://url) as "label\nurl" (own line, because these URLs are
often long) and leaves non-http targets such as mailto: or relative
paths label-only, since iMessage won't linkify those either. The default
is unchanged, so the SMS, IRC, Feishu and QQ callers keep dropping the
target as before.
plugins/platforms/line/adapter.py already carries a private
strip_markdown_preserving_urls() for exactly this reason ("LINE
auto-links bare URLs only"). This moves the behaviour behind the shared
helper instead, so Photon's three plain-text paths -- the downgrade,
_send_plain_fallback(), and PHOTON_MARKDOWN=false -- all agree.
gateway/platforms/bluebubbles.py is the same iMessage surface with the
same loss; left alone here to keep this change to one platform.
Deleted: the seven files whose subject was plugins/memory/hindsight itself
(provider, config schema, env perms, local-runtime hint, templates, health
grace timeout, root guard) and the hindsight-only cases inside the multiplex
identity-scope, provider-thread, BOM-tolerance and session-switch suites.
Retargeted: the dashboard memory-provider config-surface tests used hindsight
as the only bundled provider with a flat `<home>/<name>/config.json` declared
schema; they now install a synthetic user plugin (`flatprov`) into the isolated
HERMES_HOME so the generic declared/live PUT + GET paths stay covered. The
lazy_deps plugin-owned-range test keeps mem0ai as its subject.
Parse memory-provider manifests through the declared hermes_yaml layer so manifest-name deny-list entries cannot fail open when PyYAML is absent. Exercise enabled and disabled profiles A→B→A to lock down config and module cache isolation.
Salvage of #66777 (@chrisyoung2005): the dashboard toggle toast now says a restart is needed
(#71595). Moved the ``restart_required`` stamp from a dashboard router wrapper into
dashboard_set_agent_plugin_enabled so the ``plugins.manage`` RPC (TUI/Desktop) carries it too;
contract + generated TS regenerated. The salvaged endpoint tests stubbed the toggle they were
checking — replaced by invariant tests that run the real commands against a temp HERMES_HOME
(all red on origin/main): platform disable gates the loader, alias toggle writes the canonical
key, status honours bundled defaults + memory.provider, remove forgets config / resets
memory.provider / unlinks a symlink only, disabled memory provider is not loaded.
test_deferred_platform_client_tools: bundled a2a is keyed ``platforms/a2a`` now.
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995