Commit Graph

26895 Commits

Author SHA1 Message Date
Teknium
75bf672a78 test: managed-runtime source scan survives a vanishing sdist dir (TOCTOU)
Path.rglob raises FileNotFoundError when a directory disappears between
listing and scandir — a sibling CI job creating/removing its sdist
extraction (hermes_agent-<ver>/) killed test_allowlist_has_no_stale_entries
on run 33531869442. Switch to os.walk (tolerates vanishing dirs) with
top-level pruning of exempt and packaging dirs; file set is byte-identical
(886 files verified old==new) and a 30-scan churn harness that reliably
exercised the window shows zero errors.
2026-09-01 10:52:42 -07:00
Teknium
83cde7f31d test(gateway): widen pending-drain chain wait budget (4s -> 20s)
Same loaded-runner class: the 12-turn drain chain completed only 11
turns inside the 400x0.01s poll budget on main run 33455779041. The
loop still exits early on success, so the wider budget costs nothing
on healthy runs.
2026-09-01 10:52:42 -07:00
Teknium
28834a2098 test: raise tight wall-clock bounds that flaked on loaded CI runners
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).

Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
2026-09-01 10:52:42 -07:00
Teknium
67de93862c fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).

The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).

Fixes #98028
Fixes #100325
2026-09-01 10:52:23 -07:00
Teknium
894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
Teknium
09b88bab88 fix(state): stop the on-write identity probe cancelling our own POSIX locks (#100368) 2026-09-01 10:51:52 -07:00
teknium1
46e7ad8e12 fix(gateway): gate record-less visible-text match on _already_sent
Draft frames set _last_sent_text for dedupe without setting
_already_sent (they are ephemeral); an ungated has_delivered_text match
let a draft-only preview count as durable delivery and regressed
test_relay_seal_failure's dead-transport guarantee on CI.
2026-09-01 10:51:37 -07:00
teknium1
fd998120c1 fix(gateway): judge delivery success against final content, not flag trust (#95382, #98552)
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.

- delivered_final_matches: record-less flags are now reconciled against
  the FINAL content via has_delivered_text; only the explicitly-marked
  ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
  record their delivered payload (the last record-less flag setters);
  the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
  closed/reset) are classified as send_path_degraded (retryable) so the
  delivery-obligation ledger's reconnect sweep replays the stranded
  final response instead of losing it until a process restart.

Fixes #95382; closes the #98552 false-positive class.
2026-09-01 10:51:37 -07:00
Teknium
043c258ac2 fix(web): stale removed-backend config warns at startup and errors by name
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).

- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
  removed_backend_note(); selection_error() swaps in the specific
  removal explanation (removed in v0.21.0, keyless alternatives) while
  keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
  search_backend / extract_backend against the registry and emits a
  startup warning (deduped per stale value), surfaced by the existing
  print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
  per-capability keys, dedupe, healthy-config negative, live-backend
  failure text preserved.
2026-09-01 10:43:20 -07:00
rainbowgits
622883bad7 fix(agent): accept marker-only finish_reason after stream supersession
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:27:06 -07:00
kshitijk4poor
b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm
0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
xxxigm
43470980bf test(auth): cover keyed providers.<key> credential-pool lookup
New-style providers store keys under the durable config slug, but
runtime still looks up custom:<display-name> and sends a placeholder.
2026-09-01 22:42:27 +05:30
Teknium
6ddafd34f0 test(proxy): cover multi-line SSE data joins, truthy lastOne, EOF-without-blank-line dispatch 2026-09-01 10:12:21 -07:00
Teknium
93591eccb5 fix(proxy): harden SSE DONE tracker — spec multi-line data joins, truthy lastOne, EOF-write guard
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
  consecutive data: lines per the SSE spec (a split JSON event no longer
  reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
  proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
  widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
2026-09-01 10:12:21 -07:00
Jon Nielsen
d304422b3d fix(streaming): extract finish_reason/usage before content-shape continues
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.

Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.

Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
2026-09-01 10:12:21 -07:00
loulanyue
66d42e0dba fix(stream): do not misclassify stream with final usage chunk as mid-stream drop (#91373)
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.

If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.

Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
2026-09-01 10:12:21 -07:00
rainbowgits
ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
Hermes Agent
219cb714a1 fix(gateway): persist api_server async-delegation completions as deliveries, not user-turn wakes
On the stateless api_server platform, a background delegate_task
completion was delivered by self-POSTing /v1/chat/completions with
role=user after the parent turn had already completed (event.complete,
finish_reason=stop). That starts an unauthorized new agent turn the
client never sent, persists the completion as an ordinary role=user row
(display_kind NULL), and can blow through a pending human-confirmation
gate — the exact skip-ahead reported in #85957.

Fix: async_delegation completions targeting a non-push api_server
session are now written into the session transcript as a durable
DELIVERY row (role=user + display_kind=async_delegation_complete +
display metadata — the same bookkeeping shape the TUI/desktop delivery
path persists). No agent turn runs; clients polling
GET /api/sessions/{id}/messages see the result immediately, and the
next real client turn carries it as context. Persist failures return
False so the durable claim is released and the completion retried.

Watch-pattern notifications keep the existing self-post wake behavior;
push-capable adapters are untouched.

Fixes #85957
2026-09-01 10:11:07 -07:00
Teknium
e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
hermes-seaeye[bot]
e5e71d8c46 fmt(js): npm run fix on merge (#100517)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 16:59:53 +00:00
Teknium
045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
David Metcalfe
75f8ab2f87 style(desktop): satisfy eslint and prettier in profile migration
- curly: brace all single-line if statements in profile-migration.ts and
  profile-migration.test.ts (17 errors in CI check:lint)
- perfectionist/sort-imports: node:fs builtin import before vitest external
- padding-line-between-statements: blank lines after block statements
- prettier: normalize formatting (fmt script style) in the three touched files

All 967 electron project tests still pass.
2026-09-01 09:54:22 -07:00
David Metcalfe
b6a0360f81 fix(desktop): clarify which profile reads the migration precedes
Polish from antigravity review of the rebase resolution (GPT-OSS):
the previous comment said "BEFORE the first primaryProfileKey() /
primaryBackendIsRemote() read" but those two calls live at different
points — primaryBackendIsRemote() is the very next line, primaryProfileKey()
is inside the connection IIFE. Be explicit about which is where so a
future reader who moves one of them knows what to preserve.
2026-09-01 09:54:22 -07:00
David Metcalfe
285ced674f test(desktop): cover active-profile migration helpers
Addresses teknium1's review (#64195) finding #2: the multi-rung resolver
needs Electron tests covering precedence, stale-PID rejection, fallback
behavior, and the remote boot path. The pure decision helpers are now
covered by 29 unit tests in `profile-migration.test.ts` (vitest electron
project).

Coverage:
- precedence: legacy > single-running-gateway > state.db heuristic
- stale-PID rejection: recycled PIDs not owned by hermes are dropped
- malformed pid files: JSON parse errors, non-integer PIDs, zero/negative
- scoring edge cases: ancient files (recency floored at 0.1), tiny files
  (size floored at MIN_SIZE), larger DB beats smaller at similar recency
- single-profile fallback: best === 'default' suppresses the write
- no-op cases: preference file already exists, missing profiles root

The remote boot path is verified by code review of the call-site move
(commit preceding this one) — `migrateActiveProfileIfMissing()` now runs
before `primaryProfileKey()` is first read in `startHermes()`.

The pure decision logic that the orchestrator relies on is covered end-
to-end below; this matches the repo's testable-helper pattern (see
`profile-delete-routing.test.ts`).
2026-09-01 09:54:22 -07:00
David Metcalfe
5760c65be2 fix(desktop): migrate active profile before first primaryProfileKey() read
Addresses teknium1's review (#64195) finding #1: the previous PR placed
the migration inside the connection IIFE, AFTER
`resolveRemoteBackend(primaryProfileKey())`. When the preference file
was missing, `primaryProfileKey()` resolved to 'default' and the remote
branch returned immediately without ever reaching the migration. Remote-
mode users got no migration at all.

Move the call site to the top of `startHermes()`, before the connection
IIFE that reads `primaryProfileKey()`. Both remote and local branches now
flow through this path before any profile-dependent resolution, so the
migration runs on first boot regardless of mode.

The inlined implementation is replaced with a thin wrapper that builds a
`MigrationDeps` bag and delegates to `migrateActiveProfileIfMissing` from
`profile-migration.ts`. No production behavior change beyond the call-
site move.

Tests added in a separate commit.
2026-09-01 09:54:22 -07:00
David Metcalfe
a0511fd884 refactor(desktop): extract pure active-profile migration helpers
Addresses teknium1's review (#64195) — the migration decision logic should
be unit-testable without Electron. Pull the ladder (legacy sticky file,
running-gateway scan, state.db heuristic) into pure helpers in a new
`profile-migration.ts` module, following the dep-injection pattern
already established by `profile-delete-routing.ts`.

The helpers take an injected `MigrationDeps` bag so tests can exercise
precedence, stale-PID rejection, fallback behavior, and the single-profile
case without touching `/proc`, `ps`, or the host filesystem. The default
profile is explicitly rejected from the legacy rung because the regex
matches `default` and accepting it would suppress the heuristic that is
the whole point of the migration.

The wrapper in main.ts is unchanged in behavior — the same `MigrationDeps`
fields get filled in from `fs`/`path` and `isHermesProcess`. The atomic
write + parent-dir-create that the wrapper performs matches
`writeActiveDesktopProfile`'s semantics so the migration produces a file
indistinguishable from a user-driven profile switch.

No production behavior change; pure code organization.

Tests added in a separate commit.
2026-09-01 09:54:22 -07:00
David Metcalfe
4980a5ceae fix(desktop): migrate active profile preference from legacy signals on first boot
When active-profile.json does not exist (fresh install or first boot after
update), seed it from the best available signal so the Desktop launches
into the user's primary profile instead of always defaulting to "default".

Priority ladder:
1. Legacy ~/.hermes/active_profile (explicit CLI choice via hermes profile use)
2. Running gateway (gateway.pid with verified liveness + hermes identity check
   via /proc/cmdline or ps -o args= to avoid PID recycling false positives)
3. state.db heuristics — hybrid recency×size score picks the primary workspace
   (e.g. a 409MB coder DB beats a 28MB default DB even if touched at similar times)

The stored JSON includes _migrated:true for priority 3 (heuristic guess) so
the renderer can optionally surface a one-time notification. Priority 1 and 2
are higher-confidence signals and skip the flag.

The migration is a no-op once active-profile.json exists, and only writes
when a non-default profile is confidently identified — preserving the legacy
fallback-to-default behavior for single-profile users.

Fixes #64160 (active-profile half).
2026-09-01 09:54:22 -07:00
Konstantin Khlopkov
aea2daee29 fix(backup): import into the active HERMES_HOME instead of the root home
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.

Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.

Fixes #99839
2026-09-01 09:28:46 -07:00
joaomarcos
ecdbcef7af fix(compression): roll the live transcript back when an in-place compaction commit fails (#99477)
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.

That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.

The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 09:28:31 -07:00
Hermes Agent
3e56911edb fix(i18n): add sessionExpiredNoError to all fully-typed web locales
The salvaged commit added the key to web/src/i18n/types.ts and en.ts
only; the 15 web locales typed as full Translations (fr/de/es/it/pt/
ru/tr/uk/ko/ja/zh/zh-hant/af/ga/hu) then failed 'tsc -b' (the real
build path) with TS2741. web/ar.ts uses defineLocale and falls back
to en, so it needs nothing. Follow-up for salvaged PR #98716.
2026-09-01 09:28:17 -07:00
kokhlo
55181a7d79 fix(oauth): surface actionable guidance when sign-in windows lapse
The local expires_in countdown killed OAuth sessions with a bare
"Session expired" before the backend poller's enriched message (Portal
sign-in stalled in the opened tab, retry/API-key fallback) could reach
the UI, and the desktop onboarding poller had no local expiry at all —
a dead session polled forever. Both surfaces now lapse with guidance
naming the common cause, prefer the backend error_message when it has
one, and keep polling when the backend still reports pending (clock
skew).
2026-09-01 09:28:17 -07:00
fangliquanflq
26476a978a test(installer): cover managed Python timeout 2026-09-01 09:28:01 -07:00
fangliquanflq
c454fbc5bd fix(installer): harden managed Python resolution 2026-09-01 09:28:01 -07:00
fangliquanflq
a2895e9968 fix(installer): preserve managed Python fallback version 2026-09-01 09:28:01 -07:00
fangliquanflq
60c7c0c667 fix(installer): require Hermes-managed Python on Windows 2026-09-01 09:28:01 -07:00
Teknium
ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium
09cddb6063 chore: map contributor email for #100352 salvage 2026-09-01 09:27:00 -07:00
Teknium
4789fa1066 test(browser): isolate chrome-fallback sandbox test from the shared tmpdir
The test exercised _run_chrome_fallback_command against the REAL
/tmp agent-browser namespace, so any concurrent hermes/pytest process
running the orphan reaper could rmtree the fresh pidless socket dir
mid-command (FileNotFoundError on _stdout_open — the recurring CI
flake, 3 hits incl. two main runs on 2026-09-01). Route it through a
private tmp_path like every sibling browser test.
2026-09-01 09:27:00 -07:00
Hermes Agent
4cca38be86 fix(browser): preserve starting sessions during orphan reap 2026-09-01 09:27:00 -07:00
Adrià Arrufat
2e2cdc2927 fix(browser): claim the Chrome-fallback socket dir before using it
_run_chrome_fallback_command creates agent-browser-<session> in the shared
tmpdir and then opens stdout/stderr files inside it, but never writes the
<session>.owner_pid marker. _reap_orphaned_browser_sessions rmtree's any
agent-browser-* dir that carries no live owner and is not tracked in the
calling process, so a second hermes process — or a parallel test worker —
deletes the directory between the makedirs and the first os.open, and the
command dies with FileNotFoundError on _stdout_open.

Write the owner marker immediately after creating the directory, which is
what every other socket-dir user already does.

Deterministic repro on main: run any test that exercises the fallback while
a second process calls _reap_orphaned_browser_sessions() in a loop —
5/5 fail before, 6/6 pass after.
2026-09-01 09:27:00 -07:00
kshitijk4poor
58472d803a refactor(state): flatten registry acquire flow; restore mirror cleanup assertion
Post-merge review follow-ups on #100201:

- acquire(): drop the never-iterating 'while True' and the redundant
  'existing is not generation' half of the race check — after the retire
  path, generation is always None on the fresh-open leg, so 'existing is
  not None' is the complete condition. Same behavior, flat flow.
- test_mirror: the conversion to the shared registry dropped the
  cleanup assertion entirely; restore it by patching
  hermes_state.release_or_close and asserting the handle is released
  exactly once after _append_to_sqlite.
2026-09-01 21:12:35 +05:30
Teknium
28044757aa fix(desktop-update): self-heal a TCC-anchor-bricked venv in the posix hand-off (#95759)
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.

posix.sh gains, before the update invocation:

* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
  a scrubbed-env `import encodings` boot probe), marker-validated
  (absolute path, outside the venv), staged with per-attempt backups and
  full rollback if the repaired interpreter still fails its probe.
  Two repair shapes:
  - alias-brick (#95541 class): the anchored copy boots — re-materialize
    python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
    onto the copy is the crash shape). Marker kept: this is exactly the
    layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
  - full brick: restore python → symlink to the marker-recorded store
    interpreter (if it boots) and aliases → symlinks; marker removed.
    The unblocked `hermes update` then re-installs a boot-gated healthy
    anchor — one-shot convergence, not a loop.
  Fail-closed on missing/unbootable source (vanished uv store class),
  missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
  boots (the launchd-gateway shape), drive the update via
  `venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
  as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
  --install-root and reports, for the test harness.

Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).

NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.

Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.

Fixes #95759
2026-09-01 08:36:03 -07:00
Teknium
9f377264aa refactor(update): consolidate the gateway drain triage into one shared helper
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.

No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
2026-09-01 08:34:51 -07:00
salch-cred
49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Teknium
71c4bcf7af fix(cron): surface missed-fire catch-up lateness in hermes cron list / status
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).

- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
  scheduled_at, dispatched_at, lateness_seconds, and kind
  (on_time / late / catch_up, classified against the ticker tolerance and
  the schedule's catch-up grace window). Manual triggers and one-shots are
  not stamped (no scheduled instant to be late against / retired beyond
  grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
  show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
  catch-up, in both the built-in ticker and external provider paths.

CLI surface only — no new tools, no policy engine.

Addresses the visibility half of #99879.
2026-09-01 08:31:37 -07:00
Teknium
269e5bde33 fix(desktop,i18n): polish Russian catalog and register ru in locale tests/docs
- Translate leftover English 'toolset(s)' strings in ru.ts
- Cover ru aliases/config value in languages.test.ts (parity with ar)
- List Russian among desktop UI languages in website/docs/user-guide/desktop.md
2026-09-01 08:31:03 -07:00
FASTCHIP
a922dad9d8 feat(desktop): add Russian (ru) locale
Full desktop UI translation (ru.ts, 3076 string/function leaves,
mirrors en.ts 1:1) plus locale registration in types, catalog and
language list with aliases (ru, ru-ru, ru_ru, ru-by).

Russian plurals (1 / 2-4 / 5+, 11-14 exception) via RU_PLURAL/RU_NOUN
helpers; count accepts number | string to match en.ts signatures.

Verified against current main: structural validation 3076/3076 with
placeholder parity, typecheck (renderer/electron/e2e) clean,
i18n vitest 28/28, prettier + eslint clean.
2026-09-01 08:31:03 -07:00
Teknium
7d8b450a6a chore: map contributor email for overtoneblue 2026-09-01 08:30:45 -07:00
Teknium
ddd8065232 test: regression coverage for first_chunk_at in post_api_request hook
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
  first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
  call's post_api_request payload (per-attempt reset in the loop)
2026-09-01 08:30:45 -07:00