Commit Graph

8950 Commits

Author SHA1 Message Date
kshitijk4poor
41c17d9b5e refactor(cli): reuse storage tool-call strip patterns in _strip_reasoning_tags
The display stripper kept byte-identical inline copies of
_STRAY_TOOL_CALL_CLOSER_PATTERN and _UNTERMINATED_TOOL_CALL_PATTERN
(same flags), which had to be edited in lockstep (#102303). Import the
compiled patterns lazily instead.
2026-09-24 22:25:37 +05:30
AdamPlatin123
dfb0a8846e fix(agent): strip stray arg tags to line boundary, not end of text
(cherry picked from commit 443db6b1b7076fcb08fb6ddacf11baa902ac544c)
2026-09-24 22:25:37 +05:30
kshitijk4poor
1e04c64ef3 fix(stream_diag): feed the existing chunk-body upstream_provider into diag + hook payload (#90216)
Drop the duplicate chunk-reading helper, run_agent facade forward and
_last_serving_provider agent state; the chat-completions loop already captures
chunk.provider, so stamp it on the per-attempt diag there and read the hook's
upstream_provider from the assembled response.provider.
2026-09-24 22:25:06 +05:30
kouyichi
0370beae2e fix(agent): record the serving downstream provider in stream-drop diagnostics
Relay routing is re-rolled per request and the winning downstream is reported only
inside the delta chunk bodies, so the header snapshot in agent/stream_diag.py could
not attribute a mid-stream drop to a provider. The per-attempt diag now carries
serving_provider (first non-empty chunk-body provider), log_stream_retry prints it,
and the post_api_request payload exposes it as upstream_provider for plugins auditing
route compliance. Fixes #90216.

(cherry picked from commit f0cf4fe4f4b47e384532008b1bf56ad259c87dc2)
2026-09-24 22:25:06 +05:30
brooklyn!
177f275b77 fix(plugins): route inject_message to the TUI session_key queue
Ink TUI and desktop never registered an inject host, and sharing
set_gateway_message_injector with a live messaging gateway would let
the last writer win. A separate host queues the reported session_key
onto that session's prompt queue and leaves other keys for the gateway.

Fixes #87412
2026-09-24 10:56:54 -05:00
kshitijk4poor
9c0e223d96 fix(sessions): let set-journal-mode --force waive only a failed holder scan
Retiring the win32 gate left `--force` with no reader, so on a Windows host
where the Restart Manager cannot start a session (restricted/service context)
the fail-closed `(-1, scan failed)` sentinel made `set-journal-mode`
permanently unrunnable, while optimize/optimize-storage/prune kept a working
override. `--force` now drops only pid <= 0 sentinel entries — a process the
scan actually found is still refused — and the help text says exactly that.

The holder test parametrised over `force` now asserts something real: force
plus a live holder is refused, force plus a failed scan proceeds (and without
force the failed scan is refused). Rewrite the user-guide paragraph that still
described a POSIX-only scan and a no-scan Windows path. Drop the dead
`import time`. The doctor holder test compares the child-reported pid instead
of `Popen.pid`, which is the venv launcher on Windows, so un-skipping it there
does not assert a pid equality that cannot hold.
2026-09-24 18:08:24 +05:30
JoaoMarcos44
6c71ff2bd7 fix(state): detect Windows database holders before maintenance
(cherry picked from commit b006ae2dcf6b210d3db3dfbe78c5aac040044644)
2026-09-24 18:08:24 +05:30
kshitijk4poor
b62ca15875 fix(auth): stop promising an automatic retry on edge-blocked Nous refresh
The upstream_blocked message said "Hermes will retry", but nothing schedules
a retry of the refresh; the next request simply tries again. Reword it to
what actually happens (credentials kept, try again shortly).

Fold the separate Retry-After test into the existing parametrized refresh
classification row (header + one assertion) and drop the assertions that
pinned exact error wording.
2026-09-24 18:06:48 +05:30
kshitijk4poor
1086bd6ccc fix(auth): bound device-poll edge backoff and gate 403 on x-vercel-mitigated
The edge/WAF backoff in the shared Nous+xAI device-code poll loop had three
gaps found in review:

- Retry-After was parsed with a bare int() and never capped, so a
  `Retry-After: 3600` slept an hour past a 5-15 minute device code. Parse it
  with the shared agent.retry_utils.parse_retry_after_seconds and bound every
  sleep by min(60, time left before the device-code deadline).
- The backoff was written into current_interval, so after a block normal
  authorization_pending polls kept the inflated interval and slow_down grew
  from it. Keep it in its own edge_backoff, reset on any OAuth JSON response.
- Any non-JSON 403 was treated as transient, disagreeing with the refresh
  classifier from the previous commit. Only a 403 carrying
  x-vercel-mitigated is the edge speaking; a header-less non-JSON 403 raises
  as before. 408/429/5xx stay transient.

Tests reduced to the two invariants: recovery after edge blocks (each sleep
<= 60, back to the server interval afterwards) and a persistent block ends at
the deadline without oversleeping.
2026-09-24 18:06:48 +05:30
zzragida
0b5eeab93a fix(auth): keep device-code polling alive through edge/WAF non-JSON errors
The generic RFC 8628 device-token poll loop (shared by the Nous Portal and
xAI flows) aborted the whole login when the token endpoint returned a
non-JSON error body. Vercel fronts the Nous Portal and answers rate-limited
clients with a text/plain 403 (x-vercel-mitigated: deny) or 429 — no JSON
body, so such a response can never carry authorization_pending/slow_down.
One mitigation response mid-approval killed a device login the user may
still be approving in the browser.

Treat non-JSON 403/408/429/5xx as transient: back off (doubling from the
current interval, floor 5s, cap 60s; Retry-After honored when present) and
keep polling until the device code expires. Statuses outside that set keep
the existing abort behavior and JSON OAuth errors keep each caller's exact
error contract.

Co-Authored-By: Claude Code <noreply@anthropic.com>
(cherry picked from commit 13ff660c5ebcbade203faad5e63c12de0603af5f)
2026-09-24 18:06:48 +05:30
kshitijk4poor
96f6abaeec fix(auth): keep Nous credentials when Vercel's edge checkpoint blocks a token refresh
Vercel's Security Checkpoint in front of portal.nousresearch.com answers
non-browser POST /api/oauth/token with a text/plain 403
(x-vercel-mitigated: deny) or a 429 challenge page. Those responses come
from the edge, not from the token endpoint, yet `_refresh_access_token`
folded the non-JSON 403 into the 401/403 -> invalid_grant default, so
`_refresh_nous_or_quarantine` wiped a still-valid refresh token and told
every affected user to re-login (#120602).

Classify a 403/429 carrying `x-vercel-mitigated` as `upstream_blocked`
(the code #115812 established for WAF blocks on the inference path):
retryable, relogin_required False, Retry-After forwarded through the
existing `parse_retry_after_seconds` helper. Nothing else moves: a 401
stays terminal even behind the header, a header-less non-JSON 403 keeps
the b8ce8875b0 stance (dead grant -> re-login), and 5xx/429/404 without
the header are unchanged.

Tests: three new rows in the classification matrix (403 deny, 429
challenge -> non-terminal; 401 + header -> terminal), a Retry-After
forwarding test, and a runtime-resolver test asserting the on-disk
access/refresh tokens survive a 403/429 checkpoint. All five new cases
fail on the previous commit.
2026-09-24 18:06:48 +05:30
rodricksz4h5
c633cd87c5 fix(memory): key the journey card lookups to the fingerprinted node id
Two places map a node id back to its memory card by rebuilding the bare
`memory:<source>:<index>` string: the timeline chart rows and the desktop
starmap's tooltip/body map. With the fingerprint in the id both missed
every card, so a memory row rendered an empty body.

Both now key the card under whichever shapes apply, so an imported or
pre-fingerprint graph keeps working. The CLI help and the memory doc now
describe the id as `journey list` prints it.

(cherry picked from commit fa0c984c5d7d7ad0820fa2f2bde088e532f4e897)
2026-09-24 18:01:46 +05:30
kshitijk4poor
d55d85c882 refactor(import): share the restore's member filter with the integrity pre-flight 2026-09-24 18:00:44 +05:30
kshitijk4poor
d3c4c86daa fix(import): pre-flight refuses every member read error, ignores skipped members
The integrity pre-flight caught only BadZipFile/zlib.error/EOFError, so a
bzip2 member with a bad stream (OSError "Invalid data stream"), an lzma
LZMAError or a media read error still escaped as a traceback instead of the
clean "archive is damaged" refusal. Name the archive-read errors once
(_ZIP_MEMBER_READ_ERRORS) and use the same tuple, plus OSError, in the
pre-flight and in the per-member catch of the restore (PermissionError is an
OSError, so it no longer needs listing).

The pre-flight also decompressed members the restore never writes
(gateway.pid and the other runtime files, archived SQLite sidecars), so a rot
in one of those refused an otherwise restorable backup. Move the skip rule
into _import_skipped() and use it in both places so they cannot drift.

run_import's docstring now states the real contract: 1 for a damaged archive
or an incomplete restore, None on success or a declined prompt.
2026-09-24 18:00:44 +05:30
kshitijk4poor
b24b7b725c fix(import): refuse a damaged backup archive before touching the home
`hermes import` only checked the central directory (`is_zipfile`,
`namelist`), so an archive with one member whose deflate stream or CRC is
rotten passed validation and blew up mid-restore with a zlib.error
traceback -- after config.yaml and everything before the bad member had
already been replaced, with later members never written (#121258).

Add a pre-flight pass in run_import that streams every member through
1 MiB reads (zipfile verifies the CRC at EOF) and collects every
BadZipFile / zlib.error / EOFError. If any member is damaged the command
prints a capped list and returns 1 with the home untouched. Stdlib
`ZipFile.testzip()` is deliberately not used: it lets zlib.error escape
and names at most the first bad member.

Widen the per-member catch from the previous commit with EOFError so a
member that rots between the two passes still becomes a "skipped" warning
+ `Import incomplete` / exit 1 instead of a traceback. Rework that
commit's test to drive the per-member path (the archive now never reaches
it with a corrupt member), and let `_break_member` serve the pre-flight
read before failing the restore's own read.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-24 18:00:44 +05:30
John Paul Soliva
4194919f5d fix(import): exit 1 and say "incomplete" when archive members were skipped
`_import_members` catches the PermissionError/OSError of each member it
cannot publish and records it in `errors`. That includes the state.db the
live-safe restore refuses because a gateway or dashboard still holds it
(#100960, #110179). `run_import` printed those under "Warnings (N files
skipped)", then printed "Done. Your Hermes configuration has been
restored." and returned None. `cmd_import` did not forward a return value,
so the shell status was 0. A script chained on `hermes import --force`
carried on over a partial restore, and the dashboard's import action
showed the green "done" badge. For the refused state.db case, the user's
sessions were never restored.

`run_import` now returns 1 when `errors` is non-empty. It words both the
summary header and the final line as "Import incomplete", and `cmd_import`
forwards the return code, the same contract `hermes backup` has for an
incomplete archive. Runtime files the import deliberately keeps
(gateway.pid, SQLite sidecars) and the older-backup session warning stay
warnings. The gateway revive still runs, so the files that did land come
up as before.

Measured on main with a fresh HERMES_HOME through `main()`: a 3-file
backup with one read-only target directory, and a refused state.db (live
DB left at 3 sessions / 12 messages). Both used to exit 0 after "Done ...
restored". Both now exit 1 after "Import incomplete: 1 file(s) were not
restored". A clean import still exits 0 and prints "Done".

(cherry picked from commit a0f808ec9f868128f4174e4816e095ac17f18d28)
2026-09-24 18:00:44 +05:30
KoNit-K
2ad5f1ccc0 fix(backup): continue import after corrupt member
(cherry picked from commit 92c121680c086765a06efee379b88b8239775b68)
2026-09-24 18:00:44 +05:30
kshitijk4poor
f59c5df755 refactor(config): share one seed_config_file between config edit and doctor --fix
`hermes config edit` and `hermes doctor --fix` each carried their own
"copy cli-config.yaml.example, else save DEFAULT_CONFIG" block, and they had
already drifted: doctor's template copy skipped _secure_file, so the seeded
config.yaml kept the checkout's mode instead of 0600. Seeder drift is the
bug class behind #121230 (one seeder wrote DEFAULT_CONFIG verbatim and pinned
global display values over every messaging platform's defaults).

Both now call hermes_cli.config.seed_config_file(config_path, template=None),
which returns whether the template was used (doctor keeps its message).
Doctor still passes its own PROJECT_ROOT template so its tests' patching
keeps working. Drops doctor_config's now-unused shutil import.
2026-09-24 17:59:37 +05:30
kshitijk4poor
0a4cb5fe8a fix(setup): don't pin a global tool_progress when hermes setup agent gets Enter
`hermes setup agent` offered "all" as the default for an unset
display.tool_progress, so pressing Enter wrote a global
`display.tool_progress: all`. The gateway merges no DEFAULT_CONFIG, so that
explicit global beats every platform tier and Telegram/Slack went from their
`off` default to `all` -- the same leak this stack removes from
_apply_default_agent_settings and _blank_slate_minimize_config (#121230).

When the key is unset, Enter now keeps per-platform defaults and writes
nothing; an explicitly typed mode, or re-confirming an already-set value,
still saves as before.
2026-09-24 17:59:37 +05:30
John Paul Soliva
e67fd76fa2 fix(config): stop seeding global display values that override every messaging platform's defaults
The curl installer, the Windows installer, the Docker first boot and
`hermes doctor --fix` copy cli-config.yaml.example into config.yaml byte
for byte. The template had five display keys uncommented: tool_progress,
interim_assistant_messages, long_running_notifications, busy_ack_detail
and show_reasoning. The gateway reads config.yaml without a DEFAULT_CONFIG
merge, and resolve_display_setting takes a global display.<key> ahead of
_PLATFORM_DEFAULTS. So every seeded home ran with those values on every
platform. Telegram and Slack posted every tool call. Signal, email, SMS
and the other no-edit platforms got progress lines, heartbeats and
interim messages. Every messaging reply had the reasoning block prepended.

First-time `hermes setup` (quick and full) and Blank Slate setup also
wrote display.tool_progress: "all". That write was added as a Quick
Install recommended default (79aeaa97e6) nine days before the
per-platform tiers landed (#8006), and it has the same effect for
tool_progress on homes the template never touched.

`hermes config edit` on a home with no config.yaml wrote DEFAULT_CONFIG
unstripped, which pins show_reasoning (all 21 platforms),
interim_assistant_messages (12) and tool_preview_length (16). It now
seeds like the installer and `doctor --fix`: the template when the
checkout has one (a full file to edit, written owner-only), otherwise
DEFAULT_CONFIG with defaults stripped.

Measured through the real gateway loader across the 21 platforms in
_PLATFORM_DEFAULTS, a template-seeded home differed from a bare one on
tool_progress for 19 platforms, show_reasoning for 21, busy_ack_detail
for 14, long_running_notifications for 13 and interim_assistant_messages
for 12. With the pins commented out and the setup writes removed, the
diff is empty, and the same holds for both `config edit` seeds.

The CLI does not depend on these values. It defaults tool_progress to
"all" and show_reasoning to true when the keys are absent, and the TUI
defaults interim_assistant_messages to true.

Homes that were already seeded keep their values. A template value
cannot be told apart from one the operator chose, so there is no
migration. The messaging docs now say which lines to delete.

(cherry picked from commit 96450d4500613ab1ba45c7e972f31de570bc2d71)
2026-09-24 17:59:37 +05:30
brooklyn!
5bc1761cf1 fix(models): allow spaces in self-hosted and user-configured model ids
The whitespace reject ran for every provider, so a self-hosted catalog or a
user-configured base_url could not select an id that legitimately contains
spaces. Cloud providers still reject. Picker payloads drop an id that
reject will still refuse.

Fixes #43140
2026-09-24 06:24:56 -05:00
kshitijk4poor
4cc4072747 fix(dashboard): map replaced/deleted-WAL store errors to 503 on GET /api/sessions too
GET /api/sessions and _resolve_session_id caught only sqlite3.DatabaseError,
so StateDbReplacedError / DeletedWalGenerationError (RuntimeError family)
fell through to the bare `except Exception` -> 500 "Internal server error" —
the same mis-mapping #110054 fixed for the analytics reads. Route that
exception family through the existing corrupt_store_as_status so the sessions
list and detail routes return the same structured 503 payload; the sqlite
arms (busy -> 503, malformed -> latch + 503) are untouched, which is why the
mapping is added as a sibling arm instead of wrapping the read (that would
skip note_storage_error for a malformed image).

The kept invariant test now also exercises both sessions.py paths; it fails
without this change (raw DeletedWalGenerationError escapes the resolver,
the route returns 500).
2026-09-24 16:47:38 +05:30
kshitijk4poor
fd2212975c refactor(dashboard): key the 503 payload by classify_persistence_error bucket
The deleted_wal/replaced payload selection re-derived the type-ordered bucket
table that hermes_state_errors._PERSISTENCE_CAUSE_BY_TYPE already owns; a
future bucket could silently diverge from the dashboard mapping. Look the
payload up by cause bucket instead (default: the corrupt payload, which also
covers an FTS-scoped malformed image). Proven behaviour-equivalent for every
exception that passes the guard (both replaced-family types and subclasses,
malformed sqlite errors incl. StateDbCorruptError and fts_index-scoped ones);
busy/locked and unrelated errors still propagate.

Also: one `state_db_*` naming scheme for the `error` codes
(`deleted_wal` -> `state_db_deleted_wal`; no consumer keys on it —
web/src/lib/api.ts only branches on the auth codes), drop the redundant
sqlite3.DatabaseError isinstance that is_malformed_db_error already performs,
and make the test assert the invariant (every `--fix` mention is negated)
rather than the exact wording.
2026-09-24 16:47:38 +05:30
kshitijk4poor
183f4ab66f fix(dashboard): drop the doctor --fix nudge from the deleted_wal/replaced 503 payload
The contributor's 503 messages told the user to "click Recover or run
`hermes doctor --fix`". Main's deleted_wal/replaced explainers
(agent/turn_explainers.py, hermes_state_errors.py) say the opposite:
running `doctor --fix` while a holder process is alive repairs the wrong
generation in place, and the Desktop Recover button from #110054 was not
taken. Hoist the two payloads into module constants next to
CORRUPT_STORE_DETAIL and reuse the explainer guidance (quit every Hermes
process, run `hermes doctor`, do NOT run `--fix`); the log line likewise
points at plain `hermes doctor`.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-24 16:47:38 +05:30
joaomarcos
15136c7a97 fix(dashboard): map DeletedWalGenerationError/StateDbReplacedError to 503 in corrupt_store_as_status (#110054)
Analytics routes open SessionDB under corrupt_store_as_status, which only
caught sqlite3.DatabaseError. DeletedWalGenerationError and
StateDbReplacedError subclass RuntimeError, so a retired WAL generation
escaped the guard and every dashboard poll turned into a 500 storm.
Catch StateDbReplacedError alongside the malformed-image case and return
the same structured 503 payload with detail.error='deleted_wal' /
'state_db_replaced'.

Hunk applied from PR #110054 commit 41e2b12a3e (web-router mapping only;
the holder-termination path from that PR is left to the maintainer).
2026-09-24 16:47:38 +05:30
brooklyn!
fbff2b8a7a fix(auth): audit API 401s and name stale app-token mint failures
A bearer 401 from the dashboard gate was returned to the client but
never written to dashboard-auth.log. Record session_rejected with the
client-facing reason, path, and IP, and never the bearer.

When that rejection is the app's saved bearer, the desktop mint error
says the app token is invalid instead of telling the user to
re-authenticate the server OAuth session.

Fixes #103117
2026-09-24 05:33:01 -05:00
brooklyn!
687ca6ed88 fix(hardware): use MemAvailable for Linux RAM
getconf _AVPHYS_PAGES counts page cache as used, so the Desktop statusbar
overreports Linux RAM. Read MemAvailable from /proc/meminfo, fall back to
MemFree, then the existing getconf probe. A reported 0 is a real available
value, not a missing field.

Fixes #102252
2026-09-24 05:29:08 -05:00
kshitijk4poor
754ea72538 refactor(credential_pool): collapse the stale-writer early returns
`_credential_token_pair` already maps a non-dict row to (None, None), so the
separate isinstance guard and second early return in
`_merge_pool_row_generation` were the same branch. Adopting a peer generation
now goes through the existing `_replace_entry` swap primitive.
2026-09-24 15:44:48 +05:30
kshitijk4poor
d525b443f5 refactor(auth): one token-base helper with one blank-row policy; dedupe generation copy loop
The id->token-pair base map was built three ways: _persist dropped
(None, None) pairs, load_pool and _sync_entry_from_pool_store kept them.
_merge_pool_row_generation treats a missing base as "unknown" (plain
recency merge) but a (None, None) base as a known generation, so a
token-less row got the generation override on the first flush after
load_pool and the plain merge on every later one.

Policy chosen: KEEP blank bases everywhere (new auth_mod._token_pairs_by_id,
built on the existing _entry_ids). A blank base means "no pair when we last
looked", which is exactly the CAS witness the boundary needs: when a peer
lands a pair on that row, every flush of ours keeps the peer's generation
instead of writing our blank tokens back. Dropping blank bases would make
the first flush keep the peer's pair and later flushes overwrite it.
_sync_entry_from_pool_store now records the base before its no-token-
material bail-out so all three sites agree. _update_root_pool_rows reuses
_entry_ids for its incoming map as well.

Also collapse the three copy-from-disk-else-pop loops in
_merge_pool_row_generation into a local _take_from_disk helper.
failure_reason is passed alongside _POOL_STATUS_FIELDS rather than added
to the shared constant, which auth_codex/auth_oauth_grants and
_merge_disk_cooldown_state also consume.
2026-09-24 15:44:48 +05:30
kshitijk4poor
badd8f56e4 fix(auth): scope only terminal DEAD verdicts to the stale token generation
When a stale pool's persist finds the disk pair moved, #120943 copied every
status field from disk, so the stale writer's NEWER account-wide verdict
(402 billing / 429 throttle -> EXHAUSTED) was erased and the rotated pair
re-entered selection immediately. Only a terminal auth death is tied to the
pair it was observed on: discard the stale writer's DEAD onto the peer's
pair, but leave any other status for _merge_disk_cooldown_state's ordinary
recency merge.

Also travel `scope` and `inference_base_url` with the pair (a Nous refresh
rewrites them together with the tokens, so a stale writer must not stamp an
old scope/route onto the newer pair); use the always-initialised
_persisted_token_pairs directly; and re-hydrate an in-memory entry from the
written row only when the store overrode its pair, so unchanged rows keep
runtime-only fields to_dict() omits.

Tests: fold the stale-later-402 witness (rt-1 kept AND last_status ==
exhausted) into the stale-terminal-verdict test and drop the separate 429
rollback test it subsumes.

Fixes #120815

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-24 15:44:48 +05:30
JoaoMarcos44
d2f54de2cc fix(auth): bind stale pool writes to token generation
(cherry picked from commit 1f5c7bcd076cd6184fb048ea1a57c4549f5f4991)
2026-09-24 15:44:48 +05:30
kshitijk4poor
d933ac2104 refactor(auth): classify a non-dict Nous refresh error body as empty
A JSON body that is a list or string reached `.get()` and raised AttributeError
instead of an AuthError. Treat it like a non-JSON body. Rename the parametrised
test to what it now covers (terminal and non-terminal rows) and state the
401/403 rule in the comment without pointing at a sibling it only resembles.
2026-09-24 15:42:28 +05:30
kshitijk4poor
f8b6075bcc fix(auth): drop duplicated retry hint from the Nous 5xx refresh message
`format_auth_error` already appends "Please retry in a few seconds." for
`code == "temporarily_unavailable"`, and every CLI surface routes through
it, so the raised message ended up telling the user twice:
"...(HTTP 503); retry shortly. Please retry in a few seconds." Let the
code-keyed formatter own the remediation text, as it does for every
other provider's transient error. No test asserts on the message text.
2026-09-24 15:42:28 +05:30
kshitijk4poor
b8ce8875b0 fix(auth): keep 401/403 Nous refresh replies terminal without an OAuth error code
The stack stopped defaulting an unknown non-200 token-endpoint body to
`invalid_grant` so a 429/404 gateway page no longer wipes credentials.
That also silently demoted 401/403 with a non-OAuth (or non-JSON) body
from terminal to "bench and retry every cooldown", so a dead grant on a
Portal that answers 401 without an `error` key would never prompt a
re-login. A 401/403 from the token endpoint always means the refresh
token itself was rejected, so mirror the Codex sibling
(auth_codex.py `status_code in {401, 403}` -> relogin) and keep the base
grant-dead default for exactly those two statuses; every other status
without an `error` key stays non-terminal.

Folding the non-JSON branch into the same path lets the `code`
coercion collapse to one line (gate quality finding).

Test: extend the kept parametrized row set with 401 (JSON, no `error`)
and 403 (non-JSON) rows asserting terminal; both fail on the previous
commit. Non-5xx rows no longer pin `retryable is None` (an unspecified
"raiser did not say" value), only the 5xx rows assert `retryable is True`.
2026-09-24 15:42:28 +05:30
kshitijk4poor
8a11d6342f fix(auth): only an explicit OAuth grant-dead code makes a Nous refresh terminal
_refresh_access_token defaulted any non-5xx JSON error body lacking an
`error` key to `invalid_grant`, and flagged a non-JSON body with
relogin_required=True. A 429 quota body or 404 gateway body says nothing
about the refresh token, yet the default made it a terminal grant error
that wipes the OAuth state and quarantines the pool entry (sibling of the
5xx outage path fixed for #120976).

Take the code as the server sent it (None when absent), derive
relogin_required from the shared _OAUTH_GRANT_DEAD_CODES set, and leave
the non-JSON branch non-terminal so the user is not told to re-login for
a transport-shaped failure. Fold the 429/404/non-JSON cases into the
existing outage parametrize so the stack still adds two invariant tests.
2026-09-24 15:42:28 +05:30
KoNit-K
4d05940823 fix(auth): keep Nous credentials on portal outages
(cherry picked from commit 85696b0176c939f0f56762cac2f64eff3d990498)
2026-09-24 15:42:28 +05:30
teknium1
f97608f178 chore: release v0.21.5 (2026.9.24)
Some checks are pending
Install & Update E2E / Pick release tags (push) Waiting to run
Install & Update E2E / Expand combinations (push) Blocked by required conditions
Install & Update E2E / ${{ matrix.name }} (push) Blocked by required conditions
Install & Update E2E / Upload leg player (push) Waiting to run
Install & Update E2E / Result chart (push) Blocked by required conditions
Live provider canaries / Live provider canaries (push) Waiting to run
2026-09-24 03:08:47 -07:00
kshitijk4poor
5c3afeee45 refactor(memory): derive a payload's destructive ops in one place
"The replace/remove ops of a staged payload, single or batch" was spelled
three ways across apply_memory_pending, the CLI pending list and the pin
step, one of them re-typing _BG_DELETE_ACTIONS as a literal tuple. One
destructive_ops() helper beside the constant replaces them, and the CLI
drops isinstance guards on records this code itself wrote (the apply path
never checked them either).

Tests: the legacy/unpinned case had its own outcome and an early return
inside the parametrised "names the removed entry" test; it is its own
test now, so each name states what it asserts.
2026-09-24 15:38:22 +05:30
kshitijk4poor
3da1c59377 fix(memory): refuse unpinned legacy replace/remove at approval
A destructive pending record staged before entry pinning carries only
its old_text search string. Replaying that search at approve time is
exactly the #120841 hazard: the live agent may have rewritten the entry
in place, and a newer entry that still contains old_text gets deleted or
overwritten. Naming the removed entry afterwards does not undo it.

So apply_memory_pending now fails closed: a replace/remove (single or any
batch op) without matched_entry is refused with nothing applied, the
record stays pending, and the message tells the approver to reject it and
recreate the change. /memory pending flags such records as an unpinned
legacy target so this is visible before approving.

The legacy case of test_approve_names_the_entry_a_remove_deleted now
asserts refusal + preserved record + intact entry; two existing tests that
hand-build replay payloads pin them with matched_entry, as real staging
does since the previous commit.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 15:38:22 +05:30
John Paul Soliva
b96e88b245 fix(memory): pin a staged replace/remove to the entry the approver reviewed
A staged replace/remove recorded only its old_text search string, and
/memory approve re-ran that search against the file as it was at
approval time. The background review stages every replace/remove it
wants (#105921), so when the live agent updated the targeted entry in
place and the new text still contained old_text, approving deleted or
overwrote the newer entry the approver never saw, and the output said
only "Approved 1 memory write(s)."

- Both staging gates resolve each replace/remove, under the store lock,
  to the full entry it matches and record it as matched_entry. No match
  or an ambiguous match is refused at staging, as the direct write is.
- Approval matches that entry exactly. If it changed since staging, the
  write is refused and the pending record is kept, as #96059 does for
  skills.
- /memory approve lists removed entries next to overwritten ones, so a
  record staged before this change (no matched_entry, still replayed by
  old_text) never removes an entry silently.
- /memory pending shows the full entry each staged replace/remove
  targets, not just its search string.

(cherry picked from commit 28700fa750c9321ded48be77a76531b5809f5fea)
2026-09-24 15:38:22 +05:30
kshitijk4poor
ca1ef789a8 refactor(cron): own the store-import loop in cron/job_definition.py
hermes_cli still reached into cron's private _jobs_lock and hand-built the
"created paused" record one function away from the module created so callers
never duplicate cron's schema. import_job_definitions() now holds the lock,
loads, merges and saves, and labels a merge ValueError with the job name;
_merge_cron_store keeps only the temp-store parse and the DistributionError
wrap, and takes the profile home instead of deriving it from dest.parent.parent.

Also drops the dead `and key != "repeat"` (repeat is reassigned right after)
and the comment that restated the module docstring.
2026-09-24 15:37:00 +05:30
kshitijk4poor
bb0650e541 fix(profiles): reject an unschedulable shipped cron job before any file is replaced
A shipped job whose authored schedule was a past one-shot (or an unparseable
string) raised ValueError from _apply_schedule_update in the middle of
_copy_dist_payload: SOUL.md and skills were already replaced, the cron store
was not, and the message named no job. `hermes profile update` ended half-applied.

- _copy_dist_payload merges the cron store first, so the one step that can
  reject shipped content runs while the profile is still whole; the merge
  itself only writes after every record merged.
- merge_job_definition normalises string schedules via parse_schedule (a
  hand-authored store previously persisted the raw string, and `cron resume`
  then crashed on `.get`), and passes schedule_display only when the authored
  record has one so the helper's display fallback applies.
- _merge_cron_store wraps the ValueError into a DistributionError naming the job.
- Paused/created stamps use hermes_time.now() like every other cron record.

The kept update test now covers the past-one-shot rejection (SOUL.md untouched,
local schedule kept) and a corrupt target store surfacing as DistributionError,
so dropping the error wrapper goes red.
2026-09-24 15:37:00 +05:30
kshitijk4poor
0bf0e03d33 fix(profiles): narrow cron-merge error wrapper to the corrupt-store case
load_jobs raises RuntimeError for an unreadable or unrepairable jobs.json;
that is the only failure worth relabelling as DistributionError. Catching
OSError/ValueError too hid the profile's own permission/disk errors behind
"Could not merge cron jobs" and double-wrapped invalid-schedule ValueErrors
the CLI already reports on its own.

Refs #120823

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 15:37:00 +05:30
kshitijk4poor
c2063cf61d refactor(cron): move job-definition merge schema into cron/job_definition.py
cron/jobs.py is already ~3400 lines; JOB_DEFINITION_FIELDS and
merge_job_definition are only used by importers of a foreign cron store
(profile distributions), so they live in a small dedicated module instead
of growing the scheduler file. profile_distribution imports the new module
directly; jobs.py is left byte-identical to main (no re-export shim).

Refs #120823

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 15:37:00 +05:30
JoaoMarcos44
39682e32ca fix(profiles): merge distributed cron jobs safely
(cherry picked from commit 291dca0055855abcec3bfbedb92d8890baeb833a)
2026-09-24 15:37:00 +05:30
kshitijk4poor
19bd20fcb6 test(config): tie LEGACY_KEY_STEPS to the migration ladder
LEGACY_KEY_STEPS is a frozenset of bare version ints beside MIGRATIONS with
nothing checking they agree: a typo or a retired step would silently change
what an unversioned config.yaml receives. Assert in the existing registry test
that every allowlisted version names a ladder step, and point authors of new
steps at the allowlist from the MIGRATIONS header.
2026-09-24 15:34:19 +05:30
kshitijk4poor
92324790f2 refactor(config): read the version stamp once and drop has_version_stamp
migrate_config() and the docker boot script each parsed config.yaml twice:
check_config_version() coerced a missing `_config_version` to 0 and threw the
"was it present" bit away, so has_version_stamp() re-read the file to recover
it, guarded only by a call-order promise in its docstring. That promise did not
hold for the docker script, which used the tolerant check: a list-rooted
config.yaml read as "unversioned, not below the floor", ran the backup +
migrate_config() dance and exited 1 (base: floor warning, exit 0).

Factor the read into _read_config_version_stamp() -> (Optional[int], latest);
None means the mapping has no stamp. check_config_version() is a thin wrapper
(None -> 0) so its 10 callers see identical output. migrate_config() and the
docker script decide `unversioned` from that single read; the docker script
now does the strict read itself and leaves an unparseable or non-mapping file
alone with a warning and exit 0, matching its invalid-YAML posture.
has_version_stamp() is deleted. Docker tests that mocked the pre-check now
mock the new helper.
2026-09-24 15:34:19 +05:30
kshitijk4poor
49aaa52d33 fix(config): keep v41 off the unversioned ladder and drop the stamp read guard
v41 walks every profile SOUL.md and deletes any "## Messaging other
agents" section on a bare heading match. A config.yaml with no
_config_version says nothing about where that SOUL text came from, so an
unstamped current config must not authorize rewriting a user-owned file;
the versioned 40->41 path is unchanged. The unversioned regression test
now seeds a user-authored section under that heading and asserts it
survives.

has_version_stamp() loses its bare except: migrate_config() already
called check_config_version(raise_on_parse_error=True) and the docker
script returns early on the (latest, latest) a parse failure yields, so
the raw read cannot fail at this point.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 15:34:19 +05:30
John Paul Soliva
a88bef98b2 fix(config): stop the migration ladder rewriting unversioned configs
A config.yaml without _config_version reads as v0 and is exempt from the
support floor, so the first `hermes update`, profile clone,
`hermes doctor --fix` or docker boot ran every one-time migration step on
it. Installers seed config.yaml from cli-config.yaml.example, which had no
version, and targeted writers (`hermes config set`, /personality, the
TUI/Desktop config writers) never stamp one, so this is the normal state
of --skip-setup, non-TTY and Desktop (--non-interactive) installs. The
value- and absence-based steps then reset the personality, raised the
delegation caps, turned verify_on_stop off, shortened the curator windows,
dropped model_catalog.ttl_hours and enabled plugins the user had installed
but never enabled.

- A config with no _config_version now gets only the steps keyed on a
  legacy key or identifier (LEGACY_KEY_STEPS), then the stamp.
- cli-config.yaml.example carries _config_version, so every seeded
  config (install.sh, install.ps1, docker/stage2-hook.sh, doctor --fix)
  starts at the current schema.
- docker_config_migrate.py no longer refuses a version-less volume with
  the "predates version 12" warning; like migrate_config() it migrates
  and stamps it.

(cherry picked from commit 97ba11e07009f633662b0c7fa8701aa5b441bd22)
2026-09-24 15:34:19 +05:30
kshitijk4poor
b9a50ddf1e refactor(dashboard): hoist /api/env preview check out of the write thread
Why: PUT /api/env raised the preview rejection as ValueError inside the
profile-scoped thread only because _env_write_errors(http_passthrough=False)
would downgrade an HTTPException to a 500. The check needs neither the
profile scope nor the thread, so run it up front and raise HTTPException(400)
directly like the messaging and custom-endpoint sites; the original one-line
lambda is restored. Also drop the ad-hoc ${…} regex copy in favour of the
existing _ENV_REF_RE.fullmatch from hermes_cli.config (identical matches).
Behaviour unchanged: 400 and secret intact (gate/probes/S1_e2e.py CLEAN).
2026-09-24 15:21:27 +05:30