Commit Graph

7183 Commits

Author SHA1 Message Date
ethernet
7f1ddc70ce feat(python): move first-party runtime to 3.14; swap wake engine to pyopen-wakeword
requires-python >=3.14,<3.15 (was <3.14 ceiling): the old cap existed because
Rust-backed transitives lacked cp314 wheels; tflite-runtime (openwakeword's
hard linux dep) still caps at cp311, so the wake engine moves to
pyopen-wakeword 1.1.0 — py3 wheels + bundled tensorflowlite_c lib, whose
bundled melspectrogram/embedding models are byte-identical to the openWakeWord
v0.5.1 files (verified by sha256), and it loads the shipped hey_hermes.tflite.
Drops openwakeword/onnxruntime/ai-edge-litert/wake-tflite machinery entirely.
win32-arm64 gets a marker gate (no pyopen-wakeword wheel there; porcupine
covers wake). uv.lock regenerated for 3.14 (254 pkgs, all platforms).
Real-lib smoke verified: engine builds against the actual wheel, silence
scores 0, noise scores ~0.003 (threshold 0.6), scores flow 1:1 per frame.
2026-09-07 14:03:27 -04:00
ethernet
37650f810c merge: integrate desktop update acceptance tests
Preserve the PM runtime-repair module boundary. Port the incoming stderr-streaming fix without restoring the deleted managed_uv downloader. Targeted runtime/progress tests and root JS checks passed; desktop typechecks passed.
2026-09-06 21:59:45 -04:00
ethernet
0b30c2484d merge: integrate termux cli bundles into pm-clean
Merge ethie/cli-bundles at 0765ad689b.
Keep PM runtime publication, install identity, TLS policy, and module
boundaries from pm-clean.

Resolve the Node version-discovery method in its owning class. Preserve
staged tools if a repin download or publication fails. Carry extra-only
memory-provider setup through PM and retain restart-required reporting.
Keep target-specific TUI path assertions and discard obsolete self-lock
fixtures and the orphaned Windows service handler.

Verified locally with the canonical Python runner, root JS checks, TUI
checks/build, shell parsing, and workflow YAML parsing. Existing platform
pins and executable modes are unchanged. No full-suite CI, new bionic
bundle, or phone acceptance is claimed for this merge.
2026-09-06 20:13:45 -04:00
ethernet
e6527b2a34 fix(termux): use the current CLI version predicate before dispatch
A stale helper reference aborted every non-chat Termux subcommand before APT ownership could be checked. Verify update, pm doctor and gateway status fall through to normal dispatch.
2026-09-06 16:57:41 -04:00
ethernet
a27cd5902a merge: integrate upstream prompt and plugin fixes
Keep the upstream legacy Bot Mode protocol cleanup and BOM-safe reads. Pin integration at 5bd439d3ed. Scoped Bot Mode tests pass.
2026-09-06 16:48:32 -04:00
Teknium
5bd439d3ed refactor(plugins): own dual-kind hook fallback in the ledger mixin
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
2026-09-06 13:36:12 -07:00
Joey
684a2cfbd7 fix(plugins): give dual-kind memory hooks a single owner 2026-09-06 13:36:12 -07:00
686f6c61
c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
Teknium
b167e81750 fix: legacy Bot Mode section in SOUL.md no longer taxes every session or shadows the live roster
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.

- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
  the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
  prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
2026-09-06 13:08:31 -07:00
Teknium
dec1e87833 Merge pull request #103553 from NousResearch/fix/goal-set-no-repaste
fix(goal): /goal <text> kicks with a pointer when the last user message already carries the goal (11-call replay turn in one run)
2026-09-06 13:03:01 -07:00
ethernet
d2c14f4bec merge: reconcile upstream bounded search with literal patterns
Merge d86627a7f3 into pm-clean. Keep the shared bounded ripgrep transport and preserve translate_path=False for probe patterns. Targeted canonical search tests: 19 passed, 3 POSIX-only skips on Windows. Full integrated CI has not been run for this merge.
2026-09-06 16:00:55 -04:00
ethernet
203c0a96aa fix(termux): carry the Vulkan loader and finish runtime diagnostics
The clean non-root install passed native imports but ffmpeg needed libvulkan.so. Bundle Termux's real generic loader and exercise media conversion in the bare runtime. Restore doctor imports, pm-venv recognition, and current diagnostics seams.
2026-09-06 15:41:56 -04:00
kshitijk4poor
4f6ef8806f refactor(auth): positive store-identity check for the clone strip; drop the save kwarg
The alias check re-implemented most of _is_same_auth_store (#101356) inline, but
that helper answers False on a samefile() OSError — on the strip path that would
fail OPEN. Resolve identity positively (same resolved path, or samefile says so)
and treat any error as "refuse"; three lines instead of the inline stat dance.
The `preserve_symlinks=False` save path (raw os.replace) bypassed atomic_replace's
EXDEV/Windows fallback to defend against a path swap no concurrent writer performs
on a directory this process just created — removed with its mock-injected test.
Tests kept: shared-store invariant [symlink|hardlink] and two fail-closed variants.
2026-09-07 00:42:18 +05:30
GuardianZ71
5e108489a0 fix(auth): preserve shared store during clone cleanup 2026-09-07 00:42:18 +05:30
Teknium
28747b4034 Merge remote-tracking branch 'origin/main' into fix/goal-set-no-repaste
# Conflicts:
#	hermes_cli/goals.py
2026-09-06 12:06:39 -07:00
Teknium
634109b7aa Merge pull request #103534 from NousResearch/fix/goal-judge-wait-on-subagents
fix(goal): the judge can WAIT on delegated subagents (4 of 5 nudges re-poked a waiting orchestrator; live 3/3 CONTINUE → 3/3 WAIT)
2026-09-06 12:03:15 -07:00
Teknium
afb4e080f6 Merge pull request #103496 from NousResearch/fix/goal-judge-own-processes-only
fix(goal): the judge sees only its own session's background processes; pid/session waits expire after 30 min (parked 3h22m on a grandchild's poller)
2026-09-06 12:03:05 -07:00
Teknium
e5f8420be8 Merge pull request #103513 from NousResearch/fix/subagent-context-cap
fix(delegation): compression_threshold_tokens is opt-in (default off, children keep the 500K ratio trigger); validate the value
2026-09-06 12:02:49 -07:00
ethernet
af8212b1f6 fix(runtime): route remaining feature consumers through pm
Migrate callers to declared extra names instead of adding legacy alias maps. Preserve lazy-install refusal behavior, isolate reimported test homes, and exercise the real callback boundaries. The relevant integrated batch passes 607 tests.
2026-09-06 14:34:51 -04:00
Gille
485aaf69b7 fix(local-runtime): find nvidia-smi in WSL driver path 2026-09-06 23:34:20 +05:30
ethernet
9d7921cd6e fix(runtime): restore startup imports and isolate updater tests
Use the canonical version identity reader instead of importing its deleted predecessor. Handle update metadata flags before mutation. Restore terminal environment imports and make platform-parameterized TUI tests host-independent.
2026-09-06 14:00:43 -04:00
ethernet
07ee929979 fix(gateway): resolve service runtimes through PM state
Service PATH uses the selected dependency generation and the target home managed Node records. Preserve lexical Python paths in desktop launch tests and create the staging artifact in the GUI build fixture. Remove unsupported classifier action inputs.

Canonical Windows selection passed 48 tests and skipped 57. Linux service assertions require the native CI lane; skips are not execution proof.
2026-09-06 13:53:11 -04:00
Teknium
dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
ethernet
acba441012 fix(pm): align context-home publication and recovery scopes
Resolve dependency state and the plugin union from the active home without changing process environment. Use the existing home-root derivation for custom and named profiles. The journal validator checks the same root as the writer.

Verified 71 tests passed, 8 skipped across root resolution, union, recovery, and selection. The new context-only-home regression failed before the fix.
2026-09-06 13:31:24 -04:00
kshitijk4poor
b114641c88 refactor(state): fold the speculative-open retry into acquire's loop; keep real inode-swap tests
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
  the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
  type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
  those files carry no windows_only marker so they never run on Windows, and the monkeypatched
  predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
  db.close() already releases a registry-shared handle, so the described NameError leak only
  existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
  teardown; replacement not published before the last close settles) plus the raising-close
  settlement; comments say the WHY once.
2026-09-06 22:56:51 +05:30
joaomarcos
aad74f26f9 fix(state): coordinate SessionDB teardown with active writers
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.

Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.

Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.

The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.

Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.

Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.

Fixes #102827

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
2026-09-06 22:56:51 +05:30
ethernet
ce49cdbc59 merge: reconcile upstream main with pm audit closeout
Merge upstream 5e645791ac.

Retain the PM feature-flag owner and add upstream connection options.
Use the deny-only window-open policy while trusted external links keep
the existing IPC path. Keep both session-import and external-link copy.
Preserve captured timeout output when adding terminal yield handoff.
Quickstart tests patch the explicit upstream model-assignment owner.
Migrate incoming legacy OS markers to the branch's platforms gate.

Desktop renderer and Electron typechecks passed. Targeted Electron tests
passed (42 tests), Python conflict checks passed (26 tests, 3 skips),
and the plugin-compat import checker passed. CI owns the broad merge gate.
2026-09-06 13:17:07 -04:00
Teknium
5e645791ac Merge pull request #104421 from NousResearch/feat/nous-anthropic-wire-auto
feat(nous): anthropic_wire=auto picks the session's wire from the first response's upstream (gated until GMI native is measured)
2026-09-06 10:04:56 -07:00
Teknium
45646f4d09 feat(nous): anthropic_wire=auto decides a session's wire from its first response
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.

`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.

GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.

Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.

Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
2026-09-06 09:11:04 -07:00
Teknium
0168bb27ad refactor(cli): one foreign-log walker for the CLI picker and the desktop browser
hermes_cli.foreign_sessions._walk owns discovery (glob per source, per-entry
OSError tolerance, symlink-escape and regular-file checks, newest first); the
CLI listing and the desktop browser's candidate rows both consume it instead of
carrying two scans of the same directories.
2026-09-06 09:09:41 -07:00
Teknium
f3400ce745 fix(cli): blank user turn no longer crashes foreign session discovery
_first_user_line indexed splitlines()[0] on a whitespace-only user message
(image-only / tool-only turn) and raised IndexError, which aborted the whole
CLI picker and the desktop session.foreign.list RPC. partition("\n") yields
"" for blank text and the loop moves on to the first real user line.
Reported first in #92290.
2026-09-06 09:09:41 -07:00
Teknium
fe49acb670 refactor(desktop): import-session entry is a sidebar nav row; browser module named as a foreign_sessions sibling
- Sidebar: the entry joins SIDEBAR_NAV (same chrome, active state, data-tour
  handle as the other rows) instead of a one-off Button below the rail;
  i18n moves to sidebar.nav['session-import'].
- Reuse common.retry / common.refresh / common.back instead of duplicating them
  under sessionImport in six locales.
- hermes_cli/foreign_session_browser.py -> foreign_sessions_browser.py so it
  sorts as a sibling of the foreign_sessions module it extends.
2026-09-06 09:09:41 -07:00
Adolanium
dfd0660fd7 fix(desktop): skip inaccessible logs during foreign session discovery
One unreadable or rotated log must not hide the rest of the list; discovery
tolerates OSError from resolve()/stat() per entry and reuses the stat result.
2026-09-06 09:09:41 -07:00
Adolanium
9186e3ebc5 feat(desktop): session import view for foreign coding-agent transcripts
Browse the backend host's foreign CLI session logs, preview a bounded read-only
transcript, and continue a copy in Hermes under the selected profile. Reuses the
hermes_cli.foreign_sessions parsers and the portability validator/writer;
imports are transactional and deduplicated on the recorded origin.
2026-09-06 09:09:41 -07:00
ethernet
3c08d16ba7 fix(pm): close runtime publication and updater audit gaps
Dependency publication now recovers interrupted config/facts changes before
activation and leases live generations during collection. Receipts retain
update correlation and failed steps across nested command boundaries.
Doctor and desktop surfaces report those failures through shared owners.

Move checkout updates out of the desktop facade. Stage a detached Windows
relaunch waiter before shutdown, with bounded handshake and process-birth
checks. Keep packaged lifecycle tests isolated from the installed app.

Native verification exposed two production races: cron maintenance imported
the interactive CLI and rewrote TERMINAL_CWD, and install-ID reads collided
with first publication. Use the existing owners and locks. Plugin checks
now run at startup and each due-gated housekeeping tick, not after 60 ticks.

Share updater-test mutation boundaries and remove collection-root fixtures.
Separate cold MCP startup from command latency and give the real HTTP drip
test enough time to reach body handling.

Root npm check passed, including packaging. The fixed-tree Windows Python
run reported 44557 passed, one failed, and 1404 skipped, plus one retry-only
HTTP test. Those final failures now pass in a 35-test bounded batch. A real
isolated gateway wrote startup and periodic plugin-check receipts.

Full final-tree CI, bundled Sandbox deployment, and actual App Installer
relaunch remain unverified. docs/pm-audit-status.md records these limits.
2026-09-06 11:45:41 -04:00
kshitijk4poor
a262b2e372 refactor(update): drop the tombstone comment at the old matcher site and a duplicated exactness leg 2026-09-06 20:47:52 +05:30
mengtanx
b4b6235239 fix(update): credit launchd ai.hermes.gateway in fleet reconciliation (#103679)
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.

Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
2026-09-06 20:47:52 +05:30
Teknium
5106e939e0 fix(tui): heartbeat ticks share the /loop idle poll and rewind when no turn starts
Folds the /heartbeat driver into the existing /loop poll slot (one cadence, one
try/except table) and closes the consumed-tick gap: a dispatch that raises OR is
refused by _admit_prompt_turn (returns False) releases the session claim and rewinds
the persisted fire via HeartbeatManager.abandon_fire(), so the tick stays due for
the next poll instead of advancing fire_count with no turn. abandon_fire refuses to
overwrite a pause/resume/clear that landed between claim and dispatch (mirrors
LoopManager.abandon_tick). Tests use the real SessionDB under a temp HERMES_HOME and
drive the poller loop itself; docs note the TUI/Desktop surface.

Rollback-on-failure idea credited to jerrygooch (#104011, 51df1ed4d04 / 283fa972169);
the driver placement is Halldrix's (#102118). Reported in #102056 / #103044 (Vksh07).
2026-09-06 07:20:24 -07:00
Teknium
03e9f4caee Merge pull request #104284 from NousResearch/fix/nous-anthropic-chat-wire
fix(nous): route anthropic/* over chat/completions by default; the native Portal route re-writes the cache on 14-20% of calls (nous.anthropic_wire)
2026-09-06 07:17:48 -07:00
Teknium
54e24ae1fa fix(nous): route anthropic/* over chat/completions by default (nous.anthropic_wire)
Portal serves anthropic/* on two routes. The native /v1/messages wire, which
Hermes has used since 02d5e23085, re-writes the previous turn's prompt cache
on 14-20% of consecutive calls in concurrent tool loops; the chat/completions
route does not. Measured 2026-09-06, 20 concurrent sessions x 6 tool calls on
Fable 5.1, same account, same hour, same first-party pin:

  Nous /v1/messages          180 pairs, 25 stuck (13.9%); 3 earlier 40-session
                             runs 15.0 / 17.3 / 20.2%; unchanged by the pin
  Nous /v1/chat/completions  320 pairs (2 runs), 0 stuck
  OpenRouter direct, pinned  161 pairs, 0 stuck

"stuck" = cache_read on call k+1 equals cache_read on call k instead of
call k's prompt size: the last write was not visible, the turn (~28K) was
written twice. Every stuck pair had byte-identical system, tools and
per-message shas, so it is not client-side mutation. At $10-20/M for
cache writes that is 15-20% of a fan-out's write bill.

The cause is inside the portal's native route (every element of its
outgoing request reproduces clean from outside; NousResearch/api#227
carries the diagnostics). Until it is fixed, anthropic/* rides
chat/completions. nous.anthropic_wire: native opts back in. Cost of
chat: prior-turn thinking travels as OpenAI-style reasoning fields, and
cache_control scopes are translated by the portal's adapter.

Tests: new test_nous_anthropic_wire_default.py reads the knob through the
real config loader in a temp HERMES_HOME and through
resolve_runtime_provider (both settings); the existing native-wire
contract suite selects native via an autouse fixture so the wire keeps
working for the flip-back. Live: a real AIAgent on this branch,
provider=nous + anthropic/claude-fable-5.1, dispatches to
chat/completions with the session_id intact.
2026-09-06 05:45:44 -07:00
tylman
b0adce1cbf feat(models): add gemini-3.7-flash and gemini-3.8-flash support
- Generalize Gemini 3 thinking config model prefix match to gemini-3*
- Add pricing snapshot entries for gemini-3.7-flash and gemini-3.8-flash
- Add model identifiers to Google, OpenRouter, and Vertex CLI lists
- Add test coverage for gemini-3.8-flash thinkingLevel configuration
2026-09-06 05:42:22 -07:00
Teknium
91b4d9fb67 test(fallback): keep the two persisted-config invariants; PTY harness for the picker error path
Trim the salvaged suite to the behaviour contracts: a picker exception or Ctrl+C leaves
config.yaml's model exactly as snapshotted (parametrized), and an absent active_provider
stays absent through snapshot+restore. The mock-dispatch tests (asserting calls into our own
helpers) are dropped. evals/cli_fallback_add_picker_error.py drives the real
`hermes fallback add` under a Linux PTY into an ordinary picker OSError and checks the
persisted model: stranded on base, restored after.
2026-09-06 05:41:46 -07:00
Xipong
33cefdd8e5 fix(fallback): restore primary route after picker errors 2026-09-06 05:41:46 -07:00
Teknium
3a15c39e8e fix(banner): deferred update notice keeps prompt_toolkit routing; drop dead console arg; PTY harness
Follow-up on the cherry-picked fix: _defer_update_notice no longer takes the Rich console it
can never safely use (the notice always lands after patch_stdout owns stdout), the unit test
asserts the contract at prompt_toolkit's boundary (an ANSI fragment reaches
print_formatted_text, with no raw ESC/markup in the visible text) instead of patching our own
cprint, and evals/cli_deferred_notice.py drives the real CLI under a Linux PTY with a FIFO-gated
update cache to A/B the late-notice cases (garbled on base, clean after).
2026-09-06 05:36:26 -07:00
turingcat
a366a0bb17 fix: route deferred update notice through cprint to avoid ANSI garbling
The deferred update notice ("N commits behind") was rendered via
console.print which writes ANSI escapes directly to stdout. Under
patch_stdout (active in the interactive CLI), the StdoutProxy strips
ESC bytes, leaving visible [1;33m...[0m artifacts.

Fix: render Rich markup to an ANSI string via a captured Console, then
route through cprint (which uses prompt_toolkit ANSI parser) to
bypass the StdoutProxy mangling.

This is the same class of bug as #2262 (fixed by #2448 for agent._print_fn
and display.py), but the _defer_update_notice callsite in banner.py was
not covered by that fix.

Fixes #83969
2026-09-06 05:36:26 -07:00
Teknium
0edab1ac4d test(profiles): two invariants for the truncated-JSON refusal, comment trimmed to the why
Parametrized refusal test (bare/fenced/uppercase-fenced truncated objects) asserts
profile.yaml is byte-identical afterwards; one prose test keeps the lenient
fallback contract. Replaces five near-duplicate tests from #104075.

Campaign tracker: https://github.com/NousResearch/hermes-agent/issues/104154
2026-09-06 05:34:01 -07:00
Edizzier
1f16bbf108 fix(cli): refuse to persist a truncated JSON reply as profile description
hermes_cli/profile_describer.py's aux-LLM response parser is documented as
"lenient, never raises": when the reply doesn't parse as JSON via
_extract_json_blob, it falls back to treating the WHOLE raw reply as plain
prose and persists it (truncated to 280 chars) as profile.yaml's
description.

That fallback exists for models that ignore the JSON-output instruction and
just answer in prose -- reasonable. But it doesn't distinguish that case
from a reply that DID start out as the requested JSON object and was cut
off mid-object by the aux model or transport. _extract_json_blob requires a
matching closing brace, so a truncated object (no `}` at all, or one cut
off mid-string) returns None just like plain prose does, and the fallback
then persists the raw JSON fragment verbatim: literal leading `{`, `\n`
escapes, a sentence chopped mid-word (#104067).

Fix: when parsing fails, check whether the reply (after the existing
code-fence strip) still looks JSON-shaped (starts with `{`). If so, it's a
malformed/truncated structured reply, not prose -- refuse it and return
ok=False instead of persisting the fragment. Only fall through to the
raw-text-as-prose fallback when the reply never looked like JSON to begin
with, so genuinely prose-only aux replies keep working exactly as before.
Also logs one INFO line on the refusal path, per the issue's note that the
silent fallback made this undiagnosable.

Update (review feedback from @jerrygooch): profile_describer._FENCE_RE was
case-sensitive, so an uppercase ```JSON fence wasn't stripped -- stripping
left "JSON\n{..." which doesn't start with "{", so a truncated response
under that fence variant still fell through to the prose fallback and got
persisted, defeating the fix for that case. Added re.IGNORECASE to
_FENCE_RE, aligning it with kanban_specify._FENCE_RE (which already used
re.IGNORECASE), and added a regression test for the uppercase-fence case.

Added tests/hermes_cli/test_profile_describer.py coverage: a truncated
object cut off mid-sentence (the real-world shape from the issue), one cut
off right after the opening brace, the same truncation under a ```json
fence, the same truncation under an uppercase ```JSON fence, and a
regression guard that a genuine plain-prose reply still hits the existing
lenient fallback unchanged.

Fixes #104067

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LUtGTop5iGKi5vfWnKnwET
2026-09-06 05:34:01 -07:00
kshitijk4poor
bc1330eebc refactor(update): name the success invariant in _receipt_looks_unfinished; one predicate contract
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
2026-09-06 14:53:33 +05:30
kshitijk4poor
dad698d88c test(update): pin the refused-receipt shape the stop_reason clause exists for
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).
2026-09-06 14:53:33 +05:30
tachyon-r
6e0fba4e58 refactor(update): share command-boundary receipt reason 2026-09-06 14:53:33 +05:30