Commit Graph

686 Commits

Author SHA1 Message Date
teknium1
e56ec8c9f7 fix(chronos): a 403 invalid_client from NAS hands cron fires to the built-in ticker (#97494)
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.

NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.

Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
2026-09-18 20:07:21 -07:00
teknium1
255b4fd9da fix(image-gen): record token usage as soon as the billed HTTP 200 lands
OpenRouter (_generate_via_chat, _generate_via_image_api) and OpenAI gpt-image
called record_token_usage only after image extraction and save succeeded. A
token-billed 200 that returned text but no image (the "returned no image"
fallback case), an empty `data` list, or a failed save therefore consumed
billed tokens that never reached session_model_usage; on the fallback chain
only the fallback model's tokens were recorded.

Move the call to immediately after a successful post_json / SDK call, before
extraction/save, in all three paths. One call per HTTP success, so no double
count; failure responses (non-2xx, timeouts) still record nothing because
there is no body to read usage from.

Tests: the parametrized invariant in test_openrouter_compat_provider gains the
no-image-with-usage case for both surfaces (result is `empty_response`, one row
still recorded); the OpenAI test is parametrized on has_image the same way.
2026-09-18 19:56:51 -07:00
teknium1
6babdc96b8 fix(image-gen): every token-billed image backend records session usage (chat path, Image API, OpenAI gpt-image)
Widen the salvaged Image API hunk (#114340) to the whole class:

- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
  aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
  "image_generation", the billing provider and the priced model id. Dict and
  SDK usage objects alike; a body without tokens is a no-op, as is a call
  outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
  chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
  token-billed) now records too; the Image API path uses the shared helper
  (the contributor's local _record_image_api_usage is folded into it) and both
  pass base_url so pricing resolves the route. Task renamed image_gen ->
  image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
  Images API usage block under the API model (gpt-image-2), not the Hermes
  quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
  unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
  over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.

Fixes #114324
2026-09-18 19:56:51 -07:00
Yagna Vudathu
98ad0d8148 fix(image-gen): record token-billed OpenRouter calls to session usage
OpenRouter token-billed image calls parsed usage into extra only, so
session_model_usage never saw them. Record prompt/completion tokens via the
ambient aux accounting on success; flat-fee responses without token usage
write nothing.

Fixes #114324.
2026-09-18 19:56:51 -07:00
teknium1
74c7515625 fix(tests): install the hindsight_client_api fake only when the SDK is absent
The autouse fixture shadowed an installed hindsight-client SDK with a
fake module (no __path__), so 'import hindsight_client' failed and
test_retain_timestamp_is_serialized_by_pinned_client skipped permanently
in venvs that have the extra. Gate the fake on
importlib.util.find_spec('hindsight_client_api') is None.
2026-09-18 10:37:36 -07:00
KoNit-K
42cf2e5f78 test: skip optional SDK test dependencies 2026-09-18 10:37:36 -07:00
KoNit-K
08946095d0 fix(kanban): preserve shared attachment blobs 2026-09-18 10:34:31 -07:00
teknium1
b44a481334 fix(video-gen): trust the operator-configured origin on the first download hop
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.

save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.

security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
2026-09-18 10:22:14 -07:00
teknium1
ea05583ad5 test: one class invariant across every remote-fetch site
Replace the per-site refusal tests with a single parametrized invariant
driven through a real loopback listener: each of the nine call sites
refuses a loopback target with zero connections recorded, and a
safe-looking first hop that 302s to the metadata endpoint is stopped at
the hop with nothing cached. save_url's redirect contract (caller headers
first-hop only, missing-Location 3xx fails closed) stays in
test_provider_media. Existing save_url tests keep their httpx/allow-
private adaptations; the codex edit test keeps its guarded-client stub.
2026-09-18 10:22:14 -07:00
beardthelion
3933fdf63b fix(security): route remote-party-supplied URL fetches through the SSRF guard
Provider response URLs, model-supplied image refs, manifest-derived pet
URLs, and remote sitemap <loc> entries were fetched with raw
requests/httpx/urllib — bypassing tools/url_safety while every platform
media path already uses it. A hostile or compromised provider/manifest
endpoint could steer a server-side fetch at internal or metadata
addresses; several sites cache the body where it is deliverable back.

Apply the canonical is_safe_url + create_ssrf_safe_client pattern at
every site: per-hop revalidation at TCP connect (closing the
DNS-rebinding window), bounded redirect chains that fail closed on
missing Location, and caller headers scoped to the first hop only —
matching the openrouter provider's own documented contract that its
bearer key must never leave the operator-selected host.

Operator-configured endpoints and pinned release assets are out of
scope — those URLs are operator-selected, not remote-party-controlled.

Fixes #114468
Closes #44728

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Ray <rayjun0412@gmail.com>
Co-authored-by: zapabob <1920071390@campus.ouj.ac.jp>
2026-09-18 10:22:14 -07:00
teknium1
c2bf40d1ce fix(kanban): docs say a bound Project keeps the directory when the field is cleared; UI test pins only the unbind seam
With a bound project the settings dialog omits a blank directory from the PATCH and rename_board mirrors the project's primary folder, so clearing the field cannot fall back to scratch until the project is unbound; the Settings bullet now says so. The bundle test drops five wording-level substring asserts that pinned the payload shape already covered behaviourally by test_kanban_board_project_api.py, keeping only the switcher's unbind seam.
2026-09-18 10:19:54 -07:00
teknium1
0985aeb732 fix(kanban): board switcher shows the bound project with an unbind action
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.

Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.

Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
2026-09-18 10:19:54 -07:00
liuhao1024
bc3da64fe8 fix(kanban): expose the board's project binding in the dashboard dialogs
The REST API already accepts project_id on board create/update and
GET /boards annotates project_id/project_name, but the dashboard UI
never wired any of it: a board's project binding could only be set
through raw API calls. Add a project selector to the New board and
Board settings dialogs, populated from GET /projects, mirroring the
existing default_workdir wiring. Settings PATCHes send project_id
unconditionally when the selector rendered ("" clears the binding,
server-validated); when the projects store is unreachable the field
is omitted so saving unrelated settings cannot wipe a binding.

Fixes #114652
2026-09-18 10:19:54 -07:00
teknium1
3ffee76761 fix(gateway): race IPv6/IPv4 on every cold-start WebSocket dial
The bootstrap racer covers sync connects only. The gateway's WebSocket dials
(relay connector, Yuanbao, Buzz) go through ``websockets.connect`` →
``loop.create_connection``, whose ``happy_eyeballs_delay`` defaults to ``None``:
a serial walk that burns the full connect timeout on every blackholed AAAA
record before IPv4 answers — the same stall class #114265 reports, one layer up.

``websockets`` forwards unknown kwargs to ``loop.create_connection``, so each
call site passes ``happy_eyeballs_delay=0.25`` (the RFC 8305 delay anyio and
the sync racer already use). One invariant test per call site captures the
kwargs at a mocked ``websockets.connect``.
2026-09-18 09:56:28 -07:00
teknium1
ee9c0f1a32 fix(photon): evaluate the group mention gate before caching inline attachments
`_on_sidecar_message` ran `_normalize_content` (which base64-decodes and writes
inline attachment/voice bytes into the media cache via `_cache_inbound_attachment`)
before the `chat_type == "group" and self.require_mention` check, so an
unmentioned group attachment was persisted and then dropped. Gate on the
user-typed text extracted without touching attachment bytes (`_mention_gate_text`:
text / richlink / group text items), then normalise and strip the wake word only
for messages that pass.

Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group attachment -> 0 cache writes, mentioned -> 1 and dispatched.
2026-09-18 09:55:55 -07:00
teknium1
d7f2644141 fix(codex): profile declarations follow the host and never override OpenAI's own ladder
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes #114249 but has two
edges the transport must keep:

- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
  non-None custom declaration short-circuited the host lookup: a
  ``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
  Router catalog clamp and would send ``max`` to a gateway that 400s on it.
  Resolve by endpoint host first, then by name — the host is the endpoint's
  truth; the config-entry name is only a label (the existing Router test now
  uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
  codex_responses) is still OpenAI: its per-model ladder is known, so skip
  profile declarations on the official origin and the Codex backend.
  ``_is_openai_api_origin`` is the shared exact-host check;
  ``_is_official_openai_responses_route`` reuses it.

Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.

Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
2026-09-18 09:40:31 -07:00
liuhao1024
f111297af1 fix(custom): declare the OpenAI-compat effort vocabulary so Responses keeps a configured max
CustomProfile did not override supported_reasoning_efforts, so on the
Responses transport a custom:<name> relay fell through to the OpenAI
per-model ladder (codex_supported_efforts) and a configured effort=max
was silently clamped to xhigh — while the same provider over
chat-completions forwarded max unchanged, because that path already
clamps onto OPENAI_COMPAT_WIRE_EFFORTS.

Declare the same wire set from the profile so the two transports agree:
max survives, ultra still clamps down to max, and the official OpenAI
backend ladder (gpt-5.5 rejecting max) is untouched.

Fixes #114249
2026-09-18 09:40:31 -07:00
teknium1
06350c79d5 fix(kanban): dashboard done/review refusals name the open parents
complete_task and request_review return bare False when the dependency
gate refuses, and _patch_status only enriched the 409 for status=ready,
so PATCH status=done/review on a gated card answered "not valid from
current state" and the bulk entry said "transition refused". Consult
unsatisfied_parents on a refused done/review and name the parents in
the 409 detail and the bulk entry error, matching the CLI/tool wording.
2026-09-18 09:29:44 -07:00
teknium1
33f3d96b99 fix(disk-cleanup): never rmtree protected top-level dirs; never track kanban/
The tracked-item delete path in quick() called shutil.rmtree() on any tracked
directory without consulting the protection list, which only the empty-dir
sweep used. guess_category() files every path under cache/ as "temp", so a
terminal command that merely mentioned $HERMES_HOME/cache tracked the
directory itself, and 7 days later quick() removed cache/ wholesale — taking
cache/terminal (terminal snapshots) with it and breaking every later command
with a mktemp "No such file or directory".

Fix: _is_protected_dir() — a tracked DIRECTORY that is HERMES_HOME itself or
sits under an _EMPTY_DIR_PROTECTED_TOP_LEVEL tree is never tracked by
guess_category(), never listed by dry_run(), and skipped (logged SKIPPED,
entry dropped) by quick(). Files under cache/ still age out as before.

kanban/ (task attachments and workspaces have their own lifecycle) is added to
both _NEVER_TRACK_TOP_LEVEL and _EMPTY_DIR_PROTECTED_TOP_LEVEL; stale pre-fix
"test" entries under it are dropped by the existing re-validation instead of
deleted. The kanban row was first proposed in #80842 (@nicha16).

Fixes #114552
2026-09-18 09:22:13 -07:00
beardthelion
72a2dae25f test(tools,tui_gateway,cli,plugins): a parseable non-dict JSON file is skipped by the sibling scans
Per-subsystem invariants for the eight ported sites (relay outbox claim,
lightpanda reaper, write-approval pending store, spawn_tree list/load,
local-runtime manifest, A2A conversation replay, batch runner scans,
trajectory compressor entry), trimmed from PR #114241's test hunks to at most
two per subsystem (batch_runner's dataset-load and resume-scan cases collapsed
into one). Each was red on the previous head of this branch.

(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
2026-09-18 09:19:04 -07:00
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
7f03867523 fix(aux): parameter rungs hand a 429 on, and Fireworks gets top-level reasoning_effort
The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
2026-09-17 08:45:46 -07:00
Changhyun Min
98f758ae7e fix(upstage): default new Solar models to the 512K context window
Upstage /v1/models carries no context field, so Solar ids resolve through
DEFAULT_CONTEXT_LENGTHS, whose keys are substring matches. solar-mini4 hit
the legacy solar-mini key (32K, below the 64K minimum), so the agent refused
to start; solar-pro4 matched nothing and got the 256K fallback. The same
substring deny-list in the Upstage profile also treated solar-mini4 as
non-reasoning, silently dropping reasoning_effort for a model that reasons.

Solar keys now match only on an id boundary (key followed by end or one of
"-:.@"), after folding aggregator slugs such as solar-pro-3 into the native
solar-pro3. Legacy keys keep covering their dated, quant and variant ids
(solar-mini-250422, solar-pro3@q4) but not later generations. A "solar-"
family key gives Solar lineup ids (bare name "solar-<letter>...", org/
prefix allowed) 524288 tokens, per Upstage /v1/solar/models max_model_len;
open-weight ids like solar-10.7b-instruct keep the generic fallback, and an
explicit key still wins by longest-key-first order. The boundary-matched
prefixes are a module constant, so non-Solar keys are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 17:10:58 +05:30
teknium1
44657ff628 fix(achievements): apply the sticky-unlock floor to in-flight snapshots too
_compute_from_scan(is_partial=True) evaluated against an empty unlock
set, and _run_background_scan publishes those partials to
_SNAPSHOT_CACHE every progress_every sessions. So for the duration of
every background rescan an already-earned badge rendered as locked and
unlocked_count collapsed then climbed back — the exact 39<->40 flicker
from #112273 that the finished-scan floor alone did not remove.

Read state.json for partials as well (the floor applies); partials
still never record new unlocks or save_state, since an unlock time from
half a scan could be invalidated by a later session.

Also:
- tests/plugins fake SessionDB accepts include_compacted, which the
  scanner now passes (this was the red 'Python tests' check).
- engine tests patch get_hermes_home so _data_file's legacy migration
  cannot copy the developer's real state.json into the temp dir.

Part of #112273
2026-09-16 17:03:47 -07:00
teknium1
2db11f03a9 fix(tests): restore the trailing newline in test_disk_cleanup_plugin.py
The last hunk dropped the file's final newline, which shows up as a
'\ No newline at end of file' marker in every future diff touching the tail.
2026-09-16 16:54:19 -07:00
teknium1
0631710ae4 fix(disk-cleanup): protect every per-profile user tree and trim tests to two invariants
Extends the salvaged #112860 entry (`workspace`) to the whole class: `plans` and
`home` are bootstrapped alongside `workspace` by `hermes_cli/profiles.py::_PROFILE_DIRS`
and listed as user data in `profile_distribution.py::USER_OWNED_EXCLUDE`, so a
`test_*`/`tmp_*` file inside any of them is a user's file, never scratch.

Trims the contributor's five tests to two invariants: the end-to-end
post_tool_call -> on_session_end path (workspace file survives, a root-level
`tmp_scratch.py` control is still removed) and the empty-dir sweep leaving a
directory under `workspace/` alone. Documents the protected trees in the plugin
README and the built-in-plugins page.

Dropped: test_workspace_project_tree_never_tracked, test_quick_keeps_workspace_test_file,
test_root_level_test_file_still_auto_deleted (folded into the E2E test as its control).

Part of #112859.
2026-09-16 16:54:19 -07:00
kevin
3b6b0e24e4 fix(disk-cleanup): keep test_*/tmp_* files under $HERMES_HOME/workspace
``workspace/`` is a user-owned project tree: every profile bootstraps it
(``profiles.py::_PROFILE_DIRS``), ``profile_distribution.py`` excludes it as
user data, and bundled plugins keep durable state there (``google_meet``
writes auth/registry JSON under ``workspace/meetings/``).

It was missing from ``_NEVER_TRACK_TOP_LEVEL``, so ``guess_category()``
returned "test" for ``workspace/<project>/tests/test_*.py``;
``_is_auto_delete("test", age)`` accepts any age, so the ``on_session_end``
sweep unlinked those files seconds after pytest ran them green - 20 file
deletions and 476 empty-dir removals (chrome-profile data dirs among them)
over three days on one install. It was also missing from
``_EMPTY_DIR_PROTECTED_TOP_LEVEL``, so meaningful empty dirs inside a project
tree were swept too.

Add "workspace" to both sets, beside the sibling user project trees
(``projects``, ``patches``, ``skins``, ``themes``, ``contributors``) that
#75403 / #32164 / #37721 already protect. No ``tracked.json`` migration is
needed: ``quick()`` re-validates stored categories through
``guess_category()`` and drops mismatches.

Regression tests: 4 of the 5 new tests fail against the unpatched plugin,
including the end-to-end ``post_tool_call`` -> ``on_session_end`` path; the
fifth pins that genuine ephemeral test files (root-level ``test_*.py``) are
still cleaned.
2026-09-16 16:54:19 -07:00
teknium1
e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1
b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
KoNit-K
4fe3d288eb fix(providers): send DeepInfra reasoning effort 2026-09-15 18:22:40 -07:00
fangliquan
584fc99a53 test(photon): cover inbound NDJSON framing 2026-09-15 04:34:13 -07:00
fangliquan
992b517fdd fix(photon): preserve Unicode NDJSON separators 2026-09-15 04:34:13 -07:00
teknium1
021ab58a23 fix: hindsight update-time dep resolution survives a BOM in config.json
`_provider_pip_dependencies` still read ~/.hermes/hindsight/config.json with
strict utf-8 inside a bare `except Exception`, so a Windows-editor BOM made the
`mode` lookup silently fail and `hermes update` reinstalled only
`hindsight-client`, leaving the embedded daemon broken — the exact #70636
symptom this helper exists to prevent. Route it through the shared
`read_json_or_empty` (utf-8-sig, {} on missing/corrupt) like the other memory
readers in this PR, and add the reader to the parametrized BOM invariant.
2026-09-15 03:38:29 -07:00
teknium1
c97e36db3f test: one parametrized BOM invariant over the shared reader + 2 direct readers
mem0, hindsight and the honcho CLI all read through utils.read_json_or_empty,
so the seven per-loader tests collapse to one parametrized invariant plus the
Qwen creds case. The per-loader fix landed in the shared reader (one site, not
six), which is the salvage bar: <=2 invariant tests, no change-detectors.
2026-09-15 03:38:29 -07:00
Teknium
5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00
teknium1
d8053f4806 fix(video_gen): cap LTX 2.5 at 10s for 1440p/2160p; omit unset enum durations
fal's LTX 2.5 fast endpoints accept 6-20s only up to 1080p — "At 1440p and
2160p, all frame rates support up to 10 seconds" — so a 4K request with the
family's 20s ceiling was rejected by the vendor. Families can now declare
`duration_cap_by_resolution`, applied after the enum snap / range clamp on the
resolved resolution enum.

An unset duration on a duration_enum family also snapped to enum[0] (6s),
silently overriding the endpoint's own "auto" default; None now omits the key
for enum families exactly as it already did for range families.

test_managed_media_gateways asserts the alibaba/happy-horse/ namespace by
prefix rather than the exact v1.1 literal so the next version bump doesn't
flip an unrelated gateway test.
2026-09-15 03:37:11 -07:00
teknium1
1c1980dc1f fix(video_gen): make the duration enum explicit instead of sniffing tuple shape
`durations` carried two meanings told apart only by len==2 and gap>1: a
(min, max) range to clamp, or an enum to snap. A family with exactly two
legal values would have been misread as a range (review finding on #91311).
`durations` is now always the (min, max) window (what capabilities()/
list_models() read) and families with discrete values add `duration_enum`;
_clamp_duration takes the family and branches on the key, not the shape.

Also: restore the exact v1.1 endpoint assertion in the gateway namespace
test (a startswith/endswith check would not catch a silent version drift),
add the ltx-2.5 i2v snap case, and keep happy-horse on audio_native (the
schema test forbids audio+audio_native together, and v1.1 audio is always on).
2026-09-15 03:37:11 -07:00
Teknium
37286d3064 feat(video_gen): LTX 2.5 + Kling O3 families; Happy Horse upgraded to v1.1
Adds two new FAL video families and upgrades one:

- ltx-2.5 (cheap tier): lightricks/ltx-2.5/{text,image}-to-video/fast.
  Lightricks' open-source audio-video model. Native audio, 6-20s integer
  duration enum, 720p-2160p (i2v), $0.09/s at 720p. duration_int + 2k/4k
  resolution aliases; no seed key in the schema.
- kling-o3 (premium tier): fal-ai/kling-video/o3/standard/{text,image}-to-video.
  Kuaishou's frontier multi-shot model, 3-15s, optional native audio
  ($0.084/s off, $0.112/s on). String durations, i2v drops aspect_ratio,
  no seed/resolution keys.
- happy-horse upgraded from the sparse-docs 1.0 endpoints to
  alibaba/happy-horse/v1.1/{text,image}-to-video with the full published
  schema: nine aspect ratios, 720p/1080p, 3-15s integer durations, seed
  supported, audio native (no generate_audio key), i2v drops aspect_ratio.

All flags derived from each endpoint's llms.txt schema. Payload builder
asserted locally against the schemas; test for the old Happy Horse
"prompt-only" contract updated to pin the v1.1 schema, plus new payload
tests for ltx-2.5 and kling-o3.
2026-09-15 03:37:11 -07:00
kshitijk4poor
5de9ef7b89 test(supermemory): guards for the failure sentinel, buffer caps, and shutdown no-resend
- none-returning client counts as success (sentinel contract)
- pending buffer drops oldest past the turn cap and the byte cap
- the shutdown interleaving test now pre-seeds a pending turn so the
  no-resend assert discriminates on content, not timing: with the lock
  removed it fails on assert 2 == 1 (duplicate send), verified by mutation
2026-09-15 11:55:10 +05:30
kshitijk4poor
f7e0ae2aff test(supermemory): pin shutdown flush waits for in-flight write, never re-sends
Answers the open review ask on #109359 with the corrected contract: a hung
remote add_memory is bounded by the SDK timeout (empirically 1.03s at
timeout=1.0, max_retries=0), so shutdown's flush waiting on the capture
lock is bounded, not forever. The guard: while the worker owns the write
(blocked inside add_memory), a concurrent shutdown must wait and must not
re-send the same pending batch. Mutation-checked: lock -> nullcontext
makes this test and the session-switch interleaving test fail.
2026-09-15 11:55:10 +05:30
kshitijk4poor
69d69766dd test(supermemory): freeze the capture clock for custom-id assertions
Eight tests compare a recorded custom_id against a freshly computed
_capture_custom_id; near a 4h-bucket edge the provider's write and the
test's expectation read now() in different buckets and the equality
fails. The frozen_capture_clock fixture pins the module datetime so both
reads are identical by construction.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
c6da9e0788 fix(supermemory): serialize capture writes and document at-least-once retry
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.

Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
1b76cfff83 fix(supermemory): keep pending turns across session switch until written
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
03627dbf95 fix(supermemory): write turns via documents API instead of session-end conversations ingest
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).

Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.

Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.

Metadata stays type/session_id/timestamp plus the existing sm_source.
2026-09-15 11:55:10 +05:30
teknium1
3275ca88ec fix(image_gen): Codex-auth images use the native images endpoints, no chat host model
The openai-codex image provider rode a Responses call with a hosted
image_generation tool on a pinned chat model (gpt-5.5). Two failure
classes came with that shape: when OpenAI withdrew gpt-5.5 from an
account cohort every image call 404'd while chat kept working
(#105398, #107076), and the host model was free to answer in text
instead of calling the tool, so we streamed SSE, kept partial frames
and retried on empty streams.

Post to chatgpt.com/backend-api/codex/images/generations and
images/edits instead - the route the official Codex client uses
(codex-rs/ext/image-generation). No host model, no SSE, no
partial-frame handling; the response is a plain JSON body with
b64_json. Remote source URLs are fetched client-side and inlined as
data URLs because the backend's own downloader 400s on ordinary
public images.

The backend treats model/quality/size as advisory (#107233), so the
result now reports reported_quality/reported_size next to the
requested values plus the x-codex-imagegen-request-id for support.
GPT Image 2.5 is deliberately not added to this catalog: the backend
accepts any model id, including nonexistent ones, and generates with
its server-managed engine (C2PA reports gpt-image 2.0), so a 2.5 tier
here would be a label with no effect (#106708).
2026-09-14 10:18:13 -07:00
teknium1
1ce2d25237 test: achievements scan attaches read-only; read-only handles never trip the duplicate-writer warning
Both invariants fail on origin/main. Test doubles that stood in for
SessionDB() gain the read_only kwarg the production call sites now pass.
2026-09-14 08:10:36 -07:00
teknium1
01338e88ec test(video-gen): keep the i2v-only guard real once every family is dual-modality
The rebased synthetic guard built a bare dict that main's _build_payload
no longer accepts (it indexes aspect_ratios/resolutions directly) and never
exercised generate(); it now builds the family via _family() and asserts
generate() returns modality_unsupported without submitting. The surface
matrix's i2v-only parametrization collected zero cases after Gemini Omni
Flash 1.1 gained t2v, so it is removed along with the dead branch in the
t2v matrix.
2026-09-13 21:40:26 -07:00
Teknium
aecee6f66a feat(video): Gemini Omni Flash 1.1 — text-to-video now works (was image-only)
FAL shipped google/gemini-omni-flash/v1.1/* (Aug 2026): the family gains a
text-to-video endpoint, 360p/720p/1080p/4k resolution enum, and keeps
3-10s integer durations with always-on native audio.

- plugins/video_gen/fal: bump gemini-omni-flash to the v1.1 endpoints,
  declare the resolution enum, refresh display/strengths copy
- tests: family test asserts the versioned dual-modality endpoints; the
  i2v-only clean-error guard survives via a synthetic family; catalog
  invariant now requires both endpoints on every family

Schema verified against FAL OpenAPI (t2v: prompt required, 16:9/9:16,
360p-4k, duration 3-10 int; i2v adds image_url + optional end_image_url).
Pricing: $0.03/s 360p, $0.10/s 720p, $0.15/s 1080p, $0.30/s 4K.
2026-09-13 21:40:26 -07:00
teknium1
2ff55bc895 test(video-gen): pin the Wan 3.0 audio toggle key and i2v image key
audio_param_key is a new _build_payload seam; one invariant covers it (Wan
emits `audio`, veo3.1 still emits `generate_audio`) plus start_image_url
and the integer duration.
2026-09-13 21:28:11 -07:00