24 Commits

Author SHA1 Message Date
ethernet
c13ea774e6 refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.

Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).

Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
2026-09-23 11:41:01 -04:00
teknium1
627a12a90e feat(opencode-go): show Go plan windows in /usage through the profile hook
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-19 20:52:57 -07:00
teknium1
f678ed8299 refactor(model-providers): thinking-toggle XOR effort translation lives in agent.reasoning_effort; opencode-free imports it instead of borrowing via sys.modules
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
2026-09-13 05:19:48 -07:00
Ada
7c734c7838 fix(opencode-go): send reasoning for the canonical deepseek-flash id
The version-less canonical Flash id is not matched by _is_deepseek_thinking_model, so agent.reasoning_effort and every auxiliary reasoning_effort were silently dropped on the OpenCode Go relay while the direct provider was fixed (8435a3ae00/aeecb110f8). Match it through a version-less id set, mirroring plugins/model-providers/deepseek. Verified against the Go relay: named levels are graded (low 906 / high 1371 / max >=2500 reasoning tokens on a multi-step prompt) and integer efforts are rejected, so the named level is the knob to send.
2026-09-12 08:06:48 -07:00
ericmaddox
bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Yuan Chenglu (袁成路)
6a0f519d1e fix(opencode-go): set supports_vision_tool_messages=False for Xiaomi MiMo backend
## Problem

When using the opencode-go provider with Xiaomi MiMo models (e.g.
mimo-v2.5, mimo-v2.5-pro), the Hermes agent intermittently fails with:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

This occurs specifically when tool results contain multipart content
with image_url parts (e.g. browser screenshots). The opencode-go relay
forwards these as-is to the Xiaomi MiMo backend, which rejects list-type
tool message content while still accepting multimodal user messages.

## Root Cause

The OpenCodeGoProfile inherits supports_vision_tool_messages=True from
ProviderProfile (the default). When this flag is True, the agent sends
tool results with image parts directly to the model. However, Xiaomi
MiMo's API rejects this format:

> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."
>   — providers/base.py, line 73

The direct 'xiaomi' provider profile already correctly sets this to
False (plugins/model-providers/xiaomi/__init__.py, line 13), but the
opencode-go relay profile was missing this safeguard.

The relevant code path is in run_agent.py:_tool_result_content_for_active_model()
(line 4543), which checks _provider_supports_vision_tool_messages() when
deciding whether to embed images in tool-result messages.

## Fix

Add supports_vision_tool_messages=False to the OpenCodeGoProfile
instantiation in plugins/model-providers/opencode-zen/__init__.py.

This single-line change prevents tool-result images from being sent
as multipart content to the MiMo backend, while preserving the model's
image recognition capability through user messages and vision tool
invocations (both of which use different code paths unaffected by this
flag).

## Testing

Verified with the mimo-v2.5 model via opencode-go provider:

1. Browser tool + screenshot recognition
   → Navigated to https://www.baidu.com, took screenshot, identified
     top 3 trending topics from the image
   → Result: PASSED, recognized all topics correctly

2. Direct image as user message
   → Sent a screenshot PNG directly via --image flag, asked model to
     describe the content
   → Result: PASSED, model correctly read text from the image

3. Provider profile verification
   → Confirmed get_provider_profile('opencode-go').supports_vision_tool_messages
     returns False at runtime
   → Result: PASSED

4. No regression on non-MiMo models
   → opencode-zen provider retains supports_vision_tool_messages=True
     (unaffected)

---

fix(opencode-go): 为 Xiaomi MiMo 后端设置 supports_vision_tool_messages=False

## 问题描述

使用 opencode-go provider 搭配 Xiaomi MiMo 模型(如 mimo-v2.5、
mimo-v2.5-pro)时,Hermes agent 间歇性地抛出以下错误:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

该错误发生在工具返回结果包含 image_url 类型的 multipart 内容的场景下
(如浏览器截图)。opencode-go 中继层将这些内容原样转发给 Xiaomi MiMo
后端,而 MiMo 接受多模态用户消息,但拒绝 list-type tool message 内容。

## 根因分析

OpenCodeGoProfile 继承了 ProviderProfile 的默认值
supports_vision_tool_messages=True。当此标志为 True 时,agent 会将含
图片的工具结果直接发送给模型。但 Xiaomi MiMo API 拒绝此格式:

> providers/base.py 第 73 行注释明确指出:
> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."

直接的 'xiaomi' provider profile 已正确设置了该值为 False
(plugins/model-providers/xiaomi/__init__.py 第 13 行),但
opencode-go 中继 profile 遗漏了这一安全设置。

相关代码路径:run_agent.py 的 _tool_result_content_for_active_model()
方法(第 4543 行),该方法通过检查
_provider_supports_vision_tool_messages() 来决定是否在 tool-result
消息中嵌入图片。

## 修复方案

在 plugins/model-providers/opencode-zen/__init__.py 的
OpenCodeGoProfile 实例化中添加 supports_vision_tool_messages=False。

这一行改动阻止了 tool-result 图片以 multipart 格式发送给 MiMo 后端,
同时通过用户消息和 vision tool 调用的路径(使用不同代码路径,不受
此标志影响)保留了模型的图像识别能力。

## 测试验证

使用 mimo-v2.5 模型通过 opencode-go provider 验证:

1. 浏览器截图 + 图像识别
   → 导航至 https://www.baidu.com,截取首页截图,从图片中识别出
     热搜榜前三条
   → 结果:通过,正确识别所有热搜话题

2. 用户消息直接传图
   → 通过 --image 参数直接发送截图 PNG,要求模型描述图片内容
   → 结果:通过,模型正确读取图片中的文字

3. Provider profile 运行时验证
   → 确认 get_provider_profile('opencode-go')
     .supports_vision_tool_messages 在运行时返回 False
   → 结果:通过

4. 非 MiMo 模型无回归
   → opencode-zen provider 保持 supports_vision_tool_messages=True
     不受影响

## 修改文件

  plugins/model-providers/opencode-zen/__init__.py (+5 lines)

Signed-off-by: Yuan Chenglu (袁成路) <ycl_pj@163.com>
2026-09-09 03:52:47 -07:00
Teknium
48da317211 refactor(model-providers): tighten router/openrouter bodies, drop __future__ annotations (py>=3.11) 2026-09-02 21:58:26 -07:00
Teknium
40722b322c refactor(model-providers): pack keyword-only profile constructors (comments and multi-line values keep their lines) 2026-09-02 21:45:13 -07:00
Teknium
9338e21093 refactor(model-providers): compact reasoning-translation profiles (opencode, kimi, zai, minimax, nous, qwen, custom, upstage, nebius, meta-ai, deepseek, copilot, deepinfra, vertex, gemini), fold catalog tuples 2026-09-02 21:36:03 -07:00
Teknium
313b244e8f refactor(plugins/model-providers): reuse agent.reasoning_effort clamps, per-model dict tables, compact profiles 2026-09-02 13:30:28 -07:00
Teknium
d4d04098a5 fix: Ox Alpha reasoning effort reaches the wire clamped — shared across zen and free providers
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.

Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
2026-08-21 03:04:06 -07:00
vinsew
54227416ce fix(opencode): send Ox Alpha reasoning effort through Zen
OpenCode documents x-preview-f-free as accepting low, high, and max reasoning effort on its Zen Chat Completions endpoint. Hermes previously resolved the user's per-model override to max but the plain Zen provider profile discarded it, so successful calls silently ran at the server default.

Introduce an OpenCodeZenProfile scoped only to x-preview-f-free. It forwards the normalized top-level reasoning_effort, maps xhigh to max, preserves server defaults when unset or disabled, and leaves every other Zen model untouched.

Add profile and full transport tests that prove max reaches the outgoing request and that non-target models are unaffected. Also correct the nanoid security-pin comment to match the already-locked 3.3.18 release.
2026-08-21 03:04:06 -07:00
Teknium
19c6a11924 feat(opencode): sync Zen/Go catalogs — Ox Alpha stealth model, Grok routing, new free tier
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
  listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
  flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
  hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
  muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
  glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
  endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.

Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
2026-08-20 19:41:03 -07:00
Abdulkadir Ateş
5969ea1558 fix(opencode): route Muse Spark through the Responses API
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.

- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
2026-08-20 19:41:03 -07:00
Teknium
f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Adolanium
7fa084f58e fix: send Hermes Agent attribution headers to OpenCode Zen and Go
OpenCode identifies clients by request headers, the same way OpenRouter
does. Our opencode-zen and opencode-go profiles never set any, so every
request went out with the OpenAI SDK default "OpenAI/Python x.y.z"
User-Agent and OpenCode had no way to tell the traffic was Hermes Agent.

Two changes:

- Add HTTP-Referer, X-Title, and a HermesAgent User-Agent to both
  OpenCode profiles through profile.default_headers, the same path
  Fireworks uses. This covers chat_completions, codex_responses,
  auxiliary clients, model switches, and the models catalog fetch.
- Merge the same headers in build_anthropic_client for opencode.ai
  base URLs. The Anthropic Messages route (Claude on Zen, MiniMax and
  Qwen on Go) builds its client there and never sees profile headers.

Verified against the live Go relay with a real key. Both wire formats
return HTTP 200 and the requests now carry X-Title "Hermes Agent",
HTTP-Referer, and User-Agent HermesAgent/0.20.0.
2026-08-13 02:03:40 -07:00
kshitij
b3344502f8 chore: map Axmr1 email + update stale Go routing docstring
- Add contributors/emails/Axmr1@users.noreply.github.com for CI
  attribution check (bare noreply format needs explicit mapping)
- Update opencode-zen plugin docstring: Go routing now includes
  GPT → codex_responses and Qwen → anthropic_messages (was stale,
  only listed MiniMax and GLM/Kimi)
2026-08-08 14:47:11 +05:30
Teknium
7550c594ce feat(reasoning): add max and ultra effort levels (#62650) 2026-07-12 00:26:49 -07:00
Teknium
a6079dd350 feat(providers): GLM-5.2 native reasoning_effort controls (#58884)
Port from Kilo-Org/kilocode#11555: GLM-5.2 exposes a native
reasoning_effort knob with two enabled levels (high / max) on its
OpenAI-compatible endpoints. Previously the zai profile (direct Z.AI
/api/paas/v4) used the base ProviderProfile and emitted nothing, and the
OpenCode Go profile only handled Kimi K2 / DeepSeek — so a user's effort
preference for GLM-5.2 was silently dropped on both routes.

- zai: ZaiProfile maps effort onto high/max (xhigh/max -> max, lower -> high)
- opencode-go: same mapping for GLM-5.2, alongside existing Kimi/DeepSeek
- alias spellings recognized (glm-5.2 / glm-5-2 / glm-5p2, vendor-prefixed)
- disabled / no effort leaves the server default untouched
2026-07-05 13:48:01 -07:00
teknium1
03392b67d6 fix(opencode-go): gate thinking when reasoning_effort set to avoid HTTP 400
Salvaged from #40429; re-verified on main, tightened, tested.

Co-authored-by: jimjsong <jimjsong@users.noreply.github.com>
2026-06-07 01:24:29 -07:00
teknium1
8cf6b3da9d fix(opencode-go): cap mimo-v2.5-pro max_tokens at 131072
The opencode-go relay defaults max_tokens to 262144 when none is sent,
but Xiami mimo-v2.5-pro only supports 131072 completion tokens — every
request 400s with "max_tokens is too large: 262144" before the agent
can do anything.

Add a get_max_tokens(model) hook on ProviderProfile (default returns
default_max_tokens) so profiles fronting multiple upstreams can vary
the cap per-model. Wire chat_completions transport through the hook.
Override on OpenCodeGoProfile with mimo-v2.5-pro=131072.

Only mimo-v2.5-pro is capped — other opencode-go models (kimi, glm,
qwen, minimax, other mimo variants) unchanged.
2026-05-28 20:49:53 -07:00
teknium1
70aaa774be fix(opencode-go): emit Kimi reasoning_effort, match KimiProfile shape
The Kimi K2 branch added in the prior commit only emitted extra_body.thinking
and dropped reasoning_effort entirely. KimiProfile (api.moonshot.ai/v1) sends
both fields, and OpenCode Go proxies to the same Moonshot backend. Mirror that
shape on the Go path so /reasoning effort actually reaches Kimi.

- low/medium/high pass through verbatim
- xhigh/max clamp to high (Moonshot's max supported value)
- minimal / unknown effort → omit reasoning_effort, keep thinking on
- disabled / no config → unchanged
- DeepSeek branch unchanged
2026-05-23 02:20:28 -07:00
Harish Kukreja
3589960e03 fix(provider): expose OpenCode Go reasoning controls 2026-05-23 02:20:28 -07:00
Teknium
9022804d78 feat(providers): make all 33 providers pluggable under plugins/model-providers/
Every provider profile is now a self-contained plugin under
plugins/model-providers/<name>/, mirroring the plugins/platforms/
pattern established for IRC and Teams. The ProviderProfile ABC
stays in providers/; the per-provider profile data moves out.

- plugins/model-providers/<name>/__init__.py calls register_provider()
- plugins/model-providers/<name>/plugin.yaml declares kind: model-provider
- providers/__init__.py._discover_providers() lazily scans bundled plugins
  then $HERMES_HOME/plugins/model-providers/<name>/ (user override path)
- User plugins with the same name override bundled ones (last-writer-wins
  in register_provider)
- Legacy providers/<name>.py layout still supported for back-compat with
  out-of-tree editable installs
- Hermes PluginManager: new kind=model-provider; skipped like memory
  plugins (providers/ discovery owns them); standalone plugins with
  register_provider+ProviderProfile in their __init__.py auto-coerce to
  this kind (same heuristic as memory providers)
- skip_names extended to include 'model-providers' so the general
  PluginManager doesn't double-scan the category
- 4 new tests in tests/providers/test_plugin_discovery.py covering
  bundled discovery, user override, and general-loader isolation
- Docs updated: website/docs/developer-guide/adding-providers.md,
  provider-runtime.md, providers/README.md, plugins/model-providers/README.md

No API break: auth.py / config.py / doctor.py / models.py / runtime_provider.py /
model_metadata.py / auxiliary_client.py / chat_completions.py / run_agent.py
all still consume providers via get_provider_profile() / list_providers() —
they just now see plugin-discovered entries instead of pkgutil-iterated ones.

Third parties can now drop a single directory into
~/.hermes/plugins/model-providers/<name>/ to add or override an inference
provider without touching the repo.
2026-05-05 13:40:01 -07:00