Wiring response_format into generationConfig exposed a regression main never
had: any aux call that sends tools with tool_choice=auto AND a json_schema
format (a proxy/relay shape, or a caller that declares tools and asks for
JSON) now 400s on gemini-2.5-* with "Function calling with a response mime
type: 'application/json' is unsupported", where main silently ignored the
format and answered. Only Gemini 3+ combines function declarations with
structured output (ai.google.dev/gemini-api/docs/structured-output), so the
JSON keys are dropped whenever tools are declared on a pre-Gemini-3 model,
reusing the is_gemini3 flag build_gemini_request already computes.
Tests: the seven salvaged tests are trimmed to three invariants. The
former "keeps json output when tool_choice is auto" case asserted the
regressed behaviour (model="" is pre-Gemini-3) and is replaced by the
parametrized gemini-2.5 (drops) / gemini-3 (keeps) pair; the two schema
surfaces collapse into one parametrized test; the json_object-only and
missing-schema-key fallbacks were change-detectors on the translator.
Live (gemini-2.5-flash): tools + json_schema -> HTTP 400 on the PR head,
finish=stop on this commit; json_schema without tools still returns JSON.
Translates OpenAI-style response_format into Gemini's generationConfig as
responseMimeType plus responseJsonSchema on v1beta or responseSchema elsewhere,
reusing the existing tool-schema preparation so $ref inlining and key stripping
stay consistent. The translation is skipped when tool_choice forces function
calling, since Gemini rejects mode ANY combined with a JSON response type.
`_usage_from_metadata` (agent/gemini_native_adapter.py:476) read
`candidatesTokenCount` for `completion_tokens` and never read
`thoughtsTokenCount`; a repo-wide grep found that field nowhere. Gemini
reports hidden thinking in its own counter — `candidatesTokenCount` covers
visible output only, while `totalTokenCount` already includes thoughts. So on
any thinking turn the emitted usage contradicted itself: prompt + completion
did not add up to total.
Both the non-streaming assembly (translate_gemini_response) and the streaming
one (the finish chunk in translate_stream_event) share this helper, so both
under-billed identically. `normalize_usage` takes `completion_tokens` for
`output_tokens` and looks for reasoning under
`completion_tokens_details.reasoning_tokens`, which the adapter never set, so
the reasoning column read 0 for models whose spend is mostly reasoning.
Thoughts are now folded into `completion_tokens` (OpenAI's counter includes
reasoning) and surfaced under `completion_tokens_details.reasoning_tokens`,
matching the nesting the adapter already uses for
`prompt_tokens_details.cached_tokens`. The field is absent on non-thinking and
older responses; those count 0 and their numbers do not move.
Live: for usageMetadata {prompt 10, candidates 200, thoughts 5000, total 5210}
the adapter emitted completion_tokens=200 (10 + 200 != 5210) and
normalize_usage returned output_tokens=200, reasoning_tokens=0. It now emits
completion_tokens=5200 (10 + 5200 == 5210) with
completion_tokens_details.reasoning_tokens=5000, and normalize_usage returns
output_tokens=5200, reasoning_tokens=5000.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
Google now issues AQ. keys for both Google AI Studio and Vertex AI
express mode, so the AQ. prefix no longer identifies the key family
(#115306): auto-routing every AQ. key to aiplatform.googleapis.com 403s
the whole AI Studio fleet (6293fca019 / df53cae72d).
- normalize_gemini_base_url no longer rewrites by key shape; an express
key reaches aiplatform only through an explicitly configured base,
which is still completed to the publishers/google form (#114335 path)
- gemini_http_error appends two-way 403 PERMISSION_DENIED guidance: an
AQ. key rejected on the Studio host learns about the express base_url,
a key rejected on an explicit aiplatform base learns about the default
- doctor's explicitly configured aiplatform base now also gets the
publishers completion; the OAuth Vertex .../endpoints/openapi base
stays untouched
Fixes#115306
Google issues two Gemini key families: AI Studio keys (AIza...) and Vertex AI
express-mode keys (AQ....). Express keys only authenticate against
aiplatform.googleapis.com; the native adapter hardcoded the Studio host, so an
express key had no working path (403), and an explicit aiplatform base URL was
not even recognised as native Gemini and used the wrong model path.
- normalize_gemini_base_url(base_url, api_key="") routes an AQ. key that would
land on generativelanguage to
https://aiplatform.googleapis.com/v1beta1/publishers/google; an explicit proxy
base is never rewritten. The express base carries the publishers/google
prefix so every {base}/models/{model}:... builder (chat, tier probe, Gemini
TTS) needs no path branching; an explicit aiplatform host root / v1beta1 base
is completed to that form.
- is_native_gemini_base_url accepts the express host but NOT the OAuth Vertex
provider's .../projects/{p}/locations/{r}/endpoints/openapi base, which is
OpenAI-compatible and must stay off the native adapter.
- GeminiNativeClient, probe_gemini_tier and both Gemini TTS call sites pass the
key through.
Cherry-picked from the reporter's earlier PR #96587 and reshaped onto current
main; #114343 and #101918 proposed the same routing.
Fixes#114335
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: cloim <cloimism@gmail.com>
Pass the presented key into gemini_http_error so a leftover AIza
Standard key that Google now rejects as "API key not valid" gets the
same migration text as the original 401. Also tell the user to delete a
stale Windows/shell GEMINI_API_KEY that can hide ~/.hermes/.env.
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.
normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).
Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
Slot arguments are always complete json.dumps output (Gemini re-sends full args), so the mid-stream JSON check could never fire; the id key needs no tool name; _new_call_id and the slot lookup now share _provider_call_id.
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.
Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.
Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).
Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.
Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.
The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.
- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
after the read loop; surface EOF-without-[DONE] (warn + error result
when nothing was received); extract _consume_sse_line so line parsing
and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
proven red on origin/main.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Collapse the four Gemini chat/completions wrapper classes into SimpleNamespace
shims, finish-reason/tool-choice/HTTP-error mappings become dicts, shared
_usage_from_metadata/_assistant_message between sync and streaming translation,
vertex credential loading split into _load_credentials/_needs_refresh.
Gemini request/response/schema output byte-identical against merge-base.
``create_openai_client`` was a hardcoded if-ladder: copilot-acp builds an ACP
stdio shim, gemini builds a native client, everything else gets an
``openai.OpenAI``. There was no extension point, so a provider whose wire
protocol is not OpenAI-over-HTTP could only be added by editing this function —
which is exactly why an ACP provider cannot ship outside this tree today, even
though ``providers/__init__.py`` has discovered out-of-tree profiles from
``~/.hermes/plugins/model-providers/`` and pip entry points for a while.
``ProviderProfile.create_client(**client_kwargs)`` closes that gap. It returns
``None`` by default, so every provider that wants the standard client is
unaffected and the existing ladder still runs as the fallback. copilot-acp is
migrated onto it — its hardcoded branch is gone and its profile supplies the
client in three lines, which is the same three lines an external package writes.
Resolution goes by provider name first, then by ``base_url`` prefix, so a
runtime configured only by URL still reaches its profile — matching what the
replaced ``startswith("acp://copilot")`` branch did. A profile that raises is
logged and skipped: a third-party plugin can fail to provide a client, but it
cannot take the turn down.
Also replaces the two ``isinstance`` checks in ``agent/auxiliary_client.py``
that mean "this client is complete, do not wrap it" with capability flags the
client class declares — ``HERMES_SKIP_TRANSPORT_WRAP`` and
``HERMES_SKIP_ASYNC_WRAP``, mirroring ``SUPPORTS_HERMES_TOOL_CALLS`` in
``background_review.py``. Two in-tree consumers (the ACP shim and the Gemini
native client), an out-of-tree client is covered by the same declaration, and
the hot path no longer imports those modules just to type-test.
Co-Authored-By: Junie <junie@jetbrains.com>
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.
Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.
Original PR #32352 by @hbentel, salvaged onto current main.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
(already wrapped, no 400 path), and $ref without '#/' prefix (stays
structured).
- Document the deliberate structural (false-positive-tolerant) detection and
its O(n) cost in the helper docstring.
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).
tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.
Adds regression tests for the wrap path and the unchanged structured path.
Gemini 3+ models require explicit tool call IDs on functionCall /
functionResponse parts in replayed history; without them parallel tool
calls can be rejected or mispaired. The native adapter now:
- threads the model id into request building and includes ids for
Gemini >= 3 (version-gated: 2.x rejects unexpected id fields)
- preserves provider-returned functionCall.id on both non-streaming
and streaming responses instead of always minting a random one
Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.
Refs #83915
Port from google-gemini/gemini-cli#28700: when an interrupted/failed turn
leaves history ending on an unanswered tool result and the user sends a new
message, fusing the two into one Gemini user content makes the model read the
trailing text as a continuation of the tool result — it 'finishes your
sentence' instead of answering.
Builds on #68863 (@rille111), which split the mixed functionResponse/text
merge but emitted two consecutive user contents — a shape Gemini's
alternation contract rejects with HTTP 400 on other request paths (#55125).
This follow-up interposes gemini-cli's INTERRUPTED_RESPONSE_PLACEHOLDER model
turn between the split contents so the request stays alternation-valid while
the user's message remains a turn of its own.