Files
hermes-agent/agent
Mehmet Ali 983532311f fix(agent/gemini): count thinking tokens in usage and cost
`_usage_from_metadata` (agent/gemini_native_adapter.py:476) read
`candidatesTokenCount` for `completion_tokens` and never read
`thoughtsTokenCount`; a repo-wide grep found that field nowhere. Gemini
reports hidden thinking in its own counter — `candidatesTokenCount` covers
visible output only, while `totalTokenCount` already includes thoughts. So on
any thinking turn the emitted usage contradicted itself: prompt + completion
did not add up to total.

Both the non-streaming assembly (translate_gemini_response) and the streaming
one (the finish chunk in translate_stream_event) share this helper, so both
under-billed identically. `normalize_usage` takes `completion_tokens` for
`output_tokens` and looks for reasoning under
`completion_tokens_details.reasoning_tokens`, which the adapter never set, so
the reasoning column read 0 for models whose spend is mostly reasoning.

Thoughts are now folded into `completion_tokens` (OpenAI's counter includes
reasoning) and surfaced under `completion_tokens_details.reasoning_tokens`,
matching the nesting the adapter already uses for
`prompt_tokens_details.cached_tokens`. The field is absent on non-thinking and
older responses; those count 0 and their numbers do not move.

Live: for usageMetadata {prompt 10, candidates 200, thoughts 5000, total 5210}
the adapter emitted completion_tokens=200 (10 + 200 != 5210) and
normalize_usage returned output_tokens=200, reasoning_tokens=0. It now emits
completion_tokens=5200 (10 + 5200 == 5210) with
completion_tokens_details.reasoning_tokens=5000, and normalize_usage returns
output_tokens=5200, reasoning_tokens=5000.
2026-09-23 22:31:59 -07:00
..
…
…
2026-09-11 09:31:13 -07:00