2 Commits

Author SHA1 Message Date
Teknium
6dd091a89c test: fold Gemini thinking-budget tests into invariant form 2026-09-12 21:10:02 -07:00
Shakti Prasad Mohapatra
29056b2335 fix(title): disable Gemini thinking tokens to prevent max_tokens starvation
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.

Two-part fix:

1. title_generator.py: Pass reasoning_config={"enabled": False} to
   call_llm() so thinking is explicitly disabled for title generation.

2. chat_completions.py: In _build_gemini_thinking_config, when
   reasoning is disabled (enabled=False or effort="none"), set
   thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
   includeThoughts: False only hides thought parts from the response
   while the model still reasons internally and bills thought tokens
   against maxOutputTokens. thinkingBudget: 0 truly disables thinking
   so thought tokens do not consume the max_tokens budget.

Fixes #91927
2026-09-12 21:10:02 -07:00