fix(agent): use provider-default temperature for title generation (#72351)
Title generation hardcoded temperature=0.3 in its call_llm() call. Models like GPT-5.6 only accept their server-side default temperature and reject explicit values with "Unsupported value: 'temperature'". While call_llm() has a retry that strips temperature on error, the daemon thread races with session cleanup in short-lived CLI sessions, causing the retry to fail with a connection error. Fix: pass temperature=None so the provider uses its own default. This avoids the unsupported-temperature error entirely and eliminates the need for the retry path. Fixes #72351
This commit is contained in:
@@ -320,11 +320,16 @@ def generate_title(
|
||||
"__LANGUAGE_RULE__", _LANGUAGE_RULE_PINNED.format(language=language) if language else _LANGUAGE_RULE_MATCH_USER,
|
||||
)
|
||||
try:
|
||||
# Use the provider's default temperature instead of forcing 0.3.
|
||||
# Some models (e.g. GPT-5.6) only accept their server-side default
|
||||
# and reject explicit temperature values, causing the daemon title
|
||||
# thread to fail with "Unsupported value: 'temperature'".
|
||||
# See: #72351, #51083, #51157
|
||||
response = call_llm(
|
||||
task="title_generation",
|
||||
messages=[{"role": "system", "content": prompt}, {"role": "user", "content": user_snippet}],
|
||||
# A title is a handful of tokens; a larger ceiling let chatty models burn seconds.
|
||||
max_tokens=64, temperature=0.3, timeout=timeout, main_runtime=main_runtime,
|
||||
max_tokens=64, temperature=None, timeout=timeout, main_runtime=main_runtime,
|
||||
extra_body={"response_format": _TITLE_RESPONSE_FORMAT},
|
||||
# The module contract above promises thinking-disabled operation,
|
||||
# but nothing enforced it: with the aux default reasoning_effort
|
||||
|
||||
Reference in New Issue
Block a user