fix(agent): fallback to Anthropic-style token fields in normalize_usage
Local OpenAI-compatible servers like mlx_vlm.server emit input_tokens/output_tokens in chat completion responses instead of prompt_tokens/completion_tokens. The OpenAI Python client preserves these as extra attributes, but normalize_usage() only looked at OpenAI-style field names, causing input_tokens to always be 0. This made the context progress bar stay at 0% forever and prevented auto-compression from triggering. Fix: add Anthropic-style fallback (input_tokens/output_tokens) in the default else-branch of normalize_usage, with OpenAI-style names taking priority. Fixes #14686
This commit is contained in:
@@ -1283,8 +1283,16 @@ def normalize_usage(
|
||||
)
|
||||
input_tokens = max(0, input_total - cache_read_tokens - cache_write_tokens)
|
||||
else:
|
||||
prompt_total = _to_int(getattr(response_usage, "prompt_tokens", 0))
|
||||
output_tokens = _to_int(getattr(response_usage, "completion_tokens", 0))
|
||||
# OpenAI-style names first; fall back to Anthropic-style
|
||||
# (input_tokens/output_tokens). Local OpenAI-compatible servers like
|
||||
# mlx_vlm.server emit the Anthropic names in chat_completions responses,
|
||||
# and the OpenAI Python client preserves them as extra attributes.
|
||||
prompt_total = _to_int(getattr(response_usage, "prompt_tokens", 0)) or _to_int(
|
||||
getattr(response_usage, "input_tokens", 0)
|
||||
)
|
||||
output_tokens = _to_int(getattr(response_usage, "completion_tokens", 0)) or _to_int(
|
||||
getattr(response_usage, "output_tokens", 0)
|
||||
)
|
||||
details = getattr(response_usage, "prompt_tokens_details", None)
|
||||
# Primary: OpenAI-style prompt_tokens_details. Fallback: Anthropic-style
|
||||
# top-level fields that some OpenAI-compatible proxies (OpenRouter, Vercel
|
||||
|
||||
Reference in New Issue
Block a user