fix(agent): fallback to Anthropic-style token fields in normalize_usage

Local OpenAI-compatible servers like mlx_vlm.server emit
input_tokens/output_tokens in chat completion responses instead of
prompt_tokens/completion_tokens. The OpenAI Python client preserves
these as extra attributes, but normalize_usage() only looked at
OpenAI-style field names, causing input_tokens to always be 0.

This made the context progress bar stay at 0% forever and prevented
auto-compression from triggering.

Fix: add Anthropic-style fallback (input_tokens/output_tokens) in the
default else-branch of normalize_usage, with OpenAI-style names taking
priority.

Fixes #14686
This commit is contained in:
Chengxi Zhou
2026-06-26 00:31:17 +08:00
committed by Teknium
parent eca85e81d6
commit 69bc3159e8

View File

@@ -1283,8 +1283,16 @@ def normalize_usage(
)
input_tokens = max(0, input_total - cache_read_tokens - cache_write_tokens)
else:
prompt_total = _to_int(getattr(response_usage, "prompt_tokens", 0))
output_tokens = _to_int(getattr(response_usage, "completion_tokens", 0))
# OpenAI-style names first; fall back to Anthropic-style
# (input_tokens/output_tokens). Local OpenAI-compatible servers like
# mlx_vlm.server emit the Anthropic names in chat_completions responses,
# and the OpenAI Python client preserves them as extra attributes.
prompt_total = _to_int(getattr(response_usage, "prompt_tokens", 0)) or _to_int(
getattr(response_usage, "input_tokens", 0)
)
output_tokens = _to_int(getattr(response_usage, "completion_tokens", 0)) or _to_int(
getattr(response_usage, "output_tokens", 0)
)
details = getattr(response_usage, "prompt_tokens_details", None)
# Primary: OpenAI-style prompt_tokens_details. Fallback: Anthropic-style
# top-level fields that some OpenAI-compatible proxies (OpenRouter, Vercel