Files
hermes-agent/tests/agent/test_meta_usage_cache_reporting.py
Lavie 24545418eb fix(providers): route api.meta.ai through Responses API for prompt caching
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
  api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
  93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
  api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
  (after explicit api_mode wins, before provider-name specials) via lazy
  import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
  24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
  override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
  auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
  Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
  content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
  usage cache reporting, model-switch override, and config roundtrip; extend
  test_model_switch_openai_api_mode with meta cases.
2026-08-17 12:58:51 -07:00

47 lines
1.6 KiB
Python

"""Usage cache reporting for Meta Responses path."""
from types import SimpleNamespace
from agent.usage_pricing import normalize_usage
def test_meta_cached_tokens_flows_to_cache_read():
usage = SimpleNamespace(
input_tokens=4000,
output_tokens=100,
input_tokens_details=SimpleNamespace(cached_tokens=3920, cache_creation_tokens=0),
output_tokens_details=None,
prompt_tokens=None,
completion_tokens=None,
)
cu = normalize_usage(usage, api_mode="codex_responses")
assert cu.cache_read_tokens == 3920
assert cu.input_tokens == 80 # 4000 - 3920
# rendered cache line would be cache=3920/4000 (98%)
total = cu.prompt_tokens
assert total == 4000
pct = (cu.cache_read_tokens / total * 100) if total else 0
assert 97 <= pct <= 99
rendered = f"cache={cu.cache_read_tokens}/{total} ({pct:.0f}%)"
assert rendered == "cache=3920/4000 (98%)"
def test_meta_cache_write_tokens():
first = SimpleNamespace(
input_tokens=4000,
output_tokens=100,
input_tokens_details=SimpleNamespace(cached_tokens=0, cache_creation_tokens=4000),
)
cu1 = normalize_usage(first, api_mode="codex_responses")
assert cu1.cache_write_tokens == 4000
assert cu1.cache_read_tokens == 0
second = SimpleNamespace(
input_tokens=4000,
output_tokens=100,
input_tokens_details=SimpleNamespace(cached_tokens=3920, cache_creation_tokens=0),
)
cu2 = normalize_usage(second, api_mode="codex_responses")
assert cu2.cache_read_tokens == 3920
assert cu2.cache_write_tokens == 0