The compression/aux fallback ladder decides "payment error" via
_is_payment_error(), which only knew the spaced "resource exhausted". NVIDIA NIM
and gRPC-style wrappers serialize the same quota signal as ResourceExhausted /
RESOURCE_EXHAUSTED / resource-exhausted, so a 403/429/status-less body carrying
it re-raised instead of walking the configured fallback_chain, and compression
fell to the lossy emergency path. Add the three separator variants to
_PAYMENT_KEYWORDS (same status gate, no broader "exhausted" matching).
evals/auxiliary_resource_exhausted.py drives the real call_llm/async_call_llm
against two local OpenAI-compatible listeners (nvidia profile primary, named
custom fallback) for the before/after check. #85649