fix(agent): explicit "retry after N s" wins over "resets in ..." in the shared reset table

Unifying the credential pool's `_RETRY_DELAY_PATTERNS` into `RETRY_DELAY_PATTERNS`
flipped the pool's precedence: "retry after 30s; resets in 4hr" cooled the
credential for 14400 s where the pool used to take 30. A body carrying both
describes a short throttle inside a long quota window; the explicit retry-after
is the wait the provider actually asks for, so it is tried before "resets in".

Review follow-up on #109539.
This commit is contained in:
teknium1
2026-09-12 23:43:48 -07:00
committed by Teknium
parent 8121b8f438
commit 3bef6b6a5c
2 changed files with 7 additions and 1 deletions

View File

@@ -88,10 +88,13 @@ def _resets_in_seconds(m: "re.Match[str]") -> Optional[float]:
return float(m.group(1) or 0) * 3600 + float(m.group(2) or 0) * 60 + float(m.group(3) or 0)
# An explicit "retry after N s" wins over "resets in ..." (the credential pool's precedence):
# a body carrying both describes a short throttle inside a long quota window, and the
# shorter explicit wait is the one the provider actually asks for.
RETRY_DELAY_PATTERNS = (
(_QUOTA_RESET_DELAY_RE, _quota_reset_seconds),
(_RESETS_IN_RE, _resets_in_seconds),
(_RETRY_AFTER_SECONDS_RE, lambda m: float(m.group(1))),
(_RESETS_IN_RE, _resets_in_seconds),
)

View File

@@ -57,6 +57,9 @@ class TestResetDelayOneTable:
("Limit hit; resets in 45s", 45.0),
('"quotaResetDelay": "1500ms"', 1.5),
("please retry after 12 seconds", 12.0),
# Both grammars in one body: the explicit retry-after wins (pool precedence), not the
# multi-hour quota window.
("Rate limited. Retry after 30s; resets in 4hr", 30.0),
])
def test_credential_pool_and_error_context_agree(self, message, seconds):
"""The pooled-credential cooldown and the UI's error context read the same table, so the