The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.
Reactive, destination-scoped recovery instead:
- error_classifier: new `FailoverReason.role_alternation` for the vendor
alternation wordings (checked before the request-validation table since
the body also carries `invalid_request_error`); same abort+fallback hints
as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
loop's `_merge_user_content`), `destination_key` (base_url|provider,
model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
adjacent user turns merged, remember the destination on the facade for
the session so later iterations pre-merge, never touch destinations that
accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.
Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.
Fixes#112358