The static recommended flag in catalog.json picked the dense 27B on
every machine, including unified-memory boxes where it decodes at
~13 tok/s while the 35B-A3B MoE does ~60. Replace the flag with a
per-machine derivation: decode is memory-bound, so predicted speed is
bandwidth over bytes-read-per-token, and the pick is the highest-quality
entry that runs resident and clears a 20 tok/s pleasant floor — else the
fastest resident entry, else the least-painful spill.
Catalog entries carry two authored fields in place of the flag:
quality (AA-informed ordering, editorially owned — never fetched at
runtime) and decode_fraction (share of weight bytes a token actually
reads; 1.0 dense, the active-slice ratio for MoE). The bandwidth axis
is the existing uma flag for now; measured per-machine bandwidth can
replace the class constants without touching the rule.
All three consumers derive: the pane badge and hero card through the
catalog route, quickstart's default target through the same resolver,
each gated on engine eligibility. The decision table lives on as a
checked-in test pinning every memory-class x bandwidth cell — a catalog
change flips cells in that file and the diff in review IS the editorial
sign-off. scripts/aa_quality_sync.py proposes quality updates at
authoring time; the commit decides.
The quickstart fixture's select_variant stub now constructs a real
VariantChoice — the resolver reads zero_spill, which the SimpleNamespace
stub lacked.