The llama-install.sh installer keeps its own version index on Hugging
Face (ggml-org/install.sh): a `latest` pointer (resolve/latest →
"b10679") plus a per-build tree of prebuilt llama-app binaries. That
pointer IS the installer's updater — the "what should we be on now"
signal, unauthenticated and unrate-limited.
Wire the llama.cpp package family to it:
- llama_app_latest() reads the `latest` pointer (bare build number).
- llama_app_bucket_versions() enumerates the tree API (deduped, sorted
desc). Verified: the HF bucket tree API IGNORES the offset param —
every offset returns the same first page (the oldest ~1000 paths) — so
the tree can only see the OLDEST builds; `latest` is the authoritative
source and callers put it first.
- LlamaCpp.latest_versions(): [latest, *tree] when the bucket answers,
falling back to the GitHub releases tags when it doesn't. Artifacts
still fetch from the llama.cpp GitHub releases (1:1 tag correspondence
— every bucket tag is a GitHub release tag, so a bump always has our
per-target assets); the bucket's CONFIG-coded llama-app binaries are
hardware-probe-derived and not precomputable, so they stay out.
Live: `pm update --check llamacpp-cpu` → 10362 → 10679 via the bucket.
7 new pure tests (monkeypatched fetchers) covering latest parsing,
HF_TOKEN auth, tree dedupe/sort, and the GitHub fallback.
Verified: 143 pm tests pass.