skills.sh repos that keep SKILL.md at the repo root (no skill directory, e.g.
orzcls/win-disk-cleaner) are listed by `hermes skills search` but `inspect` and
`install` failed with "Could not find '<identifier>' in any source": discovery only
matched `<name>/SKILL.md` paths and the repo-root scan skipped non-directory entries by
construction, so no identifier form could express "the skill directory IS the repo root".
Install goes through the same resolver, so those hub-listed skills were uninstallable.
- GitHubSource._find_repo_root_skill() returns the `owner/repo/` identifier (empty path
segment = repo root) when the tree holds exactly one SKILL.md and it sits at the root,
so a categorized multi-skill repo never resolves this way.
- SKILL.md, support-file, bundle-name and source_url paths go through _skill_file_path(),
which collapses an empty skill path to the repo root; the bundle is named after the repo
(a root skill has no directory to be named after).
- SkillsShSource._discover_identifier() consults the root fallback after the tree walk.
Gate review: a transient blob failure installs with a gap, and with the tree sha recorded the
update check would report that gap up_to_date until upstream moved (main re-fetched and healed
it). _collect_tree_files now reports completeness and fetch() records the sha only for a
complete bundle. _get_repo_tree caches its misses so the new probe cannot double the GETs on a
truncated tree, and a probe exception degrades one row to the full-fetch path instead of
aborting the whole check.
`check_for_skill_updates` fetched the full bundle (SKILL.md + every support
blob) for every installed hub skill just to re-hash it, even though each lock
entry already records the tree sha it was installed from. GitHubSource now
exposes `current_revision()` (one cached tree lookup per repo, no blob GETs)
and the update loop reports `up_to_date` without fetching when the recorded
`source_revision` still matches. Entries without a recorded revision, or
whose registry has no revision probe, keep the full-fetch path.
Closes#101454
Co-authored-by: gobeumsu <gobeumsu@gmail.com>
Follow-up polish on the provider-filter-before-limit fix:
- GitHubSource.search now skips taps whose repo maps to a different provider
instead of enumerating every tap and filtering afterwards. A tap's repo fixes
the provider of every result it yields (github_provider_for is the only source
of extra.provider in this adapter), so the skip is lossless and avoids up to 23
useless tap enumerations per provider-filtered search — real GitHub API calls
against the 60/hr unauthenticated budget and the 30s overall timeout when the
index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
it replaces the four inline copies of the strip/lower/membership idiom in
_select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
provider cut as do_browse; it accepted a provider value but returned
unfiltered results.
Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.
The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
~480 scientific research skills become searchable/installable through the Skills Hub
with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and
synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0).
A new optional tap-level `bucket` key stamps extra["category"] on every skill from a
tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub
category; a sidecar grouping still wins when present. Both repos stay at community
trust (not in TRUSTED_REPOS) so the guard scans every install.
Re-grafted from #60559 onto the post-split tools/skills_hub_github.py.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.