SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
SkillSeek shows standard retrieval matches LLM-mediated agent skill selection at about half the cost.
SkillSeek is an open-source two-stage retriever for Anthropic-style agent skills, pairing a BGE-base bi-encoder with a small cross-encoder exposed over MCP. Open aggregations now exceed 230,000 skills. Across an 89-task SkillsBench grid under OpenHands, it reached parity with an LLM-mediated retrieval loop, and BM25 alone matched or beat that loop on three of four settings. Per-trial cost fell from USD 51.30 to USD 27.54.
- Open skill aggregations now exceed 230,000 SKILL.md packages.
- BM25 alone matched or beat the LLM loop on three of four settings.
- A small cross-encoder closed the remaining gap on the fourth setting.
- Per-trial spend dropped from USD 51.30 to USD 27.54.
Full article208 words · extracted from huggingface.co · click to collapse
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a 4 times 11 grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.38822