Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
Code2Skill mines 19,769 GitHub repositories into CodeSkillBank, 1,006,822 verified skill records that lift agent benchmark performance by 11.7% on average.
Code2Skill transforms code units into implementation-anchored records of atomic operations, workflows, and patterns, verifying each via source-body-blind reconstruction and source-aware comparison. Applied to 19,769 popular repositories, it produces CodeSkillBank with 1,006,822 accepted records. Across 72 protocol-matched evaluations spanning nine model settings and eight benchmarks, retrieval-augmented models improve 11.7% on average and win 57 cases, beating trajectory-derived skill banks on all seven shared benchmarks; skills from AI-generated code pass at 93.50%.
- 1,006,822 verified skill records mined from 19,769 repositories
- Skills verified through source-body-blind reconstruction
- Agent benchmarks improve 11.7% on average across 72 evaluations
- AI-generated code skills pass verification at 93.50%
Full article224 words · extracted from huggingface.co · click to collapse
Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with specific environments, while document-derived skills may lack executable evidence and verification. Source code offers a complementary path: it requires no prior agent experience yet provides executable evidence for grounding abstractions. We present Code2Skill, a fully automated pipeline that transforms selected code units into implementation-anchored records of atomic operations, composite workflows, and recurring patterns, then verifies each record through source-body-blind reconstruction and source-aware comparison. Applied to 19,769 popular, actively maintained GitHub repositories, Code2Skill produces CodeSkillBank, a grounded bank of 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, models augmented with retrieved CodeSkillBank skills improve by 11.7% on average over matched baselines and outperform them in 57 cases. Under a unified downstream interface, Code2Skill also outperforms trajectory-derived skill banks on all seven shared benchmarks, showing that repository-derived skills can provide useful procedural knowledge before agents accumulate sufficient interaction experience. Skills synthesized from tested AI-generated code achieve a 93.50% pass rate, compared with 93.00% for human-written code, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software. Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.05571