SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
SkillGym automatically builds verifiable skill tasks and trains agents that outperform much larger models.
SkillGym crawls internet skills, keeps workflows that run reproducibly offline, and uses a builder-reviewer pipeline to create difficulty-controlled tasks with reference solutions and executable verifiers. The authors built 6,800 environments and collected 19,000 verified successful trajectories for supervised finetuning. Finetuning improved models from 2B to 122B parameters across four skill-use benchmarks, and a Qwen3.5-9B SFT model beat a 397B untrained model on two of them. Training raised the rate of reading the relevant skill from 28% to 96%, with gains holding for minority task types and held-out skills.
- Pipeline keeps offline-reproducible skills and builds verified tasks.
- Authors created 6.8k environments and 19k successful trajectories.
- SFT helped models from 2B to 122B across four benchmarks.
- Qwen3.5-9B SFT beat a 397B untrained model on two benchmarks.
- Relevant-skill reading rose from 28% to 96%.
Full article202 words · extracted from huggingface.co · click to collapse
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.37539