SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
SkillSpec applies Hoare-style specification reasoning to detect semantic defects in autonomous agent skill repositories, confirming 763 defects.
The paper proposes SkillSpec, a Hoare-style framework that converts heterogeneous agent skill repositories into a unified graph and derives ExpectSpecs and FactSpecs under an intent mask, then auto-validates candidate defects in an isolated sandbox. Evaluated on 515 real-world skills from SkillsBench and widely downloaded repositories, it identified 763 manually confirmed defects across 239 skills at 61.2% precision. Specification reasoning proved reliable for code nodes while plain-text nodes remain the main bottleneck, with most defects at intent-implementation boundaries.
- Hoare-style framework unifies agent skills into a specification graph
- 763 manually confirmed defects found across 239 of 515 skills
- 61.2% precision; code nodes reliable, plain-text nodes bottleneck
- Candidate defects validated automatically in isolated sandbox
Full article241 words · extracted from huggingface.co · click to collapse
Autonomous agent systems increasingly depend on reusable skill abstractions for consolidating experiential knowledge and domain expertise. These artifacts typically bundle free-form instructions with heterogeneous resources. However, ensuring their correctness remains challenging. Their failure modes transcend conventional code defects to subtle semantic inconsistencies such as intent conflicts, which manifest as silent failures masked by the underlying model. Moreover, skill correctness must be grounded in intended task boundaries and generalizability. We propose SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem. It transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts. For each node, SkillSpec derives an ExpectSpec from the surrounding declared intent, and infers FactSpecs from encoded behavior under partially disclosed intent. An intent mask regulates access to holistic, lineage, neighborhood, and local views to balance the bias introduced by excessive context against unsupported inference caused by insufficient context. SkillSpec jointly reasons over these views to flag candidate defects, and automatically validates them in an isolated sandbox. On 515 real-world skills from SkillsBench and widely downloaded repositories, SkillSpec identified 763 manually confirmed defects across 239 skills, achieving 61.2% precision. The node-level analysis across multiple model families shows that specification reasoning is consistently reliable for code nodes, whereas plain-text nodes remain a major bottleneck. Most defects arise at the boundaries between declared intent and implementation, demonstrating that explicit specifications provide a practical foundation for skill quality assurance in real-world agent ecosystems.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.06052