IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
IdeaAMBIG, a 660-instance benchmark of underspecified research-idea details, shows that across 13 LLMs the best model recovers only 9.6% of implementation-critical defects on real-world instances, identifying defect localization as the main bottleneck.
IdeaAMBIG is a benchmark of 660 evidence-grounded instances built from papers, codebases, and reproduction artifacts to test whether research-method specifications contain enough information for faithful implementation. It combines 163 real-world gaps drawn from reproducibility reports and GitHub issues with 497 controlled synthetic gaps, and evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Across 13 evaluated LLMs, the best model achieved only a 9.6% Macro Defect Recovery Rate on real-world instances, though it reached 80.6% clarification success when given the annotated defect. An oracle study showed gold resolutions raise the codification-ready rate from 14% to 98%, and defect localization emerged as the main bottleneck across all evaluated models. Both source reports (dated 2026-09-08 and 2026-09-09) agree on all figures.
- Benchmark contains 660 evidence-grounded instances: 163 real-world gaps (from reproducibility reports and GitHub issues) plus 497 controlled synthetic gaps.
- Instances were built from papers, codebases, and reproduction artifacts.
- The benchmark evaluates codification-readiness assessment, defect localization, and clarification action generation.
- 13 LLMs were evaluated; the best model achieved a 9.6% Macro Defect Recovery Rate on real-world instances.
- With the annotated defect provided, the best clarification success rate was 80.6%.
- An oracle study showed gold resolutions raise the codification-ready rate from 14% to 98%.
- Defect localization was identified as the main bottleneck across all evaluated models.
- Both source reports agree on all reported figures; no discrepancies were found.
Coverage timelineoldest first · each row is one article
- · 7d agoIdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Hugging Face daily papers· 30
IdeaAMBIG benchmark of 660 specification-gap instances shows LLMs localize implementation-critical research gaps poorly, with best model at 9.6% defect recovery.