Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
A study finds zero-shot LLM prompting extracts intercropping meta-analysis data better than staged or multi-agent workflows, but remains inaccurate.
Researchers compared direct zero-shot prompting, a staged workflow, and a multi-agent system across six open-weight models for extracting data from intercropping literature. Outputs were scored against a manually curated ground truth and checked with a downstream statistical analysis. Direct zero-shot prompting was strongest, with a mean similarity-adjusted F1 of 0.577, but no method approached full accuracy. Most model-approach pairs recovered the direction of predictor-outcome relationships without estimating magnitude accurately.
- Compared zero-shot, staged, and multi-agent extraction with six open-weight models.
- Zero-shot prompting led with mean similarity-adjusted F1 of 0.577.
- No approach was close to fully accurate extraction.
- Downstream analysis usually recovered relationship direction, not magnitude.
Full article204 words · extracted from arxiv.org · click to collapse
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35089