SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis
SAGE uses MeSH anchors so small local models can synthesize grounded medical QA training data.
SAGE is a data-synthesis framework that lets small locally deployed models create medical question-answering training data using public taxonomies such as MeSH as semantic anchors. It alternates concept-level and relation-based generation, bootstrapping from minimal seeds without large medical corpora or proprietary APIs. Across multiple medical QA benchmarks, models fine-tuned on SAGE data outperform self-derived and conventional document-based training. The authors release code on GitHub.
- MeSH-style taxonomies act as semantic anchors for generation.
- Atomic and relation-based synthesis alternate from minimal seeds.
- The method avoids large document sets and external cloud APIs.
- SAGE-tuned models beat self-derived and document-based training baselines.
Full article185 words · extracted from arxiv.org · click to collapse
Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08093