Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes
Biomedical reranking raised clinical-note Hit@10 from 46.6% to 60.6%, while Qwen3-8B correctness reached only 48.6%.
Researchers tested biomedical reranking inside a local retrieval-augmented generation pipeline over longitudinal clinical notes. Dense PubMedBERT retrieval, BM25, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking were evaluated on 1,000 questions from 200 bariatric surgery patients. Reranking raised Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B, judge-assessed answer correctness increased only from 44.8% to 48.6%.
- Pipeline combines PubMedBERT, BM25, rank fusion, and MedCPT reranking.
- Tested on 1,000 questions from 200 bariatric surgery patients.
- Hit@10 rose from 46.6% to 60.6%; MRR from 0.2371 to 0.3252.
- Qwen3-8B judged correctness improved only from 44.8% to 48.6%.
Full article156 words · extracted from arxiv.org · click to collapse
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.01324