Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
Ranking-aware prompt evolution lifts clinical AUROC on Qwen3-VL-8B and MedGemma-4B versus accuracy-based search.
The paper argues accuracy is a poor objective for class-imbalanced clinical diagnosis and instead optimizes multimodal prompts for AUROC. Ranking-PE replaces per-instance correctness in reflective prompt evolution with pairwise positive-versus-negative ordering, so the column average equals empirical AUROC without extra model calls. Across three MIMIC diseases, it beats accuracy-based evolution by 5.8 AUROC points on fine-tuned Qwen3-VL-8B and 16.2 points on MedGemma-4B. Ablations find that a medical-grade visual backbone, from tuned supervised fine-tuning or medical pretraining, is required and cannot be replaced by prompt search.
- Optimizes clinical prompts for AUROC rather than accuracy.
- Pairwise ranking makes the score average equal empirical AUROC.
- Gains of 5.8 and 16.2 AUROC points on two models.
- A medical-grade visual backbone remains a prerequisite.
Full article263 words · extracted from arxiv.org · click to collapse
Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40361