Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support
Only 23.3% of concepts cited by clinical LLMs were causally necessary on MedQA.
Researchers introduce three faithfulness metrics—Explanation Stability Index, Causal Faithfulness Score, and Perturbation Stability Score—and evaluate six LLMs on 150 MedQA-USMLE questions (900 observations). Only 23.3% of cited clinical concepts were causally necessary. Correct answers had lower CFS than incorrect answers (0.212 versus 0.398), and consistency negatively predicted CFS (Spearman r = -0.466). Model pairs could agree on answers while sharing only 8.8% of cited reasoning concepts.
- Only 23.3% of cited clinical concepts were causally necessary.
- Correct answers had lower CFS than incorrect ones (0.212 vs 0.398).
- Answer consistency negatively predicted CFS (Spearman r = -0.466).
- Agreeing model pairs shared only 8.8% of cited concepts.
- Tests use concept ablation, not internal circuit analysis.
Full article169 words · extracted from arxiv.org · click to collapse
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation Stability Index (ESI), which measures reasoning consistency across repeated queries; the Causal Faithfulness Score (CFS), which tests whether cited concepts drive predictions via concept ablation; and the Perturbation Stability Score (PSS), which measures robustness to semantic-preserving paraphrases. By evaluating six LLMs on 150 MedQA-USMLE questions (900 model-question observations), we found that only 23.3% of the cited clinical concepts were causally necessary. Correct answers had lower CFS than incorrect answers (0.212 vs. 0.398), answer consistency negatively predicted CFS (Spearman r = -0.466), and model pairs could agree on answers while sharing only 8.8% of cited reasoning concepts. These results show that accuracy, consistency, and consensus are incomplete safety signals for clinical decision-making support. The evidence is behavioral rather than mechanistic: concept ablation tests counterfactual sensitivity of outputs, not internal circuits.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.32817