ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Daniel P. Jeong

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

infoAI researchimportance 22
AI summary · glm-5.3-flash

Study shows radiology reporting-style variations in reference reports can flip rankings of chest X-ray report generation models; releases MIMIC-CXR-Ext-ReRef dataset.

The paper quantifies how variations in radiologists' reporting practices distort evaluation of radiology report generation (RRG) models, introducing a radiologist-informed taxonomy and the ReRef method for rewriting reference reports while preserving clinical meaning. On MIMIC-CXR with RadCliQ-v1, condensing normal-findings discussion caused Libra to drop from first to second while CheXOne rose from third to first among nine models. The authors release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 original/alternative reference pairs, arguing metrics conflate clinical correctness with stylistic conformity.

  • Reference report style can materially change RRG model rankings
  • ReRef rewrites references along a radiologist-informed taxonomy of practice variation
  • Example ranking flip: Libra 1st to 2nd, CheXOne 3rd to 1st on MIMIC-CXR/RadCliQ-v1
  • Releases 120 radiologist-validated reference pairs (MIMIC-CXR-Ext-ReRef)
Full article214 words · extracted from arxiv.org · click to collapse

Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts large enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using RadCliQ-v1, condensing the discussion of normal findings in the reference reports causes Libra to drop from first to second place while CheXOne rises from third to first. Our results suggest that many current metrics fail to decouple clinical interpretation from conformity to reporting practices and that choosing the ``right'' references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.19093