A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Researchers benchmark AI reviewers' rhetorical robustness and propose SciCore, averaging full-text and science-core judgments.
AI reviewers can score the same science differently after content-preserving rewrites, rewarding rhetoric over substance. The authors define rhetorical robustness as stability under rewrites plus discrimination across papers, and introduce RobustReview with 1,260 manuscript versions and 30 reviewer configurations. Content-focused prompting does not consistently improve robustness. SciCore averages a full-manuscript judgment with one based on an extracted science core; in the primary GPT-5.5 comparison it leads the joint stability-discrimination profile while remaining competitive on human alignment.
- RobustReview covers 1,260 manuscript versions and 30 reviewer configurations.
- Low rewrite sensitivity can hide score collapse across different papers.
- Human alignment and rhetorical robustness rank reviewers differently.
- SciCore with GPT-5.5 leads joint stability and discrimination.
Full article175 words · extracted from huggingface.co · click to collapse
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39027