QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge
QuranicMMLU offers 980 questions testing generative models on Quranic Arabic across five linguistic pillars.
QuranicMMLU evaluates generative models on Quranic Arabic using a five-pillar taxonomy of phonology, morphology, syntax, semantics, and pragmatics, with 31 detailed phenomena. The set contains 980 human-reviewed questions, each in open-ended and multiple-choice form, stratified by Bloom cognitive level and verse perplexity. Across 12 systems, an Islamic-specialized model leads, but average multiple-choice accuracy is 84% versus 60% open-ended quality. Rankings agree closely (Kendall tau 0.73), while multiple choice conceals failures visible only without answer options.
- 980 human-reviewed questions appear in open-ended and multiple-choice forms.
- Five pillars cover phonology, morphology, syntax, semantics, and pragmatics across 31 leaves.
- Twelve systems average 84% multiple-choice accuracy versus 60% open-ended quality.
- An Islamic-specialized model leads; rankings agree with Kendall tau 0.73.
Full article187 words · extracted from arxiv.org · click to collapse
We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall's τ=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.22038