Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging
Researchers prove off-policy evaluation under history-dependent logging requires exponentially many episodes, resolving a hardness question for model-based POMDP evaluation.
The paper constructs POMDPs with at most two latent states per stage, three actions, and a three-memory-state logger where evaluating a known deterministic target policy to accuracy 1/8 requires Θ((3/2)^H log(1/δ)) episodes for any horizon H≥3. Coverage and outcome-revealing conditions hold with constants independent of H, yet a reset erases the unknown transition that determines the target value. The authors characterize the resulting statistical experiment exactly, derive a matching optimal estimator, and validate predictions on a two-lane gridworld. This settles the history-dependent-logging, model-based case posed by Zhang and Jiang (arXiv:2503.01134).
When LLM judges agree, should we believe them?
Amazon ICML paper uses Ising models to correct correlated LLM-judge votes, beating accuracy-weighted panels by 9-14%.
Amazon Science describes an ICML paper, "Dependence-aware label aggregation for LLM-as-a-judge via Ising models," addressing how correlated judge outputs inflate majority-vote confidence. The unsupervised method models pairwise dependence between judges, learning both reliability and similarity without human reference labels. Tested on relevance, toxicity, and summarization tasks with 10 judge models at temperature zero, it outperformed accuracy-weighted voting by 9% to 14%.
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark Radar provides a living searchable database of 1,283 AI benchmark records and 12,916 score observations drawn from 37 daily discovery sources.
Benchmark Radar combines daily discovery of benchmark papers, repositories, datasets, and releases from 13 direct connectors and 24 first-party feeds into a searchable catalog with model card mentions and score histories. The catalog contains 1,283 source records drawn from 4 benchmark catalogs plus 12,916 numeric observations on 790 records. The release includes a web dashboard with leaderboard, Pareto frontier of score versus usage, saturation and trend views, daily feeds, a CLI, and reproducible analysis. The paper audits the full catalog and examines benchmark saturation and limits of score comparisons.
Piloting the world's first double-blind AI evaluations
Google DeepMind is piloting the world's first double-blind AI evaluations, a new methodology intended to improve evaluation integrity and reduce bias.
Google DeepMind announced a pilot of double-blind AI model evaluations, described as the first of its kind. The approach is designed to reduce contamination and bias in model assessments by keeping evaluators and model identities hidden from one another. Details on participating models and protocols were not provided in the announcement text.
Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
Evaluation of twelve LLMs on 222 clinical questions shows verbatim quotes rarely substantiate claims; claude-opus-5 fully substantiates only 37.1%.
The authors build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring citation attachment, verbatim quote production, and claim substantiation. Most models attach verbatim quotes to over 90% of claims from prompting alone, though lightweight models like claude-haiku-4.5 struggle. Quotes frequently fail to substantiate claims: claude-opus-5 quotes 98.0% of claims but fully substantiates only 37.1%, exposing a capability gap for verifiable clinical QA.
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
MP-Bench is the first benchmark for voice agents in multiparty conversations, finding real-time agents near chance on turn-taking.
MP-Bench is the first benchmark designed to objectively evaluate conversational speech systems as active participants in multi-party conversations. It assesses agents on turn-taking awareness and response appropriateness, with comprehension-based question-answering as a complementary evaluation. Benchmarking 12 voice agents shows real-time agents score at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking.
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
French BabyLM entry METRON-FR (125M GPT-2, 92.47M words) shows tokenizer artifacts dominate child-scale zero-shot evaluation; proposes standard diagnostics.
METRON-FR is a 125M-parameter GPT-2 pretrained on 92.47M French words, submitted to the BabyLM 2026 Strict track, scoring 85.97% on the native Quebec-French QFrBLiMP benchmark and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE protocol combining French task-data translation with rank-16 LoRA shows relational tasks gain while world-knowledge tasks regress. Bilingual Lexicon Induction reaches p@1 of 68.84%, 18x above chance, and ablations show single-token zero-shot scoring is dominated by tokenizer and template artifacts at child scale.
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
DataFlex-RL benchmark of 13 RLVR data policies on Qwen2.5-7B finds none reproducibly beats uniform sampling under matched GRPO training.
DataFlex-RL is an evaluation platform comparing rollout-selection, reweighting, and mixture data policies for RLVR under a common GRPO recipe. Across 13 configurations and 12 matched seeds with Qwen2.5-7B-Base on 12 math, logic, and science benchmarks, uniform GRPO improved domain-balanced accuracy by 7.76 points, but no alternative policy achieved a statistically significant improvement. A corrected 12-seed Llama-3.1-8B-Base extension found no consistent winner, and math-heavy evaluation summaries were negatively correlated (-0.33) with domain-balanced summaries.
Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Study shows radiology reporting-style variations in reference reports can flip rankings of chest X-ray report generation models; releases MIMIC-CXR-Ext-ReRef dataset.
The paper quantifies how variations in radiologists' reporting practices distort evaluation of radiology report generation (RRG) models, introducing a radiologist-informed taxonomy and the ReRef method for rewriting reference reports while preserving clinical meaning. On MIMIC-CXR with RadCliQ-v1, condensing normal-findings discussion caused Libra to drop from first to second while CheXOne rose from third to first among nine models. The authors release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 original/alternative reference pairs, arguing metrics conflate clinical correctness with stylistic conformity.
What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity
Pruning study across four LLM architectures finds dense models degrade sharply on smart-home tool calling while MoE models tolerate far more.
Researchers systematically study pruning-induced degradation in smart-home tool calling across four LLMs spanning dense Transformer, dense hybrid, and mixture-of-experts architectures, combining depth, width, hybrid, and expert pruning methods, and evaluate over 19,500 instances from three datasets after post-pruning supervised fine-tuning. Dense models show narrow safe pruning regions followed by sharp degradation, while MoE models tolerate substantially more pruning. Pruning degrades grounded specificity (operation, device, argument, value) before schema-level intent, and aggressive dense pruning can induce systematic over-refusal.
Can Edge-Deployable Vision-Language Models Identify Species?
Evaluation of 2-8B VLMs (Qwen3-VL, Gemma3) against BioCLIP on camera-trap species ID shows all models degrade sharply on field imagery.
The study tests whether edge-deployable 2-8B vision-language models carry genuine taxonomic knowledge, comparing Qwen3-VL 2B/4B/8B and Gemma3 4B against the 300M specialist BioCLIP on a 96-species task across clean iNaturalist photos and six LILA.science camera-trap collections. All models degrade 9.6-26.6 percentage points on field imagery, and BioCLIP outperforms every VLM by 33.2-59.2 points on an expanded 200-image sample, suggesting specialized data rather than scale drives the gap. Under open-set prompting, 5.9-9.6% of responses are syntactically valid but taxonomically nonexistent species names, with fabrication rankings replicating across evaluation sets.
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench is a 955-case benchmark testing language models as transit kiosk tool-calling runtimes across six real metro systems.
The benchmark covers 37-414-station metro systems and eleven task categories including routing, fare calculation, disruptions, accessibility, and adversarial input, with 14 deterministic and 8 semantic scoring components. Of 26 models from six vendors, a PEFT-tuned 4B Qwen 3.5 student scored 91.3 on Tier 1, exceeding GPT-5.6 (90.6/90.0), while Muse Glimmer 30B led the composite ranking. A deterministic rule-based baseline reached 84.6, and PEFT gains over base models shrank from +7.03 points at 2B to -0.91 at 27B.
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Researchers distill a compact 82M-parameter Thai TTS from synthetic OmniVoice data, enabling on-device fixed-voice synthesis without reference audio.
The paper uses a large voice-cloning model (OmniVoice) as a synthetic data source to train Wayu-Paxa-TTS-Edge, an 82M-parameter fixed-voice Thai TTS student. The model achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1), 91.4% pause precision, and CERs of 3.7% on Thai and 1.1% on English. It outperforms its teacher on pause placement and is open-sourced with its evaluation framework.
Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
Survey of 761 Nigerian healthcare professionals finds high AI awareness (92.6%) but limited knowledge, preparedness, and major training and infrastructure barriers.
A cross-sectional study of 761 healthcare professionals across Nigeria, conducted from December 2025 to March 2026, found 92.6% awareness of AI in healthcare but 40.9% reporting low knowledge and only 63.0% feeling adequately prepared. Top barriers were lack of training (84.7%), poor infrastructure (71.1%), and high tool costs (61.0%). Willingness to adopt was strong, with 92.5% interested in training and 78.7% supporting AI in undergraduate curricula; preparedness differed significantly across geopolitical zones and professions.
Thought without systematicity? Evaluating reasoning models on rule induction tasks
Study finds reasoning models often fail on structurally equivalent variants of tasks they solve, suggesting their reasoning lacks systematicity.
The paper extends rule induction tasks from cognitive science using task isomorphisms such as recombination and substitution to test systematicity in reasoning models. Despite solving tasks correctly, models frequently fail on structurally equivalent variants of the same task. The authors conclude many model behaviors lack systematicity, making it difficult to establish cognitive abilities beyond the specific evaluation contexts.
Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting
Study finds zero-shot time-series foundation models underperform on CGM forecasting; fine-tuned Chronos-Bolt cuts RMSE up to 18.4% and dietary context adds signal.
The paper evaluates time-series foundation models for continuous glucose monitoring forecasting across eight public datasets covering Type 1 diabetes, Type 2 diabetes, and non-diabetes populations. Under a unified protocol, zero-shot foundation models did not consistently outperform baselines like Elastic Net and PatchTST, but lightweight fine-tuning did, with fine-tuned Chronos-Bolt reducing RMSE by 6.5%-18.4% in the T1D cohort and 8.6%-18.2% in the non-diabetes/T2D cohort. A residual-based fusion framework adding dietary context from CGMacros reduced overall RMSE by about 3% and postprandial RMSE by about 15% versus CGM-only baselines.
Google's AI genome system evaluates every possible one-base change
Google's AlphaGenome AI system predicts functional effects of non-coding DNA variants in humans and mice.
Google's AlphaGenome AI system evaluates genomic sequences to predict gene expression, transcription factor binding, chromatin accessibility, splice site usage, and related genomic features. The system is currently limited to human and mouse sequences and a limited set of well-studied cell types, but its predictions generally match or exceed specialized software tools. Researchers can use it to assess whether non-coding variants are likely significant and generate hypotheses about their function.
Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
Audit of eight cybersecurity LLM benchmarks shows evaluation pipeline choices can swing scores by over 80 points and reshuffle most model rankings.
Researchers modeled eight cybersecurity benchmarks as configurable measurement pipelines and audited 10 proprietary, open-weight, and cybersecurity-specialized LLMs. They identified 15 systematic failure modes and showed a single pipeline choice can change a model's score by more than 80 percentage points and alter rankings; semantically similar task pairs rank the same models differently. Under a standardized harness, nine of 10 models shifted at least three ranks on at least one benchmark, motivating pipeline-aware auditing for reliable model evaluation.
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
WearableQA benchmark tests LLM health reasoning over longitudinal wearable data; the best of 14 evaluated LLMs reaches 72.9% accuracy.
WearableQA comprises 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning. Evaluation of 14 proprietary and open-source LLMs shows performance from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.
Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education
A Saudi university study finds students value ChatGPT writing feedback but treat human instructors as the final grading authority.
Thirteen male undergraduate computing students at a Saudi public university completed handwritten writing tasks that were scored by ChatGPT using a rubric-based prompt, then reflected after being told the score and feedback were AI-generated. Inductive thematic analysis identified four themes: perceived feedback usefulness, awareness of AI's contextual and pedagogical limitations, conditional trust, and reflection on the instructor's institutional role. Participants accepted GenAI feedback for surface-level revision but consistently positioned human instructors as the authority over grading decisions, distinguishing feedback utility from evaluative authority.