ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Sho Kawano

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

infoAI researchimportance 14
AI summary · glm-5.3-flash

Researchers propose prediction-powered smoothing (PP-S/PP-TS), Bayesian small-area estimators that improve disaggregated AI evaluation accuracy and validation.

The paper treats AI evaluation sets as finite populations and improves domain-level mean estimates by combining prediction-powered inference with small area estimation. PP-S fits a Bayesian model to each domain's prediction-powered estimate, while PP-TS borrows strength across a reporting taxonomy. A new approximately unbiased design-based cross-validation score selects between direct and smoothed estimators, delivering better point and interval estimates with near-nominal coverage on a curated benchmark and human-graded deployed agent traffic.

  • PP-S fits Bayesian models to prediction-powered domain estimates
  • PP-TS borrows strength across a reporting taxonomy
  • New cross-validation score selects estimators without independent samples
  • Near-nominal coverage on curated benchmark and agent traffic
Full article203 words · extracted from arxiv.org · click to collapse

Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage. At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator's error far more accurately.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.20758