TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
TasteVal finds Opus 5.5 beats expert baselines on experimental research taste with a 2.3x compute multiplier.
TasteVal measures experimental research taste—designing experiments and interpreting results—on eight novel frontier AI R&D tasks, with a fixed Coder agent so coding skill is isolated. Twenty-four human experts provided baselines, and 20 models from 2023–2026 ran until a 40 H100-hour or 120 wall-clock-hour budget was exhausted. Opus 5.5 beat the best-expert baseline with a 2.3x compute multiplier (95% CI 1.15–4.37) at roughly 1/30 of the baseliners' average per-run cost. Since December 2025 that multiplier has doubled about every 3.0 months; the tasks are not released.
- Eight tasks separate experiment design from coding via a fixed coder
- Budgets are 40 H100 hours or 120 wall-clock hours
- Opus 5.5 posts a 2.3x compute multiplier over expert baselines
- Since December 2025 that multiplier doubled about every 3.0 months
- Tasks are withheld to keep the benchmark uncontaminated
Full article282 words · extracted from arxiv.org · click to collapse
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06824