StudentBench: AI and human tutoring yield equivalent GRE learning gains
StudentBench finds AI tutoring matches human GRE gains for 2,383 learners at far lower cost.
StudentBench evaluates whether large language models produce GRE learning gains equivalent to human tutoring, using a public platform with more than 175,000 student-AI messages. Across 2,383 participants, AI tutoring was statistically equivalent to expert human tutoring (p = .015), and the best AI tutor outperformed the human tutor on average in five of seven GRE domains. A second study collected 2,028 pairwise expert ratings of lesson plans and practice problems. One AI tutor matched human learning gains (p = .044) at roughly 918 times lower cost, USD 0.0052 versus USD 4.81 per percentage point gained.
- The study includes 2,383 participants and over 175,000 student-AI messages.
- AI tutoring matched expert human GRE tutoring with p = .015.
- The best AI tutor led in five of seven GRE domains on average.
- One tutor matched human gains at about 918 times lower cost.
Full article233 words · extracted from huggingface.co · click to collapse
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.28470