Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue
LLM consensus on K-12 math failure modes is high, but human agreement shows it is not validity.
An exploratory study assesses LLM classifications of five student failure modes in K-12 mathematics tutoring dialogue using an operational diagnostic codebook. Human-LLM agreement was moderate, with kappa from 0.524 to 0.597, while cross-model agreement was substantially higher, kappa 0.755 to 0.781 and alpha 0.769. The authors conclude that consensus among models can look like correctness without proving that the inferred learner constructs are valid.
- Five modes: uncertainty, misattribution, operator selection, conceptual gap, procedural slip.
- Human-LLM agreement was moderate, kappa 0.524 to 0.597.
- Cross-model agreement was higher, kappa 0.755 to 0.781 and alpha 0.769.
- Authors say model consensus cannot replace independent validity evidence.
Full article170 words · extracted from arxiv.org · click to collapse
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08703