Reasoning Models Are Accurate but Unsound on Identification
CERTID benchmark shows frontier reasoning models give unsound causal identification answers, with false-claim rates varying seventeen-fold across Gemini Flash, Gemini Pro, and GPT-5.5.
The paper introduces CERTID, a formal pipeline that certifies causal identification queries via the sound and complete ID algorithm and verifies returned formulas against structural causal models with exactly known interventional distributions. Three frontier reasoning models (Gemini Flash, Gemini Pro, GPT-5.5) were evaluated on 1,200 certified instances spanning graphs of 4 to 50 vertices. Accuracy proved a poor proxy for soundness, with false-claim rates on non-identifiable queries varying seventeen-fold across models. Models decided identifiability with 97-100% accuracy even on graphs generated after the strongest model's training snapshot.
- CERTID certifies identifiability using the ID algorithm and verifies formulas against known SCMs
- False-claim rate on non-identifiable queries varies seventeen-fold across frontier models
- Evaluation covers 1,200 certified instances on graphs of 4 to 50 vertices
- 97-100% identifiability accuracy even on post-training-snapshot graphs
Full article208 words · extracted from arxiv.org · click to collapse
A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at https://anonymous.4open.science/r/certid-D718.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03519