Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
Clinical AI system Doctorina achieved 82.0% primary-care diagnostic concordance versus 57.0% for physicians across 150 synthetic consultations.
The study compared Doctorina, eight physicians, and four standalone frontier language models on 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 diagnostic concordance versus 57.0% for physicians (25.0-point difference, 95% CI 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Kimi K3 ranked next on diagnosis, while Claude Opus 5 led the closely spaced management estimates among Opus, Doctorina and Kimi.
- Doctorina: 82.0% Top-1 diagnostic concordance vs 57.0% for physicians
- Workup and treatment scores 89.4 vs 66.9 and 83.7 vs 61.2
- Kimi K3 ranked second on diagnosis among the groups
- A second Doctorina run reproduced the advantage across all outcomes
Full article130 words · extracted from arxiv.org · click to collapse
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.09070