WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
WearableQA is a new benchmark of 4,084 ten-option multiple-choice questions built from real longitudinal wearable data (200 users, up to 500 days of daily measurements); across 14 evaluated LLMs, accuracy spans 19.6% to 72.9% against a 10% chance baseline,…
Two reports covering the same paper (Hugging Face daily papers, 2026-09-03; arXiv cs.AI/cs.LG/cs.CL, 2026-09-04) are fully consistent and merge without conflicts. WearableQA is a benchmark of 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements, using real longitudinal wearable data that preserves device noise and variability. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning, grounded through a dual-grounding framework that combines literature with population-validated physiological patterns. Evaluations of 14 proprietary and open-source LLMs show accuracy from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%, indicating that health reasoning over real-world wearable data remains far from solved.
- 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements.
- Benchmark uses real longitudinal wearable data, preserving device noise and variability.
- 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning.
- Dual-grounding framework combines literature with population-validated physiological patterns.
- 14 proprietary and open-source LLMs evaluated; accuracy ranges from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.
Coverage timelineoldest first · each row is one article
- · 13d agoWearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Hugging Face daily papers· 35
WearableQA benchmark introduces 4,084 questions over real longitudinal wearable data, showing 14 LLMs score 19.6-72.9% on health reasoning, far from solved.