ZeroHour
Story · 1 source · 1 articlefirst updated ()

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

infoAI researchimportance 35
What's new: First consolidated summary for this story; no previous summary existed. The later arXiv report (2026-09-04) confirms every shared figure from the Hugging Face report (2026-09-03) - 4,084 questions, 200 users, up to 500 days, 16 question types, 14 LLMs, 19.6% to 72.9% versus a 10% chance baseline - and adds two details: the benchmark preserves real device noise and variability, and its…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

WearableQA is a new benchmark of 4,084 ten-option multiple-choice questions built from real longitudinal wearable data (200 users, up to 500 days of daily measurements); across 14 evaluated LLMs, accuracy spans 19.6% to 72.9% against a 10% chance baseline,…

Two reports covering the same paper (Hugging Face daily papers, 2026-09-03; arXiv cs.AI/cs.LG/cs.CL, 2026-09-04) are fully consistent and merge without conflicts. WearableQA is a benchmark of 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements, using real longitudinal wearable data that preserves device noise and variability. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning, grounded through a dual-grounding framework that combines literature with population-validated physiological patterns. Evaluations of 14 proprietary and open-source LLMs show accuracy from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%, indicating that health reasoning over real-world wearable data remains far from solved.

  • 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements.
  • Benchmark uses real longitudinal wearable data, preserving device noise and variability.
  • 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning.
  • Dual-grounding framework combines literature with population-validated physiological patterns.
  • 14 proprietary and open-source LLMs evaluated; accuracy ranges from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.
ProductsWearableQA
AI modelsWearableQA

Coverage timeline

  1. · 13d ago
    Hugging Face daily papers· 35
    WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

    WearableQA benchmark introduces 4,084 questions over real longitudinal wearable data, showing 14 LLMs score 19.6-72.9% on health reasoning, far from solved.