Researchers have introduced WearableQA, a new benchmark designed to evaluate the health reasoning capabilities of AI systems using real-world wearable data. The benchmark consists of over 4,000 multiple-choice questions derived from longitudinal wearable time series, blood biomarkers, and demographic data of 200 individuals. WearableQA is structured to test distinct reasoning skills, including data versus health reasoning and single- versus cross-signal integration, while preserving authentic wearable data distributions. Evaluations of 14 large language models showed a wide performance range, with most models scoring below 60%, indicating that the benchmark remains challenging. AI
IMPACT This benchmark could drive the development of more sophisticated AI models capable of understanding and interpreting complex health data from wearables.
RANK_REASON The cluster contains a research paper introducing a new benchmark for AI health reasoning. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →