Get Started
Home
Topics
Search
Library
Research questionHow can language models reason over noisy, longitudinal wearable records while integrating multiple physiological signals?Real-world wearable records combine long time series, device noise, and substantial variation between users. Health reasoning requires interpreting measurements over time and connecting evidence across physiological signals.
AI
Evaluation & Benchmarks
Health
Reasoning
Latest papersRecent research connected to this question, newest first.WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA provides 4,084 ten-option questions based on wearable time series, blood biomarkers, and demographics from 200 users with up to 500 days of daily measurements. Its evidence covers 16 question types grounded in physiological literature and statistically validated population patterns; evaluations of 14 proprietary and open-source LLMs range from 19.6% to 72.9% accuracy against a 10% chance baseline.research paper · Sep 4, 2026
Related questions
How can wearable foundation models organize representations for targeted longitudinal women’s-health prediction?How can noisy, partial wearable and mobile measurements be fused into a continuous physiological-state biomarker without sacrificing simple-fusion performance?How can we evaluate LLM clinical reasoning across multilingual, multimodal clinical time series?How can language models reason iteratively to diagnose complex clinical cases safely and accurately?