Get Started
Home
Topics
Search
Library
6 min read · Evaluation · Reasoning · Added Sep 13 · Paper published Sep 4, 2026

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Source: research paper via Hugging Face Daily Papers
WearableQA attacks a real gap: benchmarks can’t tell if a health LLM computes over 500 days of noisy sensor traces or just recites physiology. Across 14 models, giving GPT-5.4 Python execution jumps accuracy from 51% to 71%, while frontier models score +48 points on textbook signal pairs versus empirical ones.
TL;DR
WearableQA tests whether LLMs can reason over hundreds of days of real wearable sensor traces plus blood panels for 200 users, using 4,084 10-option multiple-choice questions whose answers are re-derived by deterministic code from each user’s actual measurements.
Why It Matters
Health-assistant products increasingly want to answer questions like “is my sleep hurting my recovery?” by ingesting months of watch data, blood work, and demographics. To trust an LLM here, you need to know if it can actually compute over noisy longitudinal signals, or if it’s just retrieving textbook health facts.
Existing benchmarks don’t tell you. Medical QA sets like MedQA test exam knowledge from text. Time-series benchmarks like TimeSeriesExam use synthetic signals. The closest prior work on wearables, PHIA, leans on simulated user trajectories and retrieval-style questions. None of them put a model in front of a real person’s 500-day record and ask questions where the correct answer is fixed by that person’s actual numbers.
That’s the gap: no one had a benchmark where the ground truth is deterministically computable from a real user’s messy sensor history, and where questions separate “can you do the math on the signals” from “do you know the physiology.”
How It Works
The dataset is built around 200 real users, each with up to 500 days of daily wearable metrics (resting heart rate, HRV, steps, sleep stages, VO2max, etc.), a 17-marker blood panel, and demographics. Questions come from two sources the authors call dual grounding:
•
Literature-grounded: 11 peer-reviewed wearable studies, each independently verified against PubMed and required to have at least two same-direction replications. Study-specific numeric cutoffs are deliberately not imported, because those don’t transfer across cohorts.
•
Population-grounded: patterns mined from a larger user cohort over 28-day windows, then filtered through three statistical gates: effect size (correlations must exceed |ρ|≥0.5 with consistent sign in both halves of the window), robustness (survive bootstrap resampling and leave-one-out), and authenticity (a cross-user null test bounding False Discovery Rate at ≤0.20).
Once a reasoning objective exists, the answer for each user is produced by a deterministic program over that user’s real measurements, not asserted from a template. The authors call this discover-then-label and build the programs from a shared library of primitives (lagged correlation, threshold flags, argmax over pairs, etc.). Pseudocode for a “which two signals are most coupled?” question:
def strongest_pair(user_series, signals): scores = {} for a, b in itertools.combinations(signals, 2): scores[(a,b)] = lagged_correlation( user_series[a], user_series[b], max_lag=3) gold = argmax_pair(scores) # keep item only if gold uniquely satisfies the computation return gold if is_unique_winner(scores, margin=0.15) else DROP
Each question is 10-way multiple choice. Distractors for data-reasoning questions are generated by running the same computation and picking wrong-but-plausible values (opposite trend, adjacent window, wrong magnitude); an item is kept only if the gold option is the unique winner. Distractors for health-reasoning questions are alternative clinical interpretations of the same physiological finding.
Questions are indexed on two axes: data vs. health reasoning (compute from raw signals vs. clinical interpretation) and single- vs. cross-signal (one metric vs. integrating two or more). That gives 16 question types. Validation includes three-model review (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) plus human check, and shortcut auditing that caught issues like same-source signal pairs (steps ↔ active burn) being answerable from priors alone.
What They Found
14 models were evaluated with chain-of-thought prompting, given each user’s 500-day history in a row-wise text format. Chance is 10%.
•
Overall accuracy spans 19.6% (Llama-3.2-3B) to 72.9% (Gemini-3.1-Pro). Most models land below 60%, so the benchmark is unsolved.
•
Data reasoning is harder than health reasoning for 13 of 14 models. GPT-4o gets 53.5% on health vs. 25.3% on data; Mistral-Small-3.1 gets 47.9% vs. 20.0%. Only Gemini-3.1-Pro reverses this (75.4% data, 67.8% health). The authors read this as: models know physiology but struggle to derive quantities from noisy longitudinal traces.
•
Cross-signal integration hurts stronger models specifically. GPT-5.4 drops from 59.1% single-signal to 45.7% cross-signal; Gemini-2.5-Pro drops 58.1% to 45.3%. Weaker open-source models are low on both, so the gap doesn’t show up.
•
Priors vs. data. On “which two signals are most strongly related?” questions, the authors split pairs into definitional (steps + active energy, knowable a priori) and empirical (resting HR + stress, only inferable from the data), controlling for actual coupling strength. Gemini-2.5-Pro is +51.5 points better on definitional; GPT-5.4 is +48.2. Gemini-3.1-Pro and Claude-Opus-4.6 show only ~4-point gaps. Interpretation: several models are pattern-matching to familiar pairs rather than reading the trace.
•
Input format barely matters; tool use matters a lot. Row/column/CSV/Markdown text formats for GPT-5.4 all land within ~2 points of each other (~50%). Image renderings of the time series hurt (down to 34-36%). Giving GPT-5.4 Python execution jumps it from 51.2% to 71.3% overall, with the biggest gain on cross-signal questions.
•
CoT helps proprietary models more. Claude-Opus-4.6 gains +19.6 points over direct-answer prompting; open-source models average only +2-3 points. Direct-answer also amplifies positional bias: Llama-3.2-3B picks a single option letter 51.5% of the time under direct prompting; CoT brings that back to 13.5%.
•
Ablating the time series (keeping only demographics, blood, cohort refs) drops Claude-Opus-4.6 from 60.2% to 17.3% overall, and from 58.1% to 13.7% on window-scoped questions, confirming answers really do depend on the signals.
What’s Useful
•
If you’re evaluating a wearable-health assistant, this benchmark (GitHub, HuggingFace) gives you a discriminative test that separates “knows physiology” from “can actually compute over the trace.” The 2×2 taxonomy lets you localize failures instead of just seeing an aggregate score.
•
The Python-tool result is the most actionable engineering finding. Text-only prompting caps GPT-5.4 near 51%; giving it code execution pushes it to 71%. If you’re building a health-reasoning agent, invest in a tool loop over a scratchpad-style long context before you invest in prompt-format tuning. Worth testing on your own health-reasoning workload.
•
Treat model-reported cross-signal insights skeptically if the pair is a well-known textbook coupling. The definitional-vs-empirical gap (up to +51.5 points) means several frontier models can produce a confident answer about steps ↔ energy expenditure without actually looking at the user’s data. Empirical couplings are where the model is forced to read the trace.
•
For smaller open-source models, always use CoT and check for positional collapse before trusting numbers. Direct-answer on a 3B model here was essentially degenerate.
Caveats
•
Questions are 10-way multiple choice, which measures discrimination among structured options, not free-form clinical reasoning or safety of open-ended advice.
•
The 200-user cohort is described as diverse but small; cohort-reference percentiles are drawn from a larger pool that the paper doesn’t fully characterize. Generalization to other populations, devices, or wearable brands isn’t established.
•
Population-grounded questions rely on the FDR≤0.20 gate. That’s a stated 80% positive predictive value, not certainty, so a fraction of “gold” answers reflect statistical patterns rather than clinically causal ones.
•
The paper evaluates hosted models at temperature 0 on a single serialization by default; results could shift with different context lengths, retrieval scaffolds, or fine-tuning on similar data.
•
The strong agentic result is reported for GPT-5.4 only; whether Python access closes the gap for weaker models isn’t shown.
Topics
Evaluation
Reasoning
Meta FAIR
NLP
Evaluation
Reasoning
Meta FAIR
NLP
Up next in Evaluation
UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Reasoning121 episodes
Evaluation174 episodes
NLP116 episodes