发表机构
Meta; KAIST; Korea University(Meta; 韩国科学技术院; 高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究推出WearableQA基准,包含4084道基于200名用户真实可穿戴数据的选择题,评估14种LLM的推理能力,发现多数模型准确率低于60%,该基准可有效区分模型能力。
AI 中文摘要
可穿戴传感技术的最新进展实现了对生理和行为信号的连续监测,但现有基准很少评估AI系统能否对真实用户的纵向可穿戴记录进行推理。我们推出WearableQA,这一基准包含4084道10选项的多项选择题,这些题目由200名真实用户的可穿戴时间序列、血液生物标志物和人口统计数据构建而成,每位用户拥有最多500天的每日测量数据。WearableQA保留了真实的可穿戴数据分布,其中包含设备噪声和个体间差异。为评估不同的推理能力,我们引入了16种问题类型,这些类型沿两个互补轴组织:数据推理与健康推理,用于区分对纵向测量的计算和对生理的解释;以及单信号推理与跨信号推理,用于区分对单个信号的推理和对多个信号的整合。为大规模构建可靠的问题,我们采用了双基准框架,结合基于文献的生理发现和经统计验证的基于人群的生理模式,这使得能够捕捉真实世界可穿戴数据中观察到的有意义关系。对14种专有和开源大语言模型(LLM)的评估表明,WearableQA能够有效区分模型能力,其性能在19.6%至72.9%之间,而10%为随机猜测基准。此外,WearableQA仍远未被解决:大多数模型的准确率低于60%。总体而言,WearableQA为评估LLM对真实世界可穿戴数据的推理能力提供了一个现实且具有诊断性的基准。
英文摘要
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.