arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27219cs.CL

BALMS:面向纵向心理健康感知的智能体大语言模型基准测试

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

  • Dartmouth College(达特茅斯学院)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell

AI总结:

该研究提出首个面向纵向心理健康感知的智能体大语言模型基准BALMS,经多数据集、任务与主干评估发现零样本智能体表现有限,思维链提示有提升但存不足,凸显了对特定智能体的需求。

AI中文摘要:

心理健康评估依赖于阶段性自我报告量表,这类量表将压力等主观状态转化为数值评分,但仅能提供福祉的稀疏快照。可穿戴设备提供纵向行为与生理信号,用于连续、低负担的监测。近期大语言模型(LLM)驱动的个人健康智能体支持对可穿戴信号的自然语言查询,但主要处理短期、基于检索的查询(如一周内最高步数),未评估智能体能否对长期信号进行推理以预测福祉评分并提供有证据支撑的理由。为填补这一空白,我们提出BALMS,首个针对基于LLM的智能体系统用于纵向心理健康感知的系统性基准。BALMS涵盖3个真实世界纵向数据集、2类任务(封闭式福祉评分预测及由LLM作为评判者自动评分的理由生成)、3类智能体范式,在5个开源与闭源LLM主干上进行评估。我们发现,零样本智能体除了在使用更强主干或紧凑语义特征时,几乎无法优于简单均值基线。思维链提示可提升面向推理的主干性能,但无法保证时间关联性或数值正确性。结合效率与时间缩放的更多分析,BALMS凸显了对选择性检索历史、基于时间证据并对可解释行为特征进行推理的纵向心理健康智能体的需求。

英文摘要:

Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiological signals for continuous, low-burden monitoring. Recent LLM-driven personal-health agents enable natural language queries over wearable signals, but mainly handle short-term, retrieval-based lookups (e.g., highest step count over a week). They do not evaluate whether agents can reason over long-term signals to predict wellbeing scores paired with evidence-grounded rationales. To address this gap, we introduce BALMS, the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing. BALMS spans 3 real-world longitudinal datasets, 2 task families (closed-form wellbeing-score prediction and rationale generation auto-graded by an LLM-as-Judge), 3 agentic paradigms evaluated across 5 open- and closed-source LLM backbones. We find that zero-shot agents rarely outperform a simple mean baseline, except with stronger backbones or compact, semantically meaningful features. Chain-of-thought prompting improves reasoning-oriented backbones, but does not guarantee temporal grounding or numerical correctness. Together with more analysis on efficiency and temporal scaling, BALMS highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.

补充信息

↑