发表机构
Macquarie University; University of Sydney; University of California, Merced(麦考瑞大学; 悉尼大学; 加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对痴呆症护理中BPSD行为理解缺乏视频基准的问题,提出DementiaCare-Bench,含94个片段和2023个问题,验证VLM视觉需求,并证明LoRA微调可修复视频依赖缺陷。
AI 中文摘要
痴呆症影响全球约5700万人,对大多数家庭而言,护理中最困难的部分并非记忆丧失,而是痴呆症的行为和心理症状(BPSD):躁动、游荡、抗拒护理、日落综合征。理解这些症状不仅需要识别行为本身,还需要了解行为发生前的情况。同一行为可能因触发因素不同而需要不同的应对方式。视频语言模型(VLMs)有潜力为护理人员提供支持,但目前尚无基准评估这一能力。为填补这一空白,我们提出了DementiaCare-Bench:包含56个专业制作的护理人员培训视频,分割为94个片段,涵盖九类BPSD,通过多智能体流程生成2023个问题,每个临床主张均基于逐字转录片段。每个问题在四种视觉条件下进行探查,并按其最低需求进行标注,从而衡量而非假设其视觉需求。测量结果与意图相悖:我们编写了77.7%的问题要求有序帧,而实际仅有34.8%的问题需要。在12个当前VLM上,模式一致。最佳模型总体达到85%,但该平均值由仅凭临床知识即可回答的语言模型问题所拉高;在需要有序片段的问题上,平均准确率下降17个百分点,且一个领先的开源模型在判断护理人员回应是否适当时,得分仅相当于随机水平。轻量级LoRA微调模型DemCare-VLM将视频依赖性从-3.3分提升至+4.5分,表明基准所揭示的问题不仅可以被测量,还可以被修复。
英文摘要
Dementia affects an estimated 57 million people worldwide, and for most families the hardest part of care is not memory loss but the behavioral and psychological symptoms of dementia (BPSD): agitation, wandering, resistance to care, sundowning. Understanding these symptoms requires more than recognizing the behavior itself; it also requires knowing what happened beforehand. The same behavior may call for a different response depending on its trigger. Video-language models (VLMs) could potentially support caregivers, yet no existing benchmark evaluates this capability. To fill this gap, we present DementiaCare-Bench: 56 professionally produced caregiver training videos segmented into 94 clips across nine BPSD categories, with 2023 questions generated by a multi-agent pipeline that grounds every clinical claim in a verbatim transcript span. Each question is then probed under four visual conditions and labelled by the least it requires, so its visual demand is measured rather than assumed. Measurement contradicts intent: we wrote 77.7% of the questions to require ordered frames, and 34.8% do. Across 12 current VLMs the pattern is uniform. The best reach 85% overall, but that average is carried by questions a language model can answer from clinical knowledge alone; accuracy falls by 17 points on average on questions that require the ordered clip, and a leading open model scores at chance on judging whether a caregiver's response was appropriate. A lightweight LoRA fine-tune, DemCare-VLM, moves video dependence from -3.3 to +4.5 points, so what the benchmark exposes can be repaired and not only measured.