arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自然场景家庭视频的自闭症谱系筛查的确定性证据层视觉-语言模型

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

Wenqi Li, Mindi Ruan, Chuanbo Hu, Shuo Wang, Xin Li

arXiv 2610.09217首次发表:更新:

发表机构

University at Albany, SUNY; West Virginia University; Washington University in St. Louis(纽约州立大学奥尔巴尼分校; 西弗吉尼亚大学; 圣路易斯华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自闭症筛查中视觉-语言模型判断不稳定的问题,提出将决策移出模型,利用接地感知约束生成事件表,经年龄校准和确定性证据评分,在家庭视频上达到86%准确率,显著优于零样本基线。

AI 中文摘要

自闭症谱系障碍(ASD)的诊断依赖于专家对儿童社交行为的观察,而获取这种专业能力是早期识别的瓶颈。视觉-语言模型(VLMs)能够很好地从视频中描述儿童的行为;但基于描述得出的判断结果不稳定:在温度设为0的情况下,在同一个骨干网络上跨八种流水线配置,16%至37%的片段在重复运行之间预测标签发生变化,原因在于服务堆栈。我们保持VLM冻结,并将决策移出模型。在接地感知约束下,VLM生成一个事件表,包含带时间戳、有术语标签的事件,记录引发刺激并记录反证;一个纯文本阶段对每一行的置信度进行年龄校准;一个确定性的证据权重评分器将证据汇总并分层为风险类别,因此每个决策可分解为具名的逐特征贡献,并可依据保存的表重新评分。在43个由照护者录制、无协议约束的学龄前儿童家庭自由玩耍片段上,该流水线在三次运行中达到AUC 0.851±0.012、86.0%的准确率和F1 71.8。它在每次运行中正确标记74.4%的片段(零样本基线为60.5%),并且在每次运行中未标记任何典型发育儿童片段(零样本为31个中的9个)。在同一骨干网络上的消融实验将增益归因于通过确定性评分器读取的接地感知约束。

英文摘要

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0\% accuracy, and F$_1$ 71.8 over three runs. It labels 74.4\% of clips correctly in every run (60.5\% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑