arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23838cs.CVcs.AIcs.LG

用于基于Transformer的干预措施分类的婴儿护理视频数据集

Infant Care Video Dataset for Classification of Interventions Using Transformers

Igor Bogdanov, James Green

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对NICU医疗文档记录难题,构建含4144个视频的ICVD数据集,用TimeSformer等视频Transformer取得93.97%等准确率,证实时间建模对婴儿护理干预分类的重要性。

中文摘要 AI 辅助

新生儿重症监护室(NICU)的医疗文档记录存在重大挑战,护士约25%的时间用于记录工作,而多达60%的干预措施未被记录。受自动从视频中检测干预措施的需求驱动,我们提出了婴儿护理视频数据集(ICVD),该数据集包含4144个视频,涵盖12类模拟干预措施,旨在开发自动文档记录系统。我们基于人体模型的方法系统地改变了相机角度、临床医生肤色等条件,同时确保符合隐私要求。我们使用视频Transformer架构TimeSformer和MotionFormer,在12类婴儿护理类别中建立了较强的基准性能,Top-1准确率分别为93.97%和93.17%。我们的消融研究将时间模型与逐帧方法(准确率23.17%)进行比较,显示出70.80%的性能差距,验证了时间建模的必要性。ICVD为开发自动文档记录系统提供了基础,以减少新生儿护理环境中的临床负担并改进现有实践。

英文摘要

Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges, with nurses spending approximately 25\% of their time on record-keeping, while up to 60\% of interventions remain undocumented. Motivated by the need to detect interventions from video automatically, we present the Infant Care Video Dataset (ICVD), a collection of 4,144 videos spanning 12 simulated intervention classes designed for developing automated documentation systems. Our manikin-based approach systematically varies conditions, such as camera angle and clinician skin tone, while ensuring privacy compliance. Using video transformer architectures (TimeSformer and MotionFormer), we establish strong baseline performance (93.97\% and 93.17\% top-1 accuracy) among the 12 infant care classes. Our ablation study comparing temporal models with a framewise approach (23.17\% accuracy) demonstrates a 70.80\% performance gap, validating the need for temporal modeling. The ICVD provides a foundation for developing automated documentation systems to reduce clinical burden in neonatal care environments and improve existing practices.

发表机构

  • Carleton University(卡尔顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑