arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06353cs.CV

ChildGaze:儿童协作行为理解基准数据集

ChildGaze: A Benchmark Dataset for Collaborative Behavior Understanding in Children

Sindhuja Penchala, Saketh Reddy Kontham, Prachi Bhattacharjee, S. Nima Mahmoodi, Daniel Fonseca, Sareh Karami, Mehdi Garemani, Sudip Mittal, Shahram Rahimi, Noo… 展开作者

Sindhuja Penchala, Saketh Reddy Kontham, Prachi Bhattacharjee, S. Nima Mahmoodi, Daniel Fonseca, Sareh Karami, Mehdi Garemani, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz

首次发表
浏览论文内容

中文总结 AI 辅助

针对儿童协作行为理解,提出基于ChildPlay视频集的ChildGaze基准数据集,含协作/非协作标签及身体部位框,经可靠性和基线实验验证,支持自然互动研究。

中文摘要 AI 辅助

理解儿童的协作行为对于分析社交参与、同伴互动、共同注意以及游戏和学习活动中的参与度至关重要。可靠地识别这些线索可以支持儿童发展研究、教育分析以及以人为中心的计算机视觉研究。然而,估计儿童注视的位置并不一定能揭示该儿童是否积极参与共享活动。为了支持这种更高层次的分析,我们引入了ChildGaze,这是一个基于ChildPlay视频集[1]构建的以儿童为中心的行为标注数据集。ChildGaze引入了两种行为标签,即协作和非协作,在每一帧中独立分配给每个儿童。该数据集提供了儿童和成人的面部、左手和右手边界框,并在行、人和帧级别组织标注。当前版本包含27个标注视频文件、10,641帧和73,268条身体部位标注行。使用两名标注者的独立标注对1,187帧进行了标注可靠性评估。协作标签达到了93.16%的原始一致率和0.8631的Cohen's kappa系数,而边界框标注实现了0.808的总体平均IoU。使用预训练的ViT和Swin Transformer模型进行的基线实验分别达到了高达97.44%的儿童人物级别准确率和96.80%的帧级别准确率。这些结果表明,ChildGaze为研究自然环境中儿童-成人和同伴互动中的协作行为提供了一个可靠的基准。

英文摘要

Understanding collaborative behavior in children is important for analyzing social participation, peer interaction, shared attention, and engagement during play and learning activities. Reliable recognition of these cues can support research in child development, educational analysis, and human-centered computer vision. However, estimating where a child is looking does not necessarily reveal whether the child is actively participating in a shared activity. To support this higher-level analysis, we introduce ChildGaze, a child-centered behavioral annotation dataset built on the ChildPlay video collection [1]. ChildGaze introduces two behavioral labels, collaborative and non-collaborative, assigned independently to each child within a frame. The dataset provides face, left-hand, and right-hand bounding boxes for children and adults and organizes the annotations at the row, person, and frame levels. The current release contains 27 annotated video files, 10,641 frames, and 73,268 body-part annotation rows. Annotation reliability was evaluated on 1,187 frames using independent annotations from two annotators. The collaboration labels achieved 93.16% raw agreement and a Cohen's kappa of 0.8631, while bounding-box annotations achieved an overall mean IoU of 0.808. Baseline experiments with pretrained ViT and Swin Transformer models achieved up to 97.44% child-person-level accuracy and 96.80% frame-level accuracy, respectively. These results show that ChildGaze provides a reliable benchmark for studying collaborative behavior in naturalistic child-adult and peer interactions.

发表机构

  • The University of Alabama(阿拉巴马大学)
  • Mississippi State University(密西西比州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑