在对我说话还是对别人说话?重新思考第一视角视频中的“对我说话”检测
Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos
浏览论文内容
中文总结 AI 辅助
本文重新定义第一视角视频中的“对我说话”检测为在线帧级任务,构建含约90万标注帧的数据集,提出融合音频、视觉和语音语义的多模态模型,帧级F1达75.5%。
中文摘要 AI 辅助
在线理解谁在对相机佩戴者说话,是第一视角社交互动的关键能力。然而,现有的“对我说话”(TTM)研究通常被表述为离线片段级识别,这与在线互动的一致性较差,并且忽视了第一视角视频中自然出现的各种非TTM说话状态。在本文中,我们通过将这一问题重新表述为在线、帧级预测任务来重新审视它。我们不再将TTM视为针对单一负类别的二分类问题,而是在存在多种且此前未被充分探索的非TTM状态(如对他人说话、自言自语和背景条件)的情况下对其进行建模。为支持这一新表述,我们构建了一个在线TTM数据集,包含406个第一视角视频片段,约90万帧带标注的帧,每帧都标注了帧级社交互动类别(如背景、TTM、对他人说话、自言自语),这是通过扩展Ego4D社交互动基准实现的。在该基准上,我们评估了五个改编的基线模型,并开发了一个整合跨模态社交线索的新模型。实验结果表明,我们的多模态模型联合利用音频、视觉和语音语义线索,在TTM上实现了75.5%的帧级F1分数,优于强基线,并能够系统分析不同说话状态如何影响TTM识别。
英文摘要
Online understanding of who is talking to the camera wearer is a key capability for egocentric social interaction. However, existing talk-to-me (TTM) studies are commonly formulated as offline clip-level recognition, which is poorly aligned with online interaction and overlooks the diverse non-TTM speaking states that naturally arise in egocentric videos. In this paper, we revisit this problem by reformulating it as an online, frame-level prediction task. Instead of treating TTM as a binary problem against a single negative class, we model it in the presence of diverse and previously underexplored non-TTM states, such as talking-to-others, self-talking, and background conditions. To support this new formulation, we construct an Online TTM Dataset consisting of 406 egocentric video clips with approximately 900K annotated frames, each labeled with frame-level social interaction categories (e.g., background, TTM, talking-to-others, self-talking), by extending the Ego4D social interaction benchmark. In this benchmark, we evaluate five adapted baselines and develop a new model that integrates social cues across modalities. Experimental results show that our multimodal model, which jointly leverages audio, visual, and speech-semantic cues, achieves 75.5% frame-level F1 on TTM, outperforming strong baselines and enabling a systematic analysis of how different speaking states affect TTM recognition.
发表机构
- The University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。