视听对话图:从自我中心-外部中心视角出发
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
浏览论文内容
中文总结 AI 辅助
本文提出从自我中心视频推断外部中心对话交互的新问题,并设计多模态 AV-CONV 框架,利用自注意力联合建模多人听说行为,实验验证其优于基线。
中文摘要 AI 辅助
近年来,自我中心视频相关研究的蓬勃发展为对话交互研究提供了独特视角,其中视觉和音频信号都发挥着关键作用。尽管以往大多数工作聚焦于学习直接涉及相机佩戴者的行为,我们提出了 Ego-Exocentric Conversational Graph Prediction(自我中心-外部中心对话图预测)问题,标志着首次尝试从自我中心视频推断外部中心对话交互。我们提出统一的多模态框架——Audio-Visual Conversational Attention(AV-CONV,视听对话注意力),用于联合预测相机佩戴者以及自我中心视频中出现的所有其他社交伙伴的对话行为,即说话和倾听。具体而言,我们采用自注意力机制对跨时间、跨主体和跨模态的表征进行建模。为验证该方法,我们在一个具有挑战性的自我中心视频数据集上开展实验,该数据集包含多说话者和多对话场景。结果表明,与一系列基线方法相比,我们的方法性能更优。我们还给出了详细的消融研究,以评估模型中每个组件的贡献。项目页面见 https://vjwq.github.io/AV-CONV/。
英文摘要
In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the camera wearer, we introduce the Ego-Exocentric Conversational Graph Prediction problem, marking the first attempt to infer exocentric conversational interactions from egocentric videos. We propose a unified multi-modal framework -- Audio-Visual Conversational Attention (AV-CONV), for the joint prediction of conversation behaviors -- speaking and listening -- for both the camera wearer as well as all other social partners present in the egocentric video. Specifically, we adopt the self-attention mechanism to model the representations across-time, across-subjects, and across-modalities. To validate our method, we conduct experiments on a challenging egocentric video dataset that includes multi-speaker and multi-conversation scenarios. Our results demonstrate the superior performance of our method compared to a series of baselines. We also present detailed ablation studies to assess the contribution of each component in our model. Check our project page at https://vjwq.github.io/AV-CONV/.
发表机构
- Georgia Tech(佐治亚理工学院)
- Meta Reality Labs(元宇宙实验室)
- GenAI, Meta(元公司生成式人工智能部门)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。