发表机构
Lund University; KTH Royal Institute of Technology(隆德大学; KTH皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过扩展VAP模型融入视觉特征,证明多模态信息能提升面对面对话中的轮流发言预测,其中面部动作单元最有效,且结合所有特征效果最佳。
AI 中文摘要
轮流发言是口语交互的基本组成部分,虽然人类自然依赖言语和非言语信号,但对话系统通常仅依赖音频线索。本文研究面对面对话中的视觉特征是否能超越仅依赖音频所能达到的效果,增强轮流发言预测。我们扩展了语音活动投影(VAP)模型——一个基于自监督Transformer的模型,用于预测未来的语音活动——通过融入从大规模Meta Seamless Interaction数据集中提取的二元面对面对话视觉特征。视觉特征包括注视方向、头部运动、身体和手部姿态以及面部动作单元(FAU)。为融入视觉特征,我们探索了拼接、交叉注意力融合、增量特征和可训练门控机制。结果表明,视觉信息相比仅音频基线提升了性能,其中FAU比其他特征组更具信息量。然而,身体和注视特征贡献了互补信息,因为结合所有特征的模型表现最佳。此外,结果表明,具体任务的性能取决于训练和测试数据来自即兴(表演)还是自然(非表演)对话。
英文摘要
Turn-taking is a fundamental component of spoken interaction, and while humans naturally rely on both verbal and non-verbal signals, dialogue systems usually depend on audio cues alone. This paper investigates whether visual features from face-to-face conversations can enhance turn-taking prediction beyond what is achievable from audio-only. We extend the Voice Activity Projection (VAP) model, a self-supervised transformer-based model for predicting future voice activity, by incorporating visual features extracted from the large-scale Meta Seamless Interaction dataset of dyadic face-to-face conversations. The visual features include gaze direction, head movement, body and hand pose, and facial action units (FAU). For incorporating the visual features, we explore concatenation, cross-attention fusion, delta features, and trainable gating mechanisms. Results show that visual information improves performance over the audio-only baseline, with FAU being significantly more informative than other feature groups. Body and gaze features nevertheless contribute complementary information, as the model combining all features performs best. Furthermore, results indicate that performance on specific tasks varies depending on whether training and test data come from improvised (acted) or naturalistic (non-acted) conversations.
CommentsAccepted to EMNLP 2026 Workshop on Multimodal Interaction in Face-to-Face Dialogue (MINT)