arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17056cs.SDcs.CL

鸡尾酒会场景中的音视频轮流说话预测

Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios

  • Trinity College Dublin(都柏林圣三一大学)

机构由 AI 辅助整理,请以论文原文为准。

Long-Vu Hoang, Naomi Harte

AI总结:

本研究评估了音视频轮流说话预测模型在鸡尾酒会嘈杂场景下的泛化与适应能力,发现性能显著下降(加权F1最高降38%),微调可改善但效果因模态和数据量而异,并公开了代码与标签。

AI中文摘要:

当前的轮流说话预测模型(PTTMs)在受控声学条件和干净音频信号的基准测试中表现出色,但其在具有重叠语音和背景干扰的对话中的泛化能力仍未得到充分探索。在本研究中,我们评估了使用干净数据训练的音视频PTTMs在源自AVCocktail数据集的具有挑战性的鸡尾酒会测试平台上的表现,并分析了它们对该新领域的适应行为。实验结果显示,在嘈杂条件下,音频和视觉模态的性能均出现持续下降,加权F1分数最高相对下降38%。在新领域上进行微调可提高鲁棒性,但收益因模态而异,并取决于可用预训练数据的大小。这些发现提供了对音频和视觉模态不同泛化与适应能力的见解,并表明需要鲁棒的建模策略以适应噪声中人类交互的复杂性。所有代码和轮流标签均已公开,以促进进一步研究。

英文摘要:

Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.

补充信息

↑