发表机构
Kyoto University(京都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究聚焦虚拟角色对话系统中倾听者点头动作,提出由时间预测与运动学参数预测模块构成的模型,基于VAP技术,能实时预测点头相关参数,优于随机时间和固定动作基线,且轻量级可实时运行。
AI 中文摘要
在人类对话中,通过适时表达眼神交流、点头和面部表情等非语言线索来实现顺畅沟通。期望对话虚拟角色能恰当表达这些线索以实现自然且类人的交互。本研究聚焦点头,其对展示积极倾听和鼓励用户进一步发言至关重要。我们提出一个实时预测表示倾听者点头动作特征的时间和运动学参数的模型。该模型由时间预测模块和运动学参数预测模块组成,各模块基于语音活动投影(VAP)技术在说话者和倾听者通道上实现二元注意力网络。与传统模型不同,此方法能基于对话特定上下文实时预测运动学参数而非仅预测时间。此外,我们证明了从训练好的时间预测模块初始化微调运动学参数预测模块的有效性。所提模型轻量级且能实时运行,并已集成到虚拟角色对话系统中。主观评估实验表明,我们的方法显著优于随机时间基线和固定动作点头基线。代码和训练模型可通过此https链接获取。
英文摘要
In human dialogue, we achieve smooth communication by expressing nonverbal cues such as eye contact, nodding, and facial expressions with precise timing. It is expected for conversational avatars to express these cues appropriately to realize natural and human-like interactions. This study focuses on nodding, which is crucial for demonstrating active listening and encouraging further user utterances. We propose a model that predicts both timing and kinematic parameters representing the motion features of listener nodding in real time. The proposed model consists of a timing prediction module and a kinematic parameter prediction module. Each implements a dyadic attention network over the speaker and listener channels based on the technique of Voice Activity Projection (VAP). Unlike conventional models, this approach enables real-time prediction of kinematic parameters based on the specific context of the dialogue rather than just predicting the timing. Furthermore, we demonstrate the effectiveness of fine-tuning the kinematic parameter prediction module initialized from the trained timing prediction module. The proposed model is lightweight and capable of real-time operation, and it has been integrated into an avatar dialogue system. Subjective evaluation experiments shows that our proposed method significantly outperforms both a baseline with stochastic timing and another with fixed-motion nodding. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.
CommentsAccepted by 28th ACM International Conference on Multimodal Interaction (ICMI '26), Long paper