多轮对话中话轮转换结果预测:结合人际亲密度的语音与注视动态可解释建模
Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness
- Research & Exploration, GN Store Nord(GN Store Nord 研究与探索部门)
- Aalborg University(奥尔堡大学)
- Technical University of Denmark(丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究结合GaMMA语料库,用含注视、语音及人际亲密度特征的逻辑回归模型预测四人对话话轮转换结果,发现注视特征与响度结合可提升抗噪预测性能。
AI中文摘要:
顺畅的说话人转换是有效对话的基础,依赖于对话者预测何时加入对话的能力,而这种能力取决于准确解读和表达表明说话人希望获得或放弃话语权的语言与非语言线索。在嘈杂、自然的多对话者场景中,由于存在多个潜在对话者,该过程会变得更为复杂。本研究在自由四人对话中,结合注视、语音与感知到的人际亲密度,对对话话语权转换的信号进行建模。研究使用GaMMA语料库,在每个话轮转换事件前提取基于行为动机的可解释特征,训练逻辑回归模型,将话语权转移结果分类为间隙或重叠。预测变量包括注视特征(如转换 motif、行为对比、熵、基于注视的受话人身份、相互注视)、基于说话人响度的语音特征,以及说话人间感知到的人际亲密度(IOS)。结果显示,注视特征捕捉到预测结构,将其与响度结合可提升性能(ROC AUC = 0.76 ± 0.04);响度反映说话人控制权,而注视分散度与指向性则体现听话人准备状态与竞争性加入;在不同噪声条件下,性能保持稳定,表明注视为话轮转换动态提供了互补的、抗噪声的线索。
英文摘要:
Smooth speaker transitions are fundamental to effective conversation and rely on an interlocutor's ability to predict when to enter the conversation. This ability depends on accurately interpreting and expressing the verbal and non-verbal cues that signal when a speaker wishes to take or relinquish the floor. The process becomes even more complex in noisy, natural, multi-party settings, with multiple interlocutors available. This study models how gaze and speech, together with perceived interpersonal closeness, signal conversational floor changes in free four-person dialogue. Using the GaMMA corpus, we trained logistic regression models using interpretable, behaviourally motivated features extracted before each turn-taking event to classify floor-transfer outcomes as gaps or overlaps. Predictors included gaze features such as transition motifs and behavioural contrasts, entropy, gaze-based addressee identity, and mutual gaze, alongside speech features derived from speaker loudness, as well as perceived interpersonal closeness (IOS) between speakers. Results show that gaze features capture predictive structure, and that combining them with loudness improves performance (ROC AUC = 0.76 +- 0.04). Loudness reflected speaker control, while gaze dispersion and addressing indexed listener readiness and competitive entry. Performance remained robust across noise conditions, indicating that gaze provides a complementary, noise-resilient cue to turn-taking dynamics.