arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08915cs.CL

视频介导对话中不同伙伴可见性条件下的多模态信息性研究

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

发表机构马克斯·普朗克心理语言学研究所 · 乌得勒支大学 · 奥斯纳布吕克大学
查看机构详情
  • Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
  • Utrecht University(乌得勒支大学)
  • Osnabrück University(奥斯纳布吕克大学)

机构由 AI 辅助整理,请以论文原文为准。

Esam Ghaleb, Hugh Mee Wong, Kristina Kobrock

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对视频介导对话,构建基于语音、手势或两者的模型识别指称对象,发现手势可单独预测指称,融合模型在转录模型不确定时效果最佳,还揭示了伙伴可见性对互动的语用效应。

中文摘要 AI 辅助

情境化语言使用具有多模态和具身性,例如手势可携带语音信号中缺失或未明确的信息,但对话模型通常仅依赖文本转录。本研究探讨在不同伙伴可见性条件下,多模态对话中手势及其与语音结合所承载的指称信息数量。我们构建模型,基于语音转录、手势的骨骼表示或两种模态,识别视频介导指称沟通游戏中的预期指称对象。结果显示,仅手势即可预测预期指称对象,且当基于转录的模型不确定时,多模态融合最为有益;仅将学习到的表示与指称图像对齐进行训练,可进一步提升融合模型的性能。与人类互动数据对比发现,对话伙伴可见性会对手势产生和信息性产生语用效应,且在多轮重复互动中,语音和多模态表现存在同步效应,但手势表现无此效应。我们为人类对话中多模态信息的技术建模,以及通过训练模型表示分析人类互动数据做出了贡献。

英文摘要

Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.

↑