arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉特征能否改善他人发起的修复检测?一种二元多模态方法

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloé Clavel

arXiv 2607.23845首次发表:更新:

发表机构

INRIA Paris; ISIR, Sorbonne University Paris; Télécom Paris, SES, Institut Polytechnique de Paris, I3-CNRS Paris; CNRS; LTCI, Institut Polytechnique de Paris(法国国家信息与自动化研究所巴黎分院; 巴黎索邦大学信息学、系统与机器人研究所; 巴黎电信学院、巴黎高等理工学院SES系、巴黎综合理工学院I3 - 法国国家科学研究中心; 法国国家科学研究中心; 巴黎综合理工学院LTCI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究他人发起的修复检测问题,提出结合视觉特征的二元多模态模型,通过在两个语料库上评估,证明视觉信息能提升检测性能,提供跨模态特征贡献见解。

AI 中文摘要

他人发起的自我修复,简称他人发起的修复(OIR),是对话交互中的一种重要机制,接收者通过它发出说话、听力或理解方面的问题信号,促使前一个说话者解决问题。对于对话代理来说,准确识别这些修复发起策略对于有效解决沟通故障至关重要。虽然对话分析研究表明,OIR发起伴随着言语和非言语信号,如目光转移、面部表情、身体姿势和手势,但现有的计算方法主要依赖文本和音频。本文引入了一种用于OIR检测和分类的新型多模态模型,纳入了从对话分析中提取的一组视觉特征。我们在两个具有不同语言和交互设置的语料库上评估了我们的方法。结果表明,视觉信息始终优于文本和音频基线,并且提供了对两个语料库中跨模态特征贡献的见解。

英文摘要

Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis studies have shown that OIR initiation is accompanied by both verbal and non-verbal signals such as gaze shifts, facial expressions, body postures, and hand gestures, existing computational approaches rely mainly on text and audio. This paper introduces a novel multimodal model for OIR detection and classification, incorporating a set of visual features drawn from conversation analysis. We evaluate our approach on two corpora with distinct languages and interaction settings. Results demonstrate that visual information consistently improves performance over text and audio baselines, and provide insights into cross-modal feature contributions across two corpora.

CommentsAccepted to ICMI 2026 (International Conference on Multimodal Interaction), October 5-9, 2026, Napoli, Italy

DOI:10.1145/3776574.3831200

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑