arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16396cs.CL

超越言语通道的否定表达:对话中的时间多模态相关性

Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue

Leon Hammerla, Patrick Schrottenbacher, Alexander Mehler

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过虚拟现实对话数据,利用时间序列模型探针,证明否定线索在非言语多模态行为中具有可测量的预测信息,且面部特征贡献最大。

中文摘要 AI 辅助

否定通常通过其语言实现来建模,尽管口语互动伴随着紧密协调的非言语行为。我们探究以口语否定线索为中心的情境是否包含可测量的多模态行为信息:这些情境能否在没有词汇或声学输入的情况下与匹配的控制情境区分开来,这些信息在时间上出现在何处,哪些模态携带这些信息,以及这些信息是否延伸至对话伙伴。我们研究了在虚拟现实中进行的27场人机对话访谈,包括时间对齐的注视、面部、头部、身体、手部和手指行为,以及964个带注释的否定线索。将分类视为预测性探针,我们比较了20个时间序列模型,同时排除词汇和声学信息,然后系统地改变时间上下文、互动来源、模态可用性和事件时间。在分组10折交叉验证中,最强的探针从说话者侧行为达到了高达0.75的平均留出AUROC。时间分析表明,预测信息集中在线索起始周围,但在更广泛的周围间隔内仍可检测到,而对话伙伴行为携带的预测信息较弱,且时间分布相对分散。消融和时间扰动进一步表明,面部特征产生了最大的模态消融效应,并且训练好的探针对观察到的事件的时间组织敏感。

英文摘要

Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior. We ask whether contexts centered on spoken negation cues contain measurable multimodal behavioral information: whether they can be distinguished from matched control contexts without lexical or acoustic input, where this information occurs in time, which modalities carry it, and whether it extends to the dialogue partner. We study 27 human-human interviews conducted in virtual reality, comprising temporally aligned gaze, facial, head, body, hand, and finger behavior and 964 annotated negation cues. Treating classification as a predictive probe, we compare 20 time-series models while excluding lexical and acoustic information, and then systematically vary temporal context, interactional source, modality availability, and event timing. Across grouped 10-fold cross-validation, the strongest probes reach up to .75 mean held-out AUROC from speaker-side behavior. Temporal analyses show that predictive information is concentrated around cue onset but remains detectable over a broader surrounding interval, while dialogue-partner behavior carries weaker predictive information with a comparatively diffuse temporal profile. Ablation and timing perturbations further show that facial features produce the largest modality-ablation effect and that the trained probe is sensitive to the temporal organization of the observed events.

发表机构

  • Goethe University(歌德大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑