发表机构
Goethe University; Leibniz Institute for Educational Trajectories (LIfBi)(歌德大学; 莱布尼茨教育轨迹研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InterView-C是一个德语多模态语料库,包含27次VR化身调查访谈,同步行为数据与人工转写,用于分析口语互动并解决ASR和迁移挑战。
AI 中文摘要
我们介绍了InterView-C,一个完全在虚拟现实环境中进行的27次调查访谈的德语多模态语料库,两位对话者均由化身代表。该语料库将口语互动与同步的行为数据对齐,包括注视、头部和身体运动、面部行为、手和手指追踪。其参考转写和语言注释为这种多模态口语互动与主要基于文本的自然语言处理方法之间提供了可靠的接口。这一接口之所以重要,是因为自动语音转写可能会扭曲语言学相关信息,而在现有资源上训练的下游模型在应用于转写的口语数据时,还可能面临迁移挑战。因此,InterView-C为所有54个录音提供了词级时间戳且经人工后期编辑的逐字转写、访谈项目时间、问卷回答以及1,422个句子的否定提示和范围注释,其中1,398个句子进行了双重注释(提示的α=0.87;范围的α=0.81)。我们通过实验证明了这两个挑战:九个开放权重自动语音识别系统对简短封闭式回答和数字词存在不成比例的错误识别,而在现有语料库上训练的否定模型在我们的转写访谈上表现低于且在InterView-C注释上训练的模型,且性能波动较大。因此,InterView-C能够在保留与丰富多模态行为对齐的同时,实现对口语互动的语言学分析。
英文摘要
We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.