arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14029cs.CL

S2Dialog:结合语义与声学风格建模的多模态对话检索

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

  • College of Computer Science, Inner Mongolia University(内蒙古大学计算机学院)
  • University of Electronic Science and Technology of China, Shenzhen Campus(电子科技大学深圳校区)

机构由 AI 辅助整理,请以论文原文为准。

Xueqi Wang, Zhigang Wang, Runqing Zhang, Zhenqi Jia, Junfeng Zhao

AI总结:

研究针对现有多模态对话检索方法无法捕捉对话全局语义与风格一致性的问题,提出S2Dialog框架,结合对话级文本与声学检索器及对比学习,在DailyTalk数据集上实现出色检索性能。

AI中文摘要:

多模态对话检索旨在从多模态对话库中检索出与目标对话在文本语义和声学对话风格两方面都相似的对话。这种对话级检索对诸多对话相关任务至关重要,包括对话情感识别、口语对话系统和对话式语音合成,外部对话示例可为这些任务提供有价值的语义和风格参考。然而,现有检索方法大多局限于语句级或单模态匹配,往往无法捕捉整个对话的全局语义连贯性和风格一致性。为解决这一差距,我们提出S2Dialog,这是一个用于从多模态对话库中进行对话级语义风格检索的统一框架。具体而言,S2Dialog由对话级文本检索器和对话级声学检索器组成,分别将对话的文本和声学模态编码为对话级表示。为进一步增强多模态检索,我们引入对话级文本-声学对比学习,该方法使语义和风格相似的对话相互对齐,同时区分不相关的对话。在多模态对话数据集DailyTalk上进行的大量实验表明,S2Dialog取得了出色的检索性能。

英文摘要:

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level retrieval is crucial for many dialogue-related tasks, including Emotion Recognition in Conversation, Spoken Dialogue Systems, and Conversational Speech Synthesis, where external dialogue examples can provide valuable semantic and stylistic references. However, existing retrieval methods are still largely limited to utterance-level or unimodal matching, and often fail to capture the global semantic coherence and stylistic consistency of an entire dialogue. To address this gap, we propose S2Dialog, a unified framework for dialogue-level semantic-style retrieval from multimodal dialogue banks. Specifically, S2Dialog consists of a Dialogue-level Textual Retriever and a Dialogue-level Acoustic Retriever, which encode the textual and acoustic modalities of a dialogue into dialogue-level representations, respectively. To further enhance multimodal retrieval, we introduce Dialogue-level Textual-Acoustic Contrastive Learning, which aligns semantically and stylistically similar dialogues while distinguishing unrelated ones. Extensive experiments on the multimodal dialogue dataset DailyTalk demonstrate that S2Dialog achieves outstanding retrieval performance.

↑