发表机构
Korea Institute of Energy Technology(韩国能源技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态RAG数据库场景,提出STeReO重排序器,构建专用数据集训练后,可有效聚合异构检索结果,显著提升下游问答性能。
AI 中文摘要
检索增强生成(RAG)系统因具备缓解大型语言模型(LLMs)幻觉的能力而受到广泛关注。尽管RAG的知识库正日益多样化,涵盖语音、文本等多种模态,但针对此类多模态数据库场景的研究仍有限。本文提出STeReO(语音与文本重排协调器),这是一种基于语音和文本检索器的重排序器,用于聚合不同模态的数据库。为解决专用训练数据缺失的问题,我们首先整理了包含查询、混合模态证据及其对应相关度排序的数据集,随后训练该重排序器并在单模态和混合模态场景中评估其有效性。结果表明,所提算法在选择最相关证据方面表现出色,从而显著提升了下游问答性能。
英文摘要
Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech and text, research on handling such multi-modal database scenarios remains limited. In this paper, we propose STeReO (Speech and Text Reranking Orchestrator), a reranker based on speech and text retrievers that aggregates disparate modality databases. To address the lack of specialized training data, we first curate a dataset comprising queries, mixed-modality evidence, and their corresponding relevance ranks. We then train the reranker and evaluate its effectiveness in both single-modality and mixed-modality scenarios. Our results demonstrate that the proposed algorithm excels at selecting the most relevant evidence, thereby significantly improving downstream question-answering performance.
CommentsAccepted to Interspeech 2026