arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究推理语言对齐在单语检索增强生成中的作用

Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation

Oliver Hauck, Mario Sanz-Guerrero, Katharina von der Wense

arXiv 2610.03136首次发表:更新:

发表机构

Johannes Gutenberg University Mainz; University of Colorado Boulder(约翰内斯·古腾堡大学美因茨; 科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建单语德语RAG问答测试平台,发现推理语言与查询及检索文档语言对齐可提升性能,但强制德语推理仅达到原生英语水平,表明需原生多语言推理。

AI 中文摘要

推理轨迹能提升大型语言模型(LLMs)的性能,但当前模型主要被训练用英语进行推理。已有研究表明,强制模型用另一种语言推理会降低准确率,即使推理语言与提示词的语言相匹配——但这仅限于模型在短提示词上进行推理的场景。在此,我们探讨在检索增强生成(RAG)中是否同样如此,该场景下模型必须阅读并整合大量目标语言的检索证据。为研究此问题,我们构建了一个完全单语的德语RAG问答测试平台,基于桌面角色扮演游戏《黑暗之眼》的虚构世界,该领域在德语中有丰富文献记载,但过于小众,模型无法凭记忆回答,因此必须依赖检索。在该测试平台上,我们改变智能体RAG系统的强制推理语言,发现将推理语言与查询和检索文档的语言对齐是有帮助的。强制德语推理优于强制法语推理,尽管模型在法语上的基准测试得分更高,因此优势来自对齐而非语言熟练度。当检索上下文更丰富且具有结构感知时,该优势会增大。然而,强制德语仅达到模型原生、不受约束的英语推理水平而未超越之,表明需要原生多语言推理。我们公开发布了该测试平台和问答基准。

英文摘要

Reasoning traces improve large language models (LLMs), but current models are trained to reason mostly in English. It has been shown that forcing a model to reason in another language degrades accuracy, even when the reasoning language matches the language of the prompt -- but only for a setting where the model reasons over a short prompt. Here, we ask whether the same holds for retrieval-augmented generation (RAG), where the model must read and integrate a large amount of retrieved evidence in the target language. To study this, we build a fully monolingual German RAG question-answering testbed over the fictional world of the tabletop role-playing game The Dark Eye, a domain that is richly documented in German but too niche for the model to answer from memory, so that it has to rely on retrieval. Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps. Forced German reasoning outperforms forced French, although the model benchmarks higher in French, so the benefit comes from alignment and not from language proficiency. The advantage grows when the retrieved context is richer and structure-aware. However, forced German only reaches the level of the model's native, unconstrained English reasoning without surpassing it, showing that native multilingual reasoning is needed. We publicly release the testbed and QA benchmark.

CommentsAccepted to the Workshop on Open Reasoning Across Cultures & Languages at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑