arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02926cs.CL

多语言提示污染下的输出语言混淆

Output Language Confusion under Multilingual Prompt Contamination

Riju Marwah, Ritvik Garimella, Khusham Bansal, Atishay Jain, Amit Sheth

首次发表
浏览论文内容

中文总结 AI 辅助

提出多语言干扰项干扰评估协议,发现混合语言提示下模型存在脚本切换与语义错误混淆,需分解幻觉率以准确评估可靠性。

中文摘要 AI 辅助

标准事实性基准假设提示为干净的单语言且采用精确匹配评分,这两个假设在现实世界的多语言部署中同时失效,从返回混合语言段落的检索增强生成管道到用户粘贴多语言网页内容均如此。我们引入了多语言干扰项干扰(MDI),这是一种轻量级且完全可复现的评估协议,无需新数据或标注,其中事实性问题前会加上一句语义无关的外语句子,并在八种干扰条件下对五个指令微调的大语言模型在TruthfulQA和TriviaQA上进行评估(共40,000次评估)。我们的核心发现是一个度量混淆:对于Llama-3.1-8B在印地语干扰下,58%的响应切换为天城文,产生0.710的原始幻觉代理指标,但人工审查显示,148个切换脚本的响应中有120个在干净条件下正确的响应尽管以错误脚本书写,语义上仍然正确,将调整后的语义幻觉率降至0.470。所有其他模型通过弃权(不执行)升级来响应,幻觉率没有增加。一段段落长度的英语干扰在所有模型上触发近乎普遍的弃权(0.806-0.998),这与阅读理解混淆一致,这一失败模式对多语言RAG管道有直接影响。TruthfulQA多项选择准确率在所有单句条件下均不受影响。这些结果表明,在混合语言设置中,精确匹配幻觉率应在得出关于模型可靠性的结论之前分解为脚本切换和语义错误两个组成部分。

英文摘要

Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-language passages to users pasting multilingual web content. We introduce Multilingual Distractor Interference (MDI), a lightweight and fully replicable evaluation protocol requiring no new data or annotation, in which factual questions are preceded by a semantically irrelevant foreign-language sentence, and evaluate five instruction-tuned LLMs across TruthfulQA and TriviaQA under eight distractor conditions (40,000 evaluations). Our central finding is a metric confound: for Llama-3.1-8B under a Hindi distractor, 58% of responses switch to Devanagari script, yielding a raw hallucination proxy of 0.710, but manual review reveals that 120 of 148 script-switched responses that were correct under clean conditions remain semantically correct despite being written in the wrong script, reducing the adjusted semantic hallucination rate to 0.470. All other models respond through abstention escalation with no hallucination increase. A paragraph-length English distractor triggers near-universal abstention (0.806-0.998) across all models, consistent with reading-comprehension confusion, a failure mode with direct consequences for multilingual RAG pipelines. TruthfulQA multiple-choice accuracy is unaffected under all single-sentence conditions. These results show that exact-match hallucination rates in mixed-language settings should be decomposed into script-switching and semantic error components before drawing conclusions about model reliability.

发表机构

  • University of Tübingen(蒂宾根大学)
  • Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所)
  • Thapar Institute of Engineering & Technology(塔帕尔工程技术学院)
  • Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)
  • Indian AI Research Organization(印度人工智能研究组织)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑