arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04653cs.CL

推理时选择合适的语言模式以提升多语言可靠性

Choosing the Right Language Mode at Inference Time for Multilingual Reliability

Ekata Mitra, Ameeta Agrawal

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言大模型低资源语言推理差的问题,本文提出RAAI框架,通过ECE感知路由与提示融合、RI控制推理,提升低资源语言准确性并降低校准误差。

中文摘要 AI 辅助

多语言大语言模型往往在低资源至中等资源语言的推理中表现不佳。已有研究表明,翻译可通过帮助模型获取以英语为中心的更强表示来提升多语言推理能力,这引发了一个核心问题:多语言大语言模型实现可靠推理需要多少翻译量?何时增加翻译反而会引发干扰和过度自信?我们使用LLaMA和Qwen模型,开展了大量改变文本范围和语言模式(仅目标语言、仅英语、双语)的实验,以同时评估准确性和可靠性。结果揭示了明确的权衡关系:英语上下文通常能提升理解并纠正非英语理解导致的错误,但添加冗余的双语上下文会加剧干扰。我们提出了可靠性感知自适应推理(RAAI),这是一种无需训练的测试时框架,包含两个核心部分:(i)感知预期校准误差(ECE)的路由与提示融合,(ii)使用中间层风险指数(RI)控制顺序推理,仅在可能有帮助时分配计算资源并抑制有害的双语冗余。在两个模型系列上,RAAI使低资源语言的准确性提升了25%-37.7%,校准误差降低了约3%-6%,在最低资源语言层级中收益最为显著。

英文摘要

Multilingual large language models often struggle to reason in low- to mid-resource languages. Prior work has shown that translation can improve multilingual reasoning by helping models access stronger English-centric representations. This raises a central question: How much translation is needed for multilingual large language models to reason reliably, and when does more translation instead trigger interference and overconfidence? Using LLaMA and Qwen models, we run extensive experiments varying text scope and language mode (target-only, English-only, bilingual) to evaluate both accuracy and reliability. Our results reveal a clear trade-off: English context often improve understanding and recover errors caused by non-English comprehension, yet adding redundant bilingual context intensifies interference. We address this trade-off with Reliability-Aware Adaptive Inference (RAAI), a training-free test-time framework that (i) performs Expected Calibration Error (ECE)-aware routing and prompt fusion, and (ii) uses a mid-layer Risk Index (RI) to gate sequential reasoning, allocating compute only when it is likely to help and suppressing harmful bilingual redundancy. Across two model families, RAAI enhances accuracy by 25-37.7% on low-resource languages and lowers calibration error by approximately 3-6%, with the most pronounced benefits in the lowest-resource language tiers.

发表机构

  • Portland State University(波特兰州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑