arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理大语言模型中的隐藏语言一致性现象

Hidden Language Consistency Phenomena in Reasoning LLMs

Muhammad Ali Shafique, Kelly Marchisio

arXiv 2608.08447首次发表:更新:

发表机构

Kansas State University; Cohere(堪萨斯州立大学; Cohere公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以PolyMath基准测试集探究推理模型的语言一致性,发现难度提升会导致多语言模型输出语言一致性骤降,量化对一致性的影响独立于准确率,多语言能力需结合准确率、一致性与难度评估。

AI 中文摘要

多语言推理模型的评估通常仅关注其是否得出正确答案,而非在推理和响应过程中是否保持预期语言,这种遗漏掩盖了任务难度提升时显现的重要多语言行为。本文使用PolyMath基准测试集,针对八种语言和四个难度等级,研究推理模型的任务难度、任务准确率、思维-语言一致性(TC)及答案-语言一致性(AC),并得出四项发现:(1)语言一致性呈现四种依赖难度的行为:输出语言与输入保持一致、持续不一致、逐渐退化或突然崩溃;(2)识别出语言一致性崩溃效应,即难度提升会导致输出语言一致性骤降,尤其在代表性较弱及非拉丁文字语言中;(3)受该崩溃效应影响,模型转向内部主导语言时,更难任务的准确率可保持甚至提升;(4)量化对输出语言一致性的影响独立于其对准确率的影响,在容忍度投票(ε=1.0)下,GPTQ和AWQ通常表现优于AutoRound。这些结果表明,多语言能力不能仅用准确率表征,可靠评估需在多语言基准测试中同时考虑任务准确率、语言一致性和任务难度。

英文摘要

Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors that emerge as tasks become harder. In this paper, we study task difficulty, task accuracy, thinking-language consistency (TC), and answer-language consistency (AC) across reasoning models using PolyMath benchmark in eight languages and four difficulty levels. We uncover four findings: (1) language consistency exhibits four difficulty-dependent behaviors: output-language consistency remains aligned with input, remains misaligned, degrades gradually, or collapses abruptly. (2) We identify the language consistency breakdown effect, where increasing difficulty can cause a sudden drop in output-language consistency, especially in less strongly represented and non-Latin-script languages. (3) Due to this breakdown effect, accuracy can be preserved or even improved at a harder difficulty level as the model shifts to its internal dominant language. (4) Quantization can improve or degrade output-language consistency independently of its effect on accuracy, with GPTQ and AWQ often outperforming AutoRound under tolerance-based voting with ε = 1.0. These results show that multilingual capability cannot be characterized by accuracy alone; reliable evaluation should jointly consider task accuracy, language consistency, and task difficulty for multilingual benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑