AI 中文总结
该研究发现多语言推理差距受输出token预算机制影响,通过固定Qwen模型的预算峰值等实验,提出应将输出上限作为独立变量,报告不同预算机制下的准确率。
AI 中文摘要
多语言评估在单一输出token上限下报告准确率,但不同语言表达相同内容所需的token数量不同,因此该上限是一个隐藏的实验变量。我们针对Qwen3-8B和Llama-3.1-8B-Instruct在四种提示策略下,测试MGSM(德语、泰语、斯瓦希里语)上的母语与翻译差距是否为token预算的人为产物。测得的差距在不同预算间波动幅度达57个百分点,在上限约束下,长度归一化可使差距变动幅度达38.9个百分点,且在严格上限下,归一化可反转哪项策略得分更高的结果。我们预先固定了该搜索的三个Qwen峰值及其在1024处的接近零值,并在540000个独立硬截断解码上对其进行评估:第二个固定系列的六项经Holm校正的检验均拒绝原假设。$B^*=1024$处的固定检验仍无法拒绝原假设,因为此处母语准确率已达到饱和;在饱和值以上,剩余差距为策略性能差距,而非已识别的推理缺陷。相同的截断通道形成了按成本排序的适配阶梯:交叉拟合的泰语词汇扩展在固定预算下消除了0.0个百分点的差距,在19%的轨迹仍被截断的情况下消除了4.9个百分点的差距。第三个固定系列在固定强制上限下仅改变公布的预算;公布128而非2048个token使泰语母语准确率变动了5.1个百分点,因此准确率并非仅为强制上限的函数。从一次长上限运行计算得出的正确输出时间恒等式与三个预先指定的MGSM峰值匹配度达0.65个百分点,且在对另外三个基准的仅Qwen探索性分析中,与保留项的匹配度达0.92个百分点,在七个单元中的五个内准确定位了峰值。应将输出上限视为独立变量,并报告各预算机制下的准确率,而非仅在单一预算下报告。
英文摘要
Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the native-vs-translate gap on MGSM (German, Thai, Swahili) is a token-budget artifact for Qwen3-8B and Llama-3.1-8B-Instruct under four prompting strategies. The measured gap swings by up to 57 points across budgets, length normalization moves it by up to 38.9 points where the cap binds, and at tight caps normalization can reverse which strategy scores higher. We prospectively froze the sweep's three Qwen peaks and its near-zero value at 1024 and evaluated them on 540,000 independently hard-capped decodes: a second frozen family of six Holm-corrected tests rejects every null. The frozen test at $B^*=1024$ still fails to reject because native accuracy has already saturated there; above saturation, the residual difference is a strategy-performance gap, not an identified reasoning deficit. The same truncation channel prices a cost-ordered adaptation ladder: a cross-fitted Thai vocabulary extension closes 0.0 points of the gap at the frozen budget and 4.9 points where 19% of traces still truncate. A third frozen family varies only the announced budget at a fixed enforced cap; announcing 128 rather than 2048 tokens moves Thai native accuracy by 5.1 points, so accuracy is not a function of the enforced cap alone. A correct-emission timing identity computed from one long-cap run matches the three pre-specified MGSM peaks to 0.65 points and, in an exploratory Qwen-only analysis of three further benchmarks, tracks held-out items to 0.92 points, locating the peak exactly in five of seven cells. Treat the output cap as an independent variable and report accuracy across the budget regime, not at a single budget.
Comments15 pages, 2 figures, 11 tables. Under review