列表计数失败并非单一现象
List Counting Failures Are Not One Phenomenon
- The George Washington University(乔治华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示不同大语言模型在列表计数任务上呈现不同错误模式,且错误模式与输入瓶颈无关,并警示幅度修正方法不可直接跨模型迁移。
AI中文摘要:
对括号列表中的项目进行计数看似简单,然而开放权重聊天模型却常常出错。先前的研究通常将其归咎于输入瓶颈,例如子词碎片化或注意力稀释,这些解释预测不同模型应以大致相同的方式失败。然而,在相同提示下对七个指令模型进行测试时,错误答案呈现出不同的模式:Qwen 和 Gemma 27B 常将奇数长度翻转为接近的偶数,OLMo 的错误集中在几个中等大小的整数上,而 Llama 则倾向于少计数。这些模式是有用的标签,而非稳定的族规律(Gemma 9B 并未重现 Gemma 27B 的奇数到偶数下降),且在我们的基准测试中,更严重的子词碎片化并未使计数变得更难。当模型回答错误时,线性探针通常仍能从残差流中恢复真实计数。匹配相同的奇数到偶数错误也并不意味着相同的后期 MLP 幅度修正:在相同协议下,缩放后期 MLP 输出对 Qwen 有适度帮助,但对 Gemma 27B 几乎无效,而残差引导只能通过以奇数增益换取偶数损失来移动两者。这些结果警示,未经迁移检查,不应将幅度修正跨模型迁移。
英文摘要:
Counting the items in a bracketed list looks trivial, yet open-weight chat models often get it wrong. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different models should fail in roughly the same way. Across seven instruct models on identical prompts, however, wrong answers form distinct modes: Qwen and Gemma 27B often flip odd lengths to a nearby even integer, OLMo concentrates errors on a few mid-sized integers, and Llama tends to under-count. These modes are useful labels rather than a stable family law (Gemma 9B does not reproduce Gemma 27B's odd-to-even drop), and heavier subword fragmentation does not make counting harder on our benchmark. When the model answers incorrectly, a linear probe can usually still recover the true count from the residual stream. Matching the same odd-to-even error also does not imply the same late-MLP magnitude fix: scaling a late MLP output helps Qwen modestly but is near null on Gemma 27B under the same protocol, while residual steering can move both only by trading odd gains for even losses. These results caution against transferring that magnitude fix across models without a transfer check.