AI 中文总结
本研究通过证据整合任务揭示LLMs在计数时依赖思考而非直接回答,思考能构建均匀加权累积计数并提升表现,但推理成本随字母数增长,且ICRL无法将计算自动化。
AI 中文摘要
有限的计算资源迫使自动化的系统1过程与昂贵的系统2思考之间进行权衡。大型语言模型(LLMs)可以在难题上花费额外的计算,但即使是计数这种人类和动物自动执行的基本操作,直接回答也难以应对。我们探究为何LLMs在计数时需要思考。证据整合长期以来被用于心理学和神经科学中探究决策过程。我们的证据整合任务在每一轮对话中呈现一个字母,并询问两个目标字母中哪一个出现得更频繁。通过平等地加权每个字母,运行中的计数差值可以最优地解决该任务;每一轮的标记可以表示并更新这一差值。直接回答则不均匀地加权证据,表现出强烈的近因效应,并且随着难度增加,分配给正确答案的概率降低。思考提升了表现并使整合权重接近均匀,然而在两种模式下,最终查询的注意力仍集中在序列末端。推理轨迹显示模型会重新审视输入、重新计数字母并检查影响答案的中间计数,这表明思考构建了直接回答所缺乏的累积计数,而非读出已经形成的计数。推理标记的成本随字母数量增长的程度远大于随连贯性增长的程度。结果反馈并未将这种计算引入直接回答:在上下文强化学习(ICRL)下,随着重复游戏,表现恶化且近因效应增强,但模型却变得更加自信。人类和动物将此类计算摊销为自动过程,而当前的LLMs仍须在每次试验中通过思考来支付这些成本。哪些操作可以通过学习直接可用,仍是未来模型如何分配计算的核心问题。
英文摘要
Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this requires thinking in LLMs. Evidence integration has long been used in psychology and neuroscience to probe decision-making. Our evidence-integration task presents one letter per conversational turn and asks which of two target letters appeared more often. A running count difference solves the task optimally by weighting every letter equally; tokens at each turn could represent and update this difference. Direct responses instead weighted evidence unevenly, with strong recency effects, and assigned less probability to the correct answer as difficulty increased. Thinking improved performance and made integration weights nearly uniform, yet final-query attention remained concentrated on the sequence ends in both modes. Reasoning trajectories showed models revisiting input, recounting letters, and checking intermediate counts that informed the answer, suggesting that thinking constructs the accumulated count that direct responses lack rather than reading out one already formed. Reasoning-token costs grew with the number of letters far more than with coherence. Outcome feedback did not bring this computation into direct responses: under in-context reinforcement learning (ICRL), performance deteriorated over repeated games and recency effects strengthened, yet models grew more confident. Humans and animals amortize such computations into automatic processes, whereas current LLMs still pay for them with thinking on every trial. Which operations learning can make directly available remains central to how future models allocate computation.