arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26845cs.LG

不确定性下的思考:语言模型中的证据使用与信息寻求

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过双臂老虎机试验探究大型语言模型在不确定性下的思考机制,发现思考可增强其基于当前证据的行动,但未产生更具信息寻求性的探索策略。

中文摘要 AI 辅助

推理时思考提升了大型语言模型的性能,但总体结果并未揭示模型是否更有效地利用可用证据,或寻求可改善未来决策的信息。我们通过在匹配的不确定性下测量行动偏好、思考长度和报告的置信度来区分这些反应。10个开放权重模型在思考与非思考模式下完成了匹配的水平式双臂老虎机试验。一个认知模型将价值引导的行动和与不确定性无关的选择噪声,与两种探索的行为特征区分开:类似UCB的对较不熟悉臂的偏好,以及随总不确定性增加的类似Thompson的选择变异性。平均而言,思考增强了价值引导的行动并减少了与不确定性无关的选择噪声,但未产生类似UCB的探索或增强类似Thompson的探索。在行动之外,信息不平衡的历史条件(其观测值也多于匹配的平衡条件)与更长的思考长度相关。报告的置信度对决策难度更敏感,且与所选任务证据的关联更紧密。我们将这些思考长度和报告的置信度模式分别解释为与元认知控制和元认知监测一致,但未证实这两种过程。解码器扫描(尤其是温度)改变了选择噪声和思考长度,但未重现联合跨输出模式。在该受控决策环境中,思考改善了模型基于当前证据的行动方式,而测量到的特征均不支持向更具信息寻求性的策略转变。

英文摘要

Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.

↑