AI 中文总结
该研究针对大语言模型的无思维行为,提出三项评估指标,发现模型存在“思维惯性”,显式推理在严格控制下仍持续,开放式任务中仅答案合规性与准确率存在权衡,需系统评估无思维能力。
AI 中文摘要
大语言模型(LLMs)越来越多地配备了明确的“思维模式”,但与之对应的“无思维”行为却受到了少得多的关注。我们从两个维度研究LLMs的无思维行为:a. 如何衡量无思维?现有工作通常通过代理指标定义无思维,例如禁用思维模式或缺乏长推理轨迹,但这些代理指标不可靠——禁用思维模式的模型仍可能输出推理内容,而长轨迹可能包含填充内容而非真正的推理。我们转而将每个响应归一化为预回答轨迹和最终答案,并在三个层面进行评估:(i)严格仅输出答案合规性的空思维率;(ii)指令感知的问题-预回答相关性,衡量问题与预回答轨迹的相似性;(iii)作为评判者的LLM显式推理率,衡量可见的显式推理。这些指标共同区分仅输出答案、相关但非推理文本以及显式推理。b. 无思维在不同任务和模型间如何变化?我们在布尔、多选和开放式问题上,对六个LLMs评估了六种提示干预措施。我们发现,显式无思维控制无法可靠消除显式推理,模型反而表现出“思维惯性”:即使在严格控制下,显式推理仍会持续存在,且随着答案空间的开放变得更普遍。布尔和多选任务的准确率保持稳定,而开放式任务则呈现仅答案合规性与任务准确率之间的权衡。在不同答案空间中重写相同问题的实验表明,提供候选答案会使仅输出答案的响应更易生成。这些发现确立了无思维是一项非平凡的能力:不能仅通过模型设置或指令假设模型能停止显式推理,需与推理能力一起进行系统评估。
英文摘要
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
CommentsAccepted by NeurIPS 2026. Website: https://thinking-inertia.github.io GitHub: https://github.com/thinking-inertia/code