AI 中文总结
本研究利用思维链动态特性,在不假设语言化CoT语义忠实性的情况下,检测大语言模型在布尔可满足性任务中的分布式推理失败,并通过针对性提示干预提升了Llama3-70B的准确率。
AI 中文摘要
思维链(Chain-of-thought, CoT)推理可提升大语言模型(Large Language Model, LLM)的性能,同时也为模型的推理过程提供了可观测的接口。然而,现有利用语言化CoT监控推理正确性的方法,大多仅评估单个中间步骤的语义正确性或一致性,而非推理过程如何沿轨迹演变。因此,分布在整个推理轨迹中的失败,而非局限于单个错误步骤的失败,仍未得到充分探索。此外,语言化CoT未必忠实地反映模型的内部推理,这促使我们开展不将单个陈述视为内部计算字面描述的分析。本研究因此探究,是否可利用可见CoT的动态特性,在不假设此类语义忠实性的情况下,系统区分成功与失败的推理。我们在可验证的布尔可满足性任务上研究了一系列LLM,任务具有可变复杂度,可在每个模型的能力边界附近进行可控比较。按推理功能对CoT语句进行标记后,我们发现SAT问题存在过早验证崩溃现象:错误轨迹更早进入子句检查,重复类似操作,且更早完成。在UNSAT问题上,模型会贸然走向错误的SAT结论,检查候选赋值而非在构造的案例间推导矛盾。随后,针对性的证明搜索提示干预将Llama3-70B的准确率从13.3%提升至85%,纠正了84.6%的此类错误。这些结果表明,能力失败可表现为可见推理结构中与任务相关的分布式变化,且不依赖于语言化轨迹是否反映模型内部计算的CoT动态,可帮助诊断并纠正失败。
英文摘要
Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning correctness, however, largely evaluate the semantic correctness or consistency of individual intermediate steps, rather than how the reasoning process evolves across the trace. As a result, failures distributed across the reasoning trajectory, rather than those localized to a single incorrect step, remain comparatively underexplored. Furthermore, verbalized CoTs need not faithfully reflect the model's internal reasoning, motivating analyses that do not treat individual statements as literal accounts of internal computation. In this work, we therefore ask whether the dynamics of visible CoT can be leveraged to systematically distinguish successful from failed reasoning without assuming such semantic faithfulness. We study a range of LLMs on verifiable Boolean satisfiability tasks with variable complexity, enabling controlled comparisons near each model's capability frontier. Tagging CoT sentences by reasoning function reveals premature verification collapse on SAT problems: incorrect traces enter clause checking earlier, repeat similar operations, and finalize sooner. On UNSAT problems, models presumptuously move towards incorrect SAT conclusions, checking candidate assignments rather than deriving contradictions across constructed cases. Subsequently, a targeted proof-search prompt intervention raises Llama3-70B accuracy from 13.3% to 85%, correcting 84.6% of these errors. These results show that capability failures can manifest as distributed, task-dependent changes in the structure of visible reasoning, and that CoT dynamics agnostic to whether the verbalized trace reflects the model's internal computations can help diagnose and correct failures.