arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型自适应元推理的认知需求引导

Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

John Scoville, Shengzhuang Chen, Yejin Bang, Stefan Winzeck, Jonathan Richard Schwarz

arXiv 2608.01319首次发表:更新:

发表机构

Thomson Reuters Foundational Research; Imperial College London(汤森路透基础研究部; 帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无需训练的元推理框架CDS,通过认知量表刻画需求,在三类大模型和六个基准上较直接调用、标准CoT分别提升21.9%、9%准确率,难数学编码任务增益显著。

AI 中文摘要

近期的元推理框架通过将思维链生成包裹在迭代控制循环中,改进了大语言模型(LLM)的推理能力,支持更有效的回溯、推理循环终止以及注入有前景的推理模式等策略调整。尽管结果可观,但现有方法常依赖后向的奖励函数、采用粗糙的搜索动作,或需要大量少样本监督来训练额外的推理控制器。我们提出认知需求引导(Cognitive Demand Steering, CDS),这是一种无需训练的元推理框架,配备残差需求评估:在每一步,基于LLM的进度评估器会刻画得出解决方案所需的残差推理量,而非仅评估前一步。这使得元控制器能够选择包含通用示例和动作(如定量推理的通用指导)的推理干预措施,直接应对这一前瞻性需求信号。这种转变消除了对任何训练组件的需求,同时实现了跨模型和任务的零样本迁移,无需适配。我们采用认知量表,基于认知科学设计干预措施,并在16个维度上刻画初始问题复杂度和残差需求信号(如注意与扫描、学习与抽象、空间物理推理),为控制器提供细粒度诊断词汇。在三个前沿LLM和六个推理基准上取平均,CDS相比直接调用准确率提升21.9%,相比标准思维链(CoT)推理提升9%,在困难的数学和编码任务上提升最大。

英文摘要

Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of reasoning loops, and injection of promising reasoning patterns, among other strategy adjustments. Despite promising results, methods often rely on backward-looking reward functions, utilize coarse search actions, or require additional reasoning controller training requiring many-shot supervision. We introduce Cognitive Demand Steering (CDS), a training-free meta-reasoning framework equipped with residual demand assessment: at each step, an LLM-based progress evaluator characterizes the residual reasoning required to arrive at a solution rather than merely evaluating the previous step. This allows a meta-controller to select reasoning interventions comprising both general-purpose exemplars and actions (e.g., general guidance for quantitative reasoning) that directly tackle this forward-looking demand signal. This shift eliminates the need for any trained component while enabling zero-shot transfer across models and tasks with no adaptation. Rather than relying on coarse characterizations, we employ cognitive scales to both design interventions as well as profile initial problem complexity and residual demand signal over 16 dimensions motivated by cognitive science (e.g., attention and scan, learning and abstraction, spatio-physical reasoning), giving the controller a fine-grained vocabulary for diagnosing. Averaged across three frontier LLMs and six reasoning benchmarks, CDS improves accuracy by $21.9\%$ over direct calls and $9\%$ over standard CoT reasoning, with the largest gains on difficult mathematics and coding tasks.

Comments21 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑