arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时该放弃:诊断与训练大语言模型以中止无效推理

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

Xinyan Guan, Jiali Zeng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, Fandong Meng

arXiv 2607.29211首次发表:更新:

发表机构

Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Weixin AI, Tencent Inc(中国科学院软件研究所; 中国科学院大学; 腾讯微信人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型在超能力任务上产生无效推理的问题,提出CaRL方法对齐模型行为与能力边界,实验验证其可减少无效推理且不牺牲任务性能。

AI 中文摘要

大语言模型在超出自身能力的任务上会产生计算成本高但语义空洞的推理,存在看似合理实则错误的推导误导用户的风险。我们通过系统分析刻画这种“无效推理”现象,揭示其普遍的能力越界,以及能力与行为间的系统性校准偏差。主要失败模式是似是而非的推理,输出表面有效但含细微错误,且随任务难度升级。为解决该问题,我们提出CaRL(能力对齐强化学习),通过奖励塑形(激励拒绝而非无效推理)和事后拒绝增强(将失败转化为拒绝监督),使模型行为与能力边界对齐。实验表明,该方法大幅减少无效推理,同时在不同难度任务上保持性能,在不牺牲效用的情况下实现了能力对齐的行为。

英文摘要

Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this \textit{futile reasoning} phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce \textbf{CaRL} (\textbf{Ca}pability-\textbf{a}ligned \textbf{R}einforcement \textbf{L}earning), which aligns model behavior with capability boundaries through reward shaping that incentivizes refusal over futile reasoning and hindsight refusal augmentation that converts failures into refusal supervision. Experiments demonstrate a substantial reduction in futile reasoning while preserving performance across task difficulties, effectively achieving capability-aligned behavior without sacrificing utility. \footnote{https://github.com/icip-cas/Knowing-When-to-Quit}

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑