发表机构
University of Tübingen; Aleph Alpha Research; Lab1141(图宾根大学; Aleph Alpha 研究院; Lab1141)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对推理模型在欠指定任务上过度推理而不弃权的问题,提出基于资源理性的GRPO奖励,微调4B模型实现弃权率提升12.8%且推理成本降低44%。
AI 中文摘要
尽管现代大型推理模型(LRMs)在许多任务中擅长提供正确答案,我们为它们常常在一种关键能力上挣扎的观察提供了额外证据:知道何时应弃权(不执行)而不作答。我们通过将LRM行为与人类研究结果进行比较来分析这一差距,揭示出人类在不可回答任务上的推理努力上限由可回答任务决定,而LRM则在不可回答提示上生成更长的思维链(CoTs),浪费计算资源。为克服这一低效,我们从人类认知的资源理性视角获得灵感,引入一种新颖的GRPO奖励,鼓励对任务是否包含解决所需全部信息进行高效推理。使用该奖励微调多个4B规模的LRM,实现了类似人类的弃权(不执行)性能提升(平均+12.8%),同时保留作答能力并提升模型效率(平均CoT缩短44%)。
英文摘要
While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).
Comments19 pages, 9 figures