arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CEDAR-GRPO:面向大语言模型通用溯因推理的过程感知强化学习

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban

arXiv 2608.14791首次发表:更新:

发表机构

Sharif University of Technology; University of Tehran(谢里夫理工大学; 德黑兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出过程感知强化学习框架CEDAR-GRPO,通过后训练提升LLM的通用溯因推理能力,在11个未见任务上实现显著性能提升,验证了其迁移有效性。

AI 中文摘要

溯因推理常被定义为对最佳解释的推理,是不确定性下解释的核心,适用于日常认知、调查及科学发现等场景。但大语言模型(LLM)研究多通过狭窄的特定任务基准研究溯因,难以判断观测到的性能提升是否能迁移至训练或评估所用基准之外的场景。本文探究强化学习(RL)后训练能否提升溯因作为可迁移推理能力的表现,提出CEDAR-GRPO这一过程感知框架,结合最终答案正确性与证据覆盖、证据到解释方向性的溯因奖励。在受控、领域无关的溯因假设生成与假设选择任务混合数据集上,对四个开放权重LLM进行后训练,在涵盖假设选择、缺失事实生成、可废止推理、长上下文调查、临床推理、代码调试及非溯因对照的11个未见任务上评估,CEDAR-GRPO在所有未见任务上均提升了每个模型的性能,相较于基础模型和仅考虑正确性的GRPO,平均提升分别为7.4和2.7个百分点,最大提升达30.8个百分点。 ablation实验证实,RL、溯因奖励设计及任务多样性均对迁移有贡献,过程级指标进一步显示模型表现出更强的溯因行为,包括探索替代方案、排除竞争项、回溯及不确定性标记。

英文摘要

Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post-training can improve abduction as a transferable reasoning capability. We introduce CEDAR-GRPO, a process-aware framework that combines final-answer correctness with abductive rewards for evidence coverage and evidence-to-explanation directionality. Four open-weight LLMs are post-trained on a controlled, domain-neutral mixture of abductive hypothesis-generation and hypothesis-selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing-fact generation, defeasible inference, long-context investigation, clinical reasoning, code debugging, and non-abductive controls. CEDAR- GRPO improves every model on every held-out task over both base models and correctness-only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process-level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.

CommentsCode and data are available at https://github.com/cedar-grpo/cedar-grpo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑