发表机构
Zhejiang University; Fudan University; The Hong Kong Polytechnic University(浙江大学; 复旦大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
REIN是一个对齐框架,通过训练大型推理模型生成结构化推理序列,引入奖励机制鼓励模型在无正确推理链时弃权,在单次前向传播中减少幻觉、提升选择性准确率并保持高覆盖率。
AI 中文摘要
大型推理模型(LRMs)易产生幻觉,这会破坏其可靠性并对安全部署构成挑战。LRMs中的幻觉源于两种不同的失败来源:推理幻觉,即有缺陷的推理步骤会传播至错误结论;知识幻觉,即模型缺乏回答查询所需的事实知识。为解决推理幻觉,我们提出REIN,这是一个对齐框架,用于训练LRMs生成结构化推理序列:<think>→<reflection>→<answer>,使其在给出最终答案前能进行明确的自我反思。为解决知识幻觉,REIN引入了奖励机制,当采样的推理链均无法得出正确答案时,鼓励模型明确弃权(例如“我不知道”),使其能避免无依据的预测。在数学和常识推理基准上的大量评估显示,与竞争基线相比,REIN持续提升选择性准确率、减少错误但自我认可的响应,并保持高覆盖率。值得注意的是,REIN在单次前向传播中实现了这些增益,无需过程监督、推理时控制器、外部搜索或多轮批评。对多个主干模型的实验表明,REIN将幻觉代理相对基础模型降低了58%~72%,同时保持86%~91%的平均覆盖率,并在已尝试的问题上提升了6.6%~14.2%的选择性准确率。
英文摘要
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps propagate to an incorrect conclusion, and knowledge hallucination, where the model lacks the requisite factual knowledge to answer the query. To address reasoning hallucination, we propose REIN, an alignment framework that trains LRMs to produce a structured reasoning sequence, $\texttt{<think>} $$\rightarrow$ $\texttt{<reflection>} $$\rightarrow$ $\texttt{<answer>}$, enabling explicit self-reflection before committing to a final answer. To address knowledge hallucination, REIN introduces a reward mechanism that encourages explicit abstention (e.g., "I don't know") when none of the sampled reasoning chains yields a correct answer, allowing the model to refrain from unsupported predictions. Extensive evaluations on mathematical and commonsense reasoning benchmarks show that REIN consistently improves selective accuracy, reduces incorrect-but-self-endorsed responses, and maintains high coverage compared with competitive baselines. Notably, REIN achieves these gains within a single forward pass, without requiring process supervision, inference-time controllers, external search, or multi-round critiques. Experiments on multiple backbones show that REIN reduces the hallucination proxy by $58\sim72\%$ relative to the base models while maintaining $86\sim91\%$ average coverage, and improves selective accuracy on attempted questions by $6.6\sim14.2\%$.
Comments26 pages and 22 figures