已知事实导致的推理错误:针对大语言模型的步骤级自一致性组相对策略优化
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM
浏览论文内容
中文总结 AI 辅助
针对大语言模型推理中因上下文干扰产生事实错误的问题,提出步骤级自一致性组相对策略优化(SSC - GRPO),通过计算自一致性分数分配步骤级奖励,在数学推理基准和幻觉排行榜上取得领先,为检测和减轻幻觉提供新视角。
中文摘要 AI 辅助
随着大语言模型(LLMs)的迅速发展,现代系统不仅具备强大的基础能力和广泛知识,还能通过长的多步推理解决复杂问题。然而,随着推理轨迹变长,LLMs在推理过程中可能产生大量幻觉内容,难以检测。本文对LLM推理中出现的幻觉进行细粒度分析,发现推理轨迹特别容易出现上下文敏感事实幻觉,即模型实际有相关知识,但因推理过程中的上下文干扰而出现事实错误。为解决此问题,我们提出步骤级自一致性组相对策略优化(SSC - GRPO),通过计算多个展开中单个步骤的自一致性分数为推理轨迹分配步骤级奖励。与先前方法相比,SSC - GRPO在数学推理基准和幻觉排行榜上均取得了领先性能。我们的结果为检测和减轻大语言模型推理过程中的幻觉提供了新视角。
英文摘要
With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex problems via long, multi-step reasoning. However, as reasoning traces become longer, LLMs may produce a substantial amount of hallucinated content during the reasoning process, which is often difficult to detect. In this work, we conduct a fine-grained analysis of hallucinations arising in LLM reasoning and find that the reasoning traces are particularly prone to Context-Sensitive Factual Hallucinations: cases where the model actually has the relevant knowledge, yet makes factual errors due to contextual interference during reasoning. To address this issue, we propose Step-level Self-Consistency Group Relative Policy Optimization (SSC-GRPO), which assigns step-level rewards to reasoning traces by computing self-consistency scores of individual steps across multiple rollouts. Compared with prior methods, SSC-GRPO achieves state-of-the-art performance on both mathematical reasoning benchmarks and hallucination leaderboards. Our results offer a new perspective for detecting and mitigating hallucinations in the reasoning process of large language models.
发表机构
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。