arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从拒答到丰富:用于长文本幻觉强化学习的评分规则奖励

From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning

Yudong Wang, Zhe Yang, Wenhan Ma, Rang Li, Qibin Yang, Weimin Xiong, Jiangshan Duo, Liang Zhao, Zhifang Sui

arXiv 2608.12337首次发表:更新:

发表机构

Peking University; Xiaomi(北京大学; 小米)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长文本幻觉强化学习的「拒答-丰富度」权衡,提出用关键要点评分规则定义奖励,发现软组合依据、评分规则覆盖率与相关性的奖励能实现最佳平衡,提升了依据性与分布外迁移能力。

AI 中文摘要

惩罚无依据主张的奖励可提升长文本生成的依据性,但也会让模型减少回答。我们在长文本幻觉强化学习中研究这种「拒答-丰富度」的权衡。我们未采用长度、主张数量、细节或成对相关性等全局丰富度代理指标,而是用关键要点评分规则来表示每个问题,该规则指定了有用答案应涵盖的必需与可选信息。这些评分规则直接定义覆盖率,同时用于评估和作为奖励信号。在仅依据、基于代理、仅评分规则及组合奖励的对比实验中,我们发现存在稳定的权衡关系:严格的依据奖励可提升主张依据性,但会抑制覆盖率;无约束的评分规则奖励可提升覆盖率,但会削弱依据性。在实验中,依据、评分规则覆盖率与相关性的软组合实现了最佳平衡,其在分布内的依据性得到提升,且相比仅依据或仅评分规则奖励,能更好地迁移到分布外的清单任务中。

英文摘要

Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑