发表机构
Peking University; Xiaomi(北京大学; 小米)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对长文本幻觉强化学习的「拒答-丰富度」权衡,提出用关键要点评分规则定义奖励,发现软组合依据、评分规则覆盖率与相关性的奖励能实现最佳平衡,提升了依据性与分布外迁移能力。
AI 中文摘要
惩罚无依据主张的奖励可提升长文本生成的依据性,但也会让模型减少回答。我们在长文本幻觉强化学习中研究这种「拒答-丰富度」的权衡。我们未采用长度、主张数量、细节或成对相关性等全局丰富度代理指标,而是用关键要点评分规则来表示每个问题,该规则指定了有用答案应涵盖的必需与可选信息。这些评分规则直接定义覆盖率,同时用于评估和作为奖励信号。在仅依据、基于代理、仅评分规则及组合奖励的对比实验中,我们发现存在稳定的权衡关系:严格的依据奖励可提升主张依据性,但会抑制覆盖率;无约束的评分规则奖励可提升覆盖率,但会削弱依据性。在实验中,依据、评分规则覆盖率与相关性的软组合实现了最佳平衡,其在分布内的依据性得到提升,且相比仅依据或仅评分规则奖励,能更好地迁移到分布外的清单任务中。
英文摘要
Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.