从偏好到原则:基于评分规则的接地知识答案对齐
From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
浏览论文内容
中文总结 AI 辅助
本研究针对开放域问答的奖励信号设计难题,提出基于检索证据生成查询特定多维评分规则的奖励框架,经实验验证其比基线方法有显著性能提升,为复杂开放域问答提供更有效的奖励监督。
中文摘要 AI 辅助
为开放域问答设计有效的奖励信号颇具挑战,因为高质量响应需同时满足多方面答案质量要求,而这些要求难以用整体标量目标捕获。我们提出一种基于评分规则的奖励框架,该框架基于检索到的证据生成查询特定评分规则,分解为多个质量维度,在后期训练中提供细粒度监督。在三个评估维度(组合性、接地性和指令遵循性)上取平均,我们的方法比指令微调基线提升6.5%,比平面评分规则变体提升4%,在所有评估数据集上均取得一致增益。将评分规则与检索证据关联可改善事实支持,将评分规则分解为特定质量维度可进一步提升连贯性、条理性及对查询要求的遵循。我们的结果表明,接地的多维评分规则可为复杂开放域问答提供更有效的奖励监督。
英文摘要
Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4%, with consistent gains across all evaluation datasets. Conditioning rubrics on retrieved evidence improves factual support, while decomposing rubrics into quality-specific dimensions further improves coherence, organization, and adherence to query requirements. Our results show that grounded, multi-dimensional rubrics provide more effective reward supervision for complex open-domain question answering.
发表机构
- Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。