发表机构
School of Computing, National University of Singapore; Xiaohongshu Inc.; School of Cyber Science and Technology, Zhejiang University; School of Intelligence Science and Technology, Peking University; School of Statistics and Data Science, Shanghai University of Finance and Economics; Institute for Artificial Intelligence, Peking University(新加坡国立大学计算学院; 小红书公司; 浙江大学网络空间安全学院; 北京大学智能科学与技术学院; 上海财经大学统计与数据科学学院; 北京大学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对构建可靠特定查询评分标准难的问题,提出“试验中的评分标准”框架,仅从查询演化评分标准集,靠合成响应对获取监督并验证,实验证明该框架有效,平均准确率最佳且在多数评估集领先。
AI 中文摘要
评分标准为训练和评估大语言模型提供结构化、细粒度的信号。然而,可靠的特定查询评分标准很难构建。现有方法通常从人工编写的评分标准、偏好数据或采样响应中获取监督。直接的查询到评分标准生成避免了这些资源,但没有明确检查合理的评分标准是否有用。这样的评分标准可能无法区分答案质量、奖励可选风格或惩罚有效的替代策略。我们引入了“试验中的评分标准”,这是一个仅基于查询的框架,它从空集演化出一组评分标准,无需外部注释或模型训练。它仅从合成的评分标准条件响应对中获取监督,并在添加每个提议的评分标准之前进行验证,筛选出无区分性、过于特定和仅风格的候选评分标准。在五个偏好基准套件上的实验证明了“试验中的评分标准”的有效性,它实现了最佳平均准确率,并在七个评估集中的六个上领先。
英文摘要
Rubric evolution offers a promising approach to improving the quality of rubrics generated by large language models (LLMs). Central to this process is rubric comparison, which identifies the better of two rubrics and guides the direction of evolution. However, accurate rubric comparison is difficult, which presents two challenges. (1) It should reflect downstream task performance, which is essential for assessing rubric utility but often prohibitively expensive to evaluate. (2) It should discourage unnecessary criteria, which increase verification costs and may dilute the influence of essential criteria. To address these challenges, we introduce Rubrics on Trial, a multi-agent framework that evolves rubrics by comparing synthetic response pairs. To address challenge 1, the framework compares synthetic responses that satisfy the respective rubrics, providing a proxy for downstream performance without training a separate policy for each rubric. To address challenge 2, it assesses the necessity of a candidate criterion by independently generating high-quality alternative responses that violate it and comparing them with edited versions that satisfy it. A rubric is favored when it improves response quality in both comparisons, and the resulting comparison signal is further incorporated for rubric evolution. Extensive experiments demonstrate that Rubrics on Trial improves the quality of generated rubrics and leads to better downstream task performance.