发表机构
Uniphore(Uniphore)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PADMÉ提出一种仅用小型语言模型、无需人工参与的偏好对齐数据合成方法,用于生成智能体评估器的元评估数据,将人类判断一致性从73%提升至85%。
AI 中文摘要
语言模型经常被用来评估其他语言模型。一个能够跨多个标准对智能体行为进行评分的LM评估器,只要其决策与人类判断一致,就是有价值的。我们将评估这种一致性的问题称为元评估。直接解决这个问题是困难的:收集人类数据成本高昂,绝对评分难以对齐,而使用LM元评估器则会递归地引发可信度问题。我们将元评估重新表述为一个偏好判断问题:不是比较人类和LM评估器对轨迹的评分,而是询问它们的隐含偏好是否一致。基于此,我们引入了PADMÉ,一种数据合成方法,能够为智能体设置生成可靠的基于标准的元评估数据。PADMÉ仅使用小型语言模型,在评估过程中不需要人类参与,并且在低计算预算下运行。我们构建了PADMÉ的原型,并在四个智能体领域和三个评估标准上合成了一个包含1,000个样本的数据集。对150个样本子集的人类验证表明,与朴素基线相比,PADMÉ将人类判断的一致性从73%提高到85%。使用我们的数据集对25个常见模型进行元评估,展示了评估性能与评分粒度、宽松度、模型大小等因素之间的相关性。
英文摘要
Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.
CommentsAccepted at the NeurIPS 2026 Workshop TAE (Trust-AI-Eval): Can We Trust AI Evaluation? 27 pages, 3 figures. Code and data at https://github.com/chc012/padme