面向可靠且实用的评估流程
Towards a Reliable and Practical Eval Pipeline
浏览论文内容
中文总结 AI 辅助
针对现有LLM软件评估仅处理局部可靠性问题的不足,提出结合清单创建与学习聚合的端到端评估流程,提升评判一致性与准确性,还提供自一致性等功能并实证验证其有效性。
中文摘要 AI 辅助
基于大语言模型(LLM)的软件系统在开发生命周期中日益需要有效的“评估”作为质量关卡。然而,现有工作通常仅处理评估可靠性的个别方面,而非全部实用需求。我们提出一种端到端评估流程,该流程结合评估清单创建与针对清单响应的学习聚合方法,以提升LLM评判者间的一致性及与人工评判的准确性。此框架还提供自一致性、解释及预测不确定性,我们通过实证研究证明了其有效性。
英文摘要
LLM-based software systems increasingly require effective "evals" as quality gates in the development lifecycle. However, existing work typically addresses individual aspects of eval reliability rather than the full set of practical requirements. We present an end-to-end eval pipeline that combines eval checklist creation, with learned aggregation for checklist responses, to improve agreement across LLM judges and accuracy against human judgments. The framework additionally pro- vides self-consistency, explanations, and prediction uncertainty, and we empirically demonstrate its effectiveness.
发表机构
- Salesforce(Salesforce(赛富时))
机构由 AI 辅助整理,请以论文原文为准。