arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为多轮业务智能体设计可靠的LLM即评判者测量系统

Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents

Kaiwen Luo, Ming Gao

arXiv 2609.33955首次发表:更新:

发表机构

Meta(Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多轮业务智能体评估,提出集成测量系统方法论,涵盖规范、模块化评判者、模拟器与人工治理,生产验证提升保真度与可维护性。

AI 中文摘要

许多LLM即评判者的评估在固定任务定义下对固定输出进行评分。生产环境中的多轮业务智能体反而需要一个持续维护的测量系统:正确性取决于业务特定的事实和流程,结果在多个轮次中逐步显现,并且失败必须在可操作之前归因于智能体能力或缺失的业务知识。我们提出了一种集成方法论,涵盖评估规范、模块化LLM评判者、意图保持的用户模拟以及人在环路的治理。该规范定义了对话级别的最终状态和可操作的失败归属。原子评判者共享带版本的证据,并输入一个显式的聚合图。模拟器仅在任务保持性和稳定性检查通过后才发布。独立的人工审计估计测量保真度,更新分层参考集,并将分歧路由到标签修正、指南修订或评判者改进。生产研究表明,系统级保真度在多次审计中持续提升,人类评审员和自动化评判者在共享反馈循环下共同改进,并且他们的联合工作流在两种报告的任务完成设置中具有最强的描述性表现。由于研究是观察性的,且人类参考本身需要修订,这些发现证明了操作上的实用性而非因果性或普遍优越性。贡献在于提供一个实用框架,使多轮智能体测量在被评估系统及其证据演化时保持可靠、可操作和可维护。

英文摘要

Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑