发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对任务导向LLM智能体评估中LLM评判者的可靠性问题,提出GAUGE协议,发现满意度与任务成功脱节,且对强智能体区分度低,并建议采用校准后信任的节奏。
AI 中文摘要
比较和选择任务导向的LLM智能体越来越依赖于一种低成本的离线评估门控:基于人物角色的LLM用户模拟器与每个候选智能体进行对话,LLM作为评判者对对话记录进行评分,得分较高的智能体被提升。我们引入了GAUGE,一种可重用的离线协议,用于衡量该门控的排名是否与基于事实的可验证奖励相匹配,该协议在来自六个提供商的25个智能体上,基于$\ au^2$-bench和SimulatorArena基准进行测试,并区分了发布实践中混淆的两种评估有效性:排名有效性和构念有效性。首先,存在满意度-成功差距:满意度几乎不携带关于任务成功的信息,因为我们盲审小组评定为满意的对话与实际成功不相关,其中57.5%的对话未能完成客户的任务,这一模式在五个评分者群体、两个基准以及我们评定的每个主观维度上均一致。其次,虽然门控的排名在广泛的能力范围内是稳健的,但在近乎相等的强智能体之间失去了分辨率:决策分歧率从宽奖励对上的<1%跃升至接近对上的31%。因此,该门控经过人工验证但锚定错误。作为补救措施,我们提出了一种先校准后信任的节奏,其中无评判者的完成位是截断回归的零成本绊线。
英文摘要
Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the $τ^2$-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from $<$1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.
CommentsAccepted to EMNLP 2026 (Industry Track)