arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:面向任务型对话智能体的基于状态、感知弃权的评估方法

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

Rayan Khoury, Shih-Yao Lin, Pratyush Mishra

arXiv 2609.00434首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGE是一种基于状态、感知弃权的任务型对话智能体评估方法,其SAGE-Core零LLM成本且在四个数据集切片上表现不劣于各类LLM评判器,经人工审计标签保真度高。

AI 中文摘要

评估面向任务的对话智能体,不仅需要判断回复是否通顺,还需判断每一轮对话是否正确推进了底层工作流状态——这一区别是传统的整体式大语言模型(LLM)评估器可能忽略的,因为它们将可用上下文作为单一单元进行评估,且每轮对话需要一次或多次完整模型调用。本文提出SAGE(State-Grounded Abstention-Aware Evaluation,基于状态、感知弃权的评估方法),该方法将工作流规范和每轮状态差异编译为原子的、基于模式的标准,并通过级联的符号验证器和编码器/自然语言推理(NLI)验证器处理每个标准,这些验证器会弃权(不执行)而非猜测,将标准裁决汇总为带有证据追踪的轮级决策。其推荐的操作点SAGE-Core仅通过编译器、符号规则和设备上的编码器即可判定81%至91%的标准,且零付费LLM成本;而SAGE-LLM则为开放类标准添加了可选的聚焦LLM回退机制。在涵盖MultiWOZ、Schema-Guided Dialogue和ABCD的四个数据集切片上,所有被评估的LLM作为评判基准——包括感知状态的GPT-4.1评判器和更便宜的GPT-4.1-mini变体——在任何切片上的表现都未显著优于SAGE-Core,尽管GPT-4.1的G-Eval评判器每1000轮成本为4.7至8.0美元,而SAGE-Core的成本仅为0美元。一项包含200个样本的双标注员人工审计(κ=0.94)证实,在转录可见的失败类别上标签保真度很高,其中排除弱显著性IUV类别后,SAGE-Core与最强的LLM评判器统计上相当,且能诚实地将被忽略的用户价值界定为具有弱广泛人类显著性的状态一致性信号。我们分析了注入失败和部分符号循环带来的构念效度限制。

英文摘要

Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as a single unit and require one or more full-model calls per turn. We propose SAGE (State-Grounded Abstention-Aware Evaluation), which compiles a workflow specification and per-turn state diff into atomic, schema-grounded criteria and routes each through a cascade of symbolic and encoder/NLI verifiers that abstain rather than guess, aggregating criterion verdicts into a turn-level decision with an evidence trace. Its recommended operating point, SAGE-Core, decides 81--91% of criteria with only the compiler, symbolic rules, and on-device encoders--at zero paid LLM cost--while SAGE-LLM adds an optional focused-LLM fallback for open-class criteria. Across four slices spanning MultiWOZ, Schema-Guided Dialogue, and ABCD, no evaluated LLM-as-a-judge baseline--including a state-aware GPT-4.1 judge and cheaper GPT-4.1-mini variants--significantly exceeds SAGE-Core on any slice, even though the GPT-4.1 G-Eval judge costs $4.7--8.0 per 1,000 turns to SAGE-Core's $0. A two-annotator human audit (n=200, $κ$=0.94) confirms strong label fidelity on the transcript-visible failure classes--where, excluding the weak-salience IUV class, SAGE-Core is statistically tied with the strongest LLM judge--and honestly scopes ignored-user-value as a state-consistency signal with weak broad-human salience. We analyze construct-validity limits from injected failures and partial symbolic circularity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑