发表机构
Sprinklr AI(Sprinklr AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对比人工与GPT-4.1、GPT-5对电信和零售语音智能体对话的评估,发现LLM评估可靠性依赖指标与配置,可作为规模化评估组成部分,支持混合评估 pipeline。
AI 中文摘要
规模化评估对话式语音智能体需要可靠的评估方法,既要捕捉可观测的交互质量,也要体现人工评估者通常具备的情境判断能力。本研究对“LLM作为评判者”的评估方法展开调查,将人工评判与GPT-4.1、GPT-5在电信和零售领域的语音智能体对话上的结果进行对比,覆盖对话质量与安全两个维度。相同的交互内容在三种评估配置p0、p1和p2下被打分,以检验自动评判是否对评估设置敏感,以及观测到的模式是否能跨配置和评判模型推广。除了总体一致性,本研究还考察了指标级相关性、评估者一致性以及人工与LLM之间的系统性分歧,以确定哪些对话属性可由自动化可靠评判,哪些仍对解释和情境敏感。有效的语音智能体评估还受 pipeline 级因素影响,如语音生成、流式传输以及ASR、推理和工具调用阶段的错误传播,这促使我们聚焦于对比人工与LLM评判者对相同端到端交互的打分情况。研究结果表明,基于LLM的评估可作为规模化语音智能体评估的有效组成部分,但其可靠性取决于具体指标和配置,而非统一不变。本研究提供了一个实证框架,用于识别哪些指标适合自动化评估,并支持混合 pipeline,其中LLM评判者负责规模化评估,而人工评估者则参与那些需要情境解释和高置信度判断的指标。
英文摘要
Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.
CommentsExtends LLM-as-a-Judge to voice agents across telecom and retail, testing GPT-4.1, GPT-5 and Claude against human raters across 10 safety and efficiency metrics. A correlation-based calibration analysis reveals domain-dependent reliability and identifies Recovery Turn Count and safety-recall metrics as unreliable for fully automated judging