发表机构
Five9(Five9)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对对话情感识别,提出置信度门控混合系统,仅在集成模型低置信时调用LLM,在三个数据集上以更低成本实现帕累托最优性能,为CCaaS平台提供可解释的LLM支出分配方案。
AI 中文摘要
对话中的情感识别(ERC)是联络中心即服务(CCaaS)平台中代理辅助提示、升级路由和通话后分析背后的生产能力,在这些平台中,成本和延迟约束与准确性同等重要。我们报告了三种用于对话上下文ERC的部署选项的系统级比较:低成本堆叠集成(句子嵌入、窗口化上下文、随机森林/XGBoost/逻辑回归堆叠)、现成的LLM提示(GPT-4o-mini;零样本、少样本、思维链),以及一种置信度门控混合系统,该系统仅将集成模型最不自信的预测升级到LLM——模拟了生产联络中心中IVA到人工代理的升级策略。在IEMOCAP上,集成模型显著优于所有LLM配置(加权F1为0.595对比0.460-0.536,p<0.0001),且成本仅为LLM的一小部分,延迟低于10毫秒;在MELD和CMU-MOSI上,排名反转,表明两种纯系统都不是安全默认选择。置信度门控混合系统通过在三个数据集上帕累托支配两种纯系统(加权F1分别为0.620、0.643、0.824),同时将大部分流量路由到近零成本的集成模型,解决了这一问题,转化为每百万话语约10-85美元的成本,而纯LLM流水线为99-170美元。升级策略并非不透明的成本/准确性调节旋钮:升级的话语不成比例地跟随情感或情绪转变,为操作员提供了可解释、可审计的路由信号,且集成模型的置信度校准良好,安全地偏向欠自信而非过度自信。该模式在三个数据集和两个LLM提供商中保持一致。置信度门控级联在通用ML系统中已确立;我们的贡献在于展示其能干净地迁移到对话上下文ERC,为CCaaS和对话式AI平台决定如何分配LLM支出提供了具体的部署方案。
英文摘要
Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service (CCaaS) platforms, where cost and latency constraints matter as much as accuracy. We report a systems-level comparison of three deployment options for dialogue-contextual ERC: a low-cost stacked ensemble (sentence embeddings, windowed context, RandomForest/XGBoost/logistic-regression stacking), off-the-shelf LLM prompting (GPT-4o-mini; zero-shot, few-shot, chain-of-thought), and a confidence-gated hybrid that escalates only the ensemble's least-confident predictions to the LLM - modeled on IVA-to-human-agent escalation policies used in production contact centers. On IEMOCAP, the ensemble significantly outperforms every LLM configuration (0.595 vs. 0.460-0.536 weighted F1, p < 0.0001) at a fraction of the cost and sub-10ms latency; on MELD and CMU-MOSI the ranking reverses, showing neither pure system is a safe default. The confidence-gated hybrid resolves this by Pareto-dominating both pure systems on all three datasets (0.620, 0.643, 0.824 weighted F1) while routing the majority of traffic through the near-zero-cost ensemble, translating to roughly $10-85 per million utterances versus $99-170 for an LLM-only pipeline. The escalation policy is not an opaque cost/accuracy dial: escalated turns disproportionately follow an emotion or sentiment shift, giving operators an interpretable, auditable routing signal, and the ensemble's confidence is well-calibrated and safely under- rather than over-confident. The pattern holds across three datasets and two LLM providers. Confidence-gated cascading is established in general ML systems; our contribution is showing it transfers cleanly to dialogue-contextual ERC, yielding a concrete deployment recipe for CCaaS and conversational-AI platforms deciding how to allocate LLM spend.