实现对话代理的多维评估:一种具有选择性重新评估和模型基准测试的可扩展、受治理的管道
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
- Lowes(劳氏公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对零售对话代理评估难题,提出GenAI Evaluation管道,通过规范化等处理生产日志,能评估多维度指标。选择性重新评估提高效率,支持审计。每日处理约50000条记录,已评估超两百万次交互,取得较好分数和准确率。
AI中文摘要:
评估零售对话代理需要超越词汇重叠指标的方法来评估意图对齐、事实性、有用性、清晰度、语气和整体响应质量。虽然基于大语言模型的评判方法为人工评估提供了可扩展的替代方案,但生产部署在治理、可重复性、成本、模式一致性、可追溯性和可靠性方面带来了挑战。我们提出了GenAI评估,这是一个用于大规模评估零售对话系统的受治理、配置驱动的管道。它通过规范化、分片、异步执行和模式约束的大语言模型评分来处理生产聊天机器人日志。该框架评估有用性、真实性、清晰度、语气对齐和特定于翻译的维度。选择性重新评估仅处理不完整、格式错误或模式无效的记录,而模式锁定、版本化配置、验证日志和记录级出处支持可审计性。该框架每天处理约50,000条记录,已评估超过200万次交互。验证使用了来自四名训练有素的注释者的12,980条分层随机人工标记记录。分类涵盖14个意图、156个子意图、18个主要领域和129个子领域。该管道实现了0.93的宏F1分数和89%的翻译人工可接受性准确率。
英文摘要:
Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation of retail conversational systems. It processes production chatbot logs through normalization, sharding, asynchronous execution, and schema-constrained LLM scoring. The framework evaluates helpfulness, truthfulness, clarity, tone alignment, and translation-specific dimensions. Selective re-evaluation processes only incomplete, malformed, or schema-invalid records, while schema locking, versioned configurations, validation logs, and record-level provenance support auditability. The framework processes approximately 50,000 records daily and has evaluated more than two million interactions. Validation used 12,980 stratified-random human-labeled records from four trained annotators. Classification covered 14 intents, 156 sub-intents, 18 major domains, and 129 sub-domains. The pipeline achieved a macro F1 score of 0.93 and 89% human-acceptability accuracy for translation.