发表机构
Lumytics(卢米提克斯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出三层评估框架,结合NPC模拟器验证LLM聊天智能体的多轮对话目标达成能力,其模拟器成本低、效率高,可用于CI/CD集成,相关资源已公开供其他团队使用。
AI 中文摘要
部署大语言模型(LLM)聊天智能体的生产团队面临一个特定的质量保证缺口:现有评估工具仅测试单个响应或模拟社交互动,但无一能系统验证真实用户是否能通过多轮对话实现其目标。我们提出一种三层内部试用框架以填补该缺口,该框架结合了标准问题库测试(第一层)、随机游走多轮评估(第二层),以及包含五类结构化目标类型和十类失败分类法的目标导向NPC(非玩家角色)模拟器(第三层)。在对生产级多智能体系统开展的为期约三个月的纵向案例研究中(共257次评估运行、1套含108个场景的NPC套件),我们发现三层框架会产生互补的回归信号:同步运行内响应质量的跨层相关性较弱(斯皮尔曼相关系数rho介于-0.15至0.14之间),而纵向序列间呈负相关(rho低至-0.46),证实标准正确性无法预测目标导向对话的成功。该NPC模拟器的目标达成率达77%,每次运行成本为0.17美元(比人工评估便宜6272倍),支持每日持续集成/持续部署(CI/CD)集成及自动PROMOTE/HOLD/ROLLBACK发布决策。我们公开了完整的提示模板、失败分类法和以Python优先的可复现性指南,以便其他团队能将该框架应用于自身的LLM聊天智能体。
英文摘要
Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through multi-turn conversation. We introduce a three-layer dogfooding framework that bridges this gap by combining canonical question-bank testing (Layer 1), random-walk multi-turn evaluation (Layer 2), and a goal-directed NPC (Non-Player Character) simulator with five structured goal types and a ten-category failure taxonomy (Layer 3). In a longitudinal case study on a production multi-agent system over roughly three months (257 evaluation runs; a 108-scenario NPC suite), we find that the three layers produce complementary regression signals: cross-layer correlation for response quality is weak within a synchronized run (Spearman rho between -0.15 and 0.14) and negative across the longitudinal series (rho down to -0.46), confirming that canonical correctness does not predict goal-directed conversation success. The NPC simulator achieves 77 percent goal achievement at 0.17 dollars per run (6,272x cheaper than human evaluation), enabling daily CI/CD integration with automated PROMOTE/HOLD/ROLLBACK release decisions. We release full prompt templates, the failure taxonomy, and a Python-first replicability guide so that other teams can adopt the framework for their own LLM chat agents.
Comments24 pages, 11 tables, 3 figures