arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21267cs.AIcs.SE

生产环境中的高效基准测试:一个不断演进的LLM智能体研究

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Yining She, Lei Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对生产LLM智能体重复评估成本高的问题,比较多种高效采样方法,发现多维2PL自适应测试保真度最佳,但最终部署难度分层固定子集,因其操作简单且可迁移,并给出实用建议。

中文摘要 AI 辅助

生产环境中的LLM智能体在演进过程中会被反复评估,但完整的智能体基准测试重新运行成本高昂。我们针对一个服务数万月度活跃用户的生产分析智能体,研究了高效的周期性评估方法,并报告了第一手的部署经验。利用生产基准测试的574次历史运行,按时间顺序划分为校准期和保留期,我们比较了随机采样、历史缓存、固定代表性子集以及基于IRT的自适应测试。结果表明,多维2PL自适应测试实现了最佳的整体评分保真度:执行200个问题,占完整运行的38.5%,产生的平均绝对误差(MAE)为1.03个百分点。尽管如此,我们最终部署了难度分层的固定子集,因为其操作简单,并展示了这些子集无需重新校准即可迁移到其他五个智能体家族,且在校准窗口短至一天的情况下仍保持稳定。基于这一部署经验,我们为生产智能体的周期性评估提出了实用建议。

英文摘要

Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching, fixed representative subsets, and IRT-based adaptive testing. The results show that multidimensional 2PL adaptive testing achieves the best overall score fidelity: executing 200 questions, 38.5% of a full run, yields 1.03 pp of MAE. We nevertheless deployed difficulty-stratified fixed subsets because of their operational simplicity, and show they transfer without recalibration to five other agent families and remain stable across calibration windows as short as one day. Drawing on this deployment experience, we report practical recommendations for recurring production-agent evaluation.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑