CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
CirrusBench:在真实云服务环境中评估基于LLM的代理超越正确性的能力
机构 * School of Mathematics and Sciences, Fudan University(复旦大学数学科学学院) ; Shanghai Center for Mathematical Sciences, Fudan University(复旦大学上海数学中心) ; Alibaba Group(阿里巴巴集团) ; Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学类脑智能科学与技术研究院) ; Center for Applied Mathematics & Shanghai Key Laboratory of Contemporary Applied Mathematics, Fudan University(复旦大学应用数学中心暨上海市现代应用数学重点实验室)
AI总结 本文提出CirrusBench框架,基于真实云服务工单数据评估LLM代理的鲁棒性和解决效率,引入客户导向指标量化服务质量,揭示现有模型在复杂多轮任务中的不足。
Comments Submitted for SIGKDD 2026