Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
Collab-Overcooked: 大语言模型作为协作代理的基准测试与评估
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Li Auto Inc
AI总结 本文提出Collab-Overcooked基准测试,评估大语言模型在协作代理中的表现,揭示其在目标解释和协作能力上的优劣。
Comments Accepted to EMNLP 2025 Main Conference. Camera-Ready Version. 30 pages, 17 figures