[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
通过可评分游戏评估大语言模型在谈判中的能力基准测试
机构 * Informatics Institute University of Amsterdam(信息学院阿姆斯特丹大学)
AI总结 本文通过可评分游戏评估大语言模型在谈判中的能力,发现基准复杂但模型比较存在歧义,强调上下文在评估中的重要性。
Comments Accepted for publication at Transactions on Machine Learning Research (TMLR) and MLRC Journal Track, 2025. Code available at: https://github.com/joshrosie/FACT29
Journal ref Transactions on Machine Learning Research (TMLR), June 2025