arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31473cs.AI

游戏竞技场:竞争环境中的战略LLM评估

Game Arena: Strategic LLM Evaluation in Competitive Environments

Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hw… 展开作者

Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文介绍Kaggle游戏竞技场,一个通过竞争性游戏评估LLM的平台,涵盖国际象棋、扑克和狼人杀,旨在防止性能饱和,确保评估的可重复性和泛化性。

中文摘要 AI 辅助

我们推出了Kaggle游戏竞技场,这是一个开放且不断扩展的平台,旨在通过竞争性游戏评估大型语言模型(LLMs)。与静态基准不同,游戏竞技场使模型能够在结构化环境中进行正面交锋,其中游戏强度随着模型的进化而自然增加,从而防止性能饱和。本技术报告详细介绍了游戏竞技场背后的基础设施,并描述了三个试点游戏环境:国际象棋、扑克和狼人杀。这些环境涵盖了完美信息、不完美信息和多人游戏设置,从而能够系统地研究模型在不确定性下的战略规划、适应性和鲁棒性。对于每个游戏,我们提供了环境的详细描述、评估指标以及跨模型运行完整比赛的结果。通过稳健的基础设施和大规模基于真实情况的评估,游戏竞技场确保了可重复性、透明性以及对新游戏和变体的泛化能力。

英文摘要

We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.

补充信息

↑