arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GameCommBench:面向AI生成游戏评论的统一基准与类型感知评估

GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary

Qirui Zheng, Zhengteng Lin, Yunyi Xiao, Junhao Li, Keyuan Cheng, Xingbo Wang, Yongyi Wang, Lingfeng Li, Yunlong Lu, Wenxin Li

arXiv 2610.11129首次发表:更新:

发表机构

Peking University; South China University of Technology(北京大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对AI生成游戏评论的碎片化问题,构建涵盖多类游戏的统一基准GameCommBench,提出类型感知评估框架TACE并验证其可靠性,揭示AI评论员在实时观察与策略分析上的瓶颈,为AI-GGC评估提供诊断基础。

AI 中文摘要

游戏评论是一项开放式生成任务,需要多模态感知、策略推理和上下文知识。现有的AI生成游戏评论(AI-GGC)研究在游戏、模态和评估协议方面仍呈碎片化,而基于重叠或整体的评估器无法捕捉评论的功能异质性。我们引入GameCommBench,这是一个涵盖棋盘游戏、体育和电子竞技的统一基准,其评论与异质游戏上下文对齐并按评论类型标注。我们进一步提出类型感知评论评估(TACE),这是一个用于评估不同类型评论的结构化框架。随后我们验证了TACE的可靠性和人类一致性,并使用它对代表性AI评论员进行基准测试。结果显示AI评论员的能力分布不均,其中实时观察和策略分析是主要瓶颈。GameCommBench与TACE共同为可比较、可解释的AI-GGC评估提供了诊断基础。

英文摘要

Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑