arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视频游戏数据标注的视觉语言模型(VLMs)

VLMs for Videogame Data Annotation

Katrin Schmid, Iuri Frosio

arXiv 2608.05949首次发表:更新:

AI 中文总结

该研究针对VLMs在视频游戏应用受限的问题,探究其用于游戏帧序列奖励信号标注的可行性,分析其在游戏问答中的不足并提出应对措施,还研究了输入参数对标注质量的影响。

AI 中文摘要

视觉语言模型(VLMs)与人工智能(AI)智能体已彻底改变工程师处理现实世界应用中复杂问题的方式,但它们在视频游戏中的应用却因合成场景的极端多变性及对现实世界物理规则的适配性较差而受限。本研究探讨将VLMs用于视频游戏帧序列的奖励信号标注,该任务有诸多潜在应用,包括条件训练与离线强化学习等。研究表明,VLMs常难以回答竞速类视频游戏的基础问题(尽管在其他游戏类型中也观察到类似表现),并讨论了VLM输出混合、提示优化等应对措施,同时探究了输入序列长度、分辨率及问题批处理对标注质量与令牌消耗的影响。

英文摘要

Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑