arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05954cs.AIcs.CVcs.LG

基于VLM标注数据集训练条件化电子游戏智能体

Training a Conditioned Video Game Agent on a VLM Annotated Dataset

Katrin Schmid, Iuri Frosio

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对电子游戏强化学习中奖励获取、加权及稀疏性等问题,提出用VLM标注数据集,结合离线RL训练条件化智能体,并分析实验中的困难与局限。

中文摘要 AI 辅助

强化学习(RL)是一种强大但远未达到易用程度的策略学习技术。在电子游戏这一特定场景中,需要访问游戏引擎以获取训练所需的奖励(例如从环境中收集奖励)。此外,对奖励进行恰当的识别与加权通常需要采用困难的试错方法。最后,奖励往往是稀疏的,理解它们最终如何影响学习到的策略是一项非平凡的工作。为缓解这些问题,我们提出使用视觉语言模型(VLM)对电子游戏数据集进行标注,指示VLM提取人类定义的奖励。我们表明,随后可以使用离线强化学习(offline RL)来训练一个条件化智能体,该智能体可对期望回报做出相应响应,同时我们也讨论了早期实验中出现的困难与局限性。

英文摘要

Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards generally requires a difficult trial-and-error approach. Lastly, rewards are often sparse and understanding how they eventually affect the learned policy is a non-trivial exercise. To ease these issues we propose annotating a video game dataset with Vision Language Models (VLMs) instructed to extract human defined rewards. We show that offline RL can then be used to train a conditioned agent that responds accordingly to the desired returns and we discuss the difficulties and limitations that emerged in our early experiments.

↑