arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从像素到按键:探索游戏逆向动力学中的空间与运动线索

Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics

Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio

arXiv 2609.37907首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在数据受限场景下探索游戏逆向动力学模型,分析空间运动特征、架构与训练目标的影响,发现架构和运动流提取关键,并揭示不平衡指标局限及动作歧义需显式建模。

AI 中文摘要

视频游戏为研究具身智能中的感知与控制提供了可扩展的环境。在线游戏视频可以提供演示,但很少包含用于训练玩家输入。因此,逆向动力学模型(IDMs)被提出用于从帧中推断输入。在约1K-2K游戏小时上训练的大型(高达10亿参数)IDMs展示了在该规模下的可行性和跨环境泛化能力,但研究人员未阐明恢复单个动作的关键组成部分,且通常仅报告可能掩盖罕见动作失败的总体准确率。我们在数据受限场景中研究该问题,以评估空间运动特征、模型架构和训练目标如何影响IDM的输出,并在每键和平衡指标(如$F_1^{macro}$)上分析我们的模型。我们在Trackmania上的实验强调了模型架构和预处理中运动流提取等因素的重要性,同时展示了通过不平衡指标评估的局限性。将相同架构和训练方案应用于Cyberpunk 2077揭示了不同游戏机制间的不均匀性能。我们的逐动作评估和失败分析突出了由相机运动、延迟效应和不平衡按键频率引起的模糊性,这些模糊性要求未来实现中显式建模3D场景结构、长期状态并采用适当的损失函数。

英文摘要

Video games offer scalable environments for studying perception and control in embodied agents. Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

CommentsAccepted at the Workshop on Multimodal Digital Agents (ECCV 2026): https://mda-workshop.allen.ai/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑