arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14200cs.LG

用于流式视频游戏中鲁棒且高效模仿学习的增强方法

Augmentations for Robust and Efficient Imitation Learning in Streamed Video Games

Somjit Nath, Abdelhak Lemkhenter, Pallavi Choudhury, Chris Lovett, Katja Hofmann, Sergio Valcarcel Macua, Lukas Schäfer

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对流式视频游戏中模仿学习面临的挑战,提出流式增强方法,基于预测逆动力学模型,在现代3D游戏三个任务中评估其效果,结果显示该增强方法能有效提升代理性能,是训练鲁棒高效游戏代理的有力工具。

中文摘要 AI 辅助

模仿学习是一种通过训练策略将视觉观察映射到人类示范动作,从而将游戏代理扩展到复杂3D环境的有吸引力的方法。然而,收集这些示范成本高昂,且现代游戏常通过流式传输进行,网络延迟和压缩会引入时空相关的视觉伪影,导致测试时的协方差偏移。为应对这些挑战,我们提出了流式增强方法,模仿在低带宽网络连接流式传输中常见的四种伪影:像素化块和擦除、全局模糊和重影。我们在预测逆动力学模型(PIDM)之上实例化我们的方法,该模型在学习的潜在空间中将未来状态条件与逆动力学策略相结合,并在现代3D电子游戏的三个任务中评估我们的增强方法的影响。在稳定的流式传输条件下,与在相同数据预算下未进行增强训练的代理相比,经过时空增强训练的代理在评估性能上提高了41%。当引入网络延迟时,经过增强训练的代理性能仅下降7.45%,而仅使用原始数据训练的代理性能下降49.82%。这些结果清楚地表明,为流式传输设置量身定制的时空增强是训练鲁棒且高效的游戏代理的一种简单而强大的工具。

英文摘要

Imitation learning is an appealing way to scale game-playing agents to complex 3D environments by training policies to map visual observations to actions from human demonstrations. However, these demonstrations are expensive to collect and modern game-playing is often done through streaming in which network delay and compression introduce spatiotemporally correlated visual artifacts that can cause a covariance shift at test time. To address these challenges, we propose streaming augmentations that mimic four types of artifacts commonly encountered during streaming with low-bandwidth network connection: pixelated blocks and scrubs, global blur, and ghosting. We instantiate our approach on top of predictive inverse dynamics models (PIDM), which combine future-state conditioning with an inverse dynamics policy in a learned latent space, and evaluate the impact of our augmentations across three tasks in modern 3D video games. Under stable streaming conditions, agents trained with spatiotemporal augmentations achieve up to 41% higher evaluation performance compared to agents trained without augmentations under an identical data budget. When network lag is introduced, agents trained with augmentations degrade by only 7.45% vs 49.82% of the original performance for agents trained only with the original data. These results clearly indicate that spatiotemporal augmentations tailored for the streaming setting are a simple yet powerful tool to train robust and efficient game-playing agents.

补充信息

↑