arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22739cs.CVcs.AIcs.LG

Cortex:具有冻结视觉特征的《雷神之锤》紧凑行为克隆

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

Dzmitry Malyshau

首次发表
浏览论文内容

中文总结 AI 辅助

研究在不添加强化学习等情况下,简单行为克隆策略在《雷神之锤》游戏中的进展。核心方法是基于冻结视觉特征的Cortex策略,在Pixels2Play语料库子集训练。主要贡献是评估了该策略表现,通过消融实验揭示影响因素,还发布了相关内容。

中文摘要 AI 辅助

我们研究了在不添加强化学习或显式记忆的情况下,一个刻意简单的行为克隆策略在视觉丰富的第一人称游戏中能取得多大进展。Cortex是一个紧凑的《雷神之锤》策略,在冻结的DINOv3编码器上的六层变压器中有1098万个可训练参数。它在公共Pixels2Play语料库的《雷神之锤》子集中进行训练,使用了6849个记录。评估了两批20个随机的、120秒的情节。Cortex虽未完成关卡,但每集都能到达开门处、按钮房和闸门下降处,且每批20集中有19集至少有一次击杀。消融实验表明了不同因素的影响,剩余失败与协变量转移一致,促使有针对性的校正数据。我们发布了策略实现、检查点和一个代表性的展示。

英文摘要

We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a six-layer transformer over a frozen DINOv3 encoder. It is trained on the Quake subset of the public Pixels2Play corpus: 6,849 recordings (about 474.7 hours), represented as 17.09 million cached decision frames with keyboard and mouse actions. One sampled training epoch uses 517,048 four-frame windows and takes 3.3 minutes of policy-head optimization on one RTX 5080, excluding one-time feature extraction. We evaluate two independent batches of 20 stochastic, 120-second episodes on Quake E1M1. Cortex does not complete the level, but every episode reaches the opening door, button room, and gate descent; 19 of 20 episodes in each batch record at least one kill. Under the same time-controlled harness, released P2P-150M and NitroGen checkpoints remain shallower in five matched-duration episodes each. These comparisons are limited by small reference samples and different native interfaces. Ablations show that denser visual tokens improve combat and survival, while longer optimization and naive action history improve offline metrics without consistently improving play. The remaining failures are consistent with covariate shift and motivate targeted corrective data. We release the policy implementation, checkpoint, and a representative rollout.

补充信息

↑