arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Marionette:预测世界状态、渲染几何、绘制外观

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang

arXiv 2608.14530首次发表:更新:

发表机构

Alaya Lab; Shanghai Innovation Institute; Huazhong University of Science and Technology(Alaya实验室; 上海创新研究院; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Marionette是面向带关节角色的交互式游戏的世界模型,通过显式建模3D世界状态、委托几何计算、合成外观,实现了可控的长时序行为,且外观生成无明显保真度损失。

AI 中文摘要

交互式游戏世界模型通常在像素或隐空间中自回归地直接生成视觉观测,迫使姿态、几何和遮挡等结构化属性由同一个生成序列隐式维护。在长时序下,这些隐式世界属性的误差会累积,导致一致性和可控性脆弱。我们显式建模演化的世界状态,将精确的几何计算委托给固定的零参数渲染器,而让神经模型合成外观。我们将该思想实例化为Marionette,这是一个针对带关节角色的交互式游戏的世界模型。首先,一个两阶段自回归动力学模型预测显式且可解释的276维3D世界状态,该状态包含多实体带关节骨架、度量根轨迹和旋转。其次,一个零参数图形桥将预测的状态转换为姿态控制视频,以闭式形式计算世界空间几何和遮挡。第三,一个控制条件的视频扩散观测模型从得到的结构化控制中合成 photorealistic RGB 观测。我们的实验确立了Marionette的两个特性:第一,预测的世界状态是直接可控的,强制不匹配的动作流会在48个保留片段中使根对齐关节误差变化31%;第二,长时序行为由状态决定,且可在状态中修复,若放任不管,两个生成的角色会漂移至相距21.2米(记录会话保持在5米附近),且三分之一的帧会出现地面穿透,对显式状态施加地形碰撞器和分离上限这两个规则,可将穿透率降低66%并保持角色互动,且无需改变观测模型。通过预测状态路由外观不会造成可检测的保真度损失,其FVD为831,而记录姿态的FVD为799。

英文摘要

Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two-stage autoregressive dynamics model predicts an explicit and interpretable 276-dimensional 3D world state comprising multi-entity articulated skeletons, metric root trajectories, and rotations. Second, a zero-parameter graphics bridge converts the predicted state into pose-control videos, computing world-space geometry and occlusion in closed form. Third, a control-conditioned video-diffusion observation model synthesizes photorealistic RGB observations from the resulting structured controls. Our experiments establish two properties of Marionette. First, the predicted world state is directly controllable. Forcing a mismatched action stream changes root-aligned joint error by 31% across 48 held-out segments. Second, long-horizon behaviour is determined in the state, and can be repaired there. Left free, the two generated characters drift to 21.2 m apart (recorded sessions stay near 5 m) and a third of frames show ground penetration. Two rules imposed on the explicit state, a terrain collider and a separation cap, cut penetration by 66% and keep the pair engaged, with no change to the observation model. Routing appearance through the predicted state costs no fidelity we can detect, at an FVD of 831 against 799 for recorded pose.

CommentsProject page: https://alayalab.github.io/Marionette/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑