Proxy2World:从轻量级代理学习生成世界而无需看到它们
Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Proxy2World提出一种从普通RGBD视频学习生成世界的可控模型,无需配对代理数据,通过跨模态流匹配联合训练深度条件RGB生成与RGBD生成,实现代理-相机混合去噪,在ProxyBench上平衡结构遵循与视觉质量。
AI中文摘要:
轻量级场景代理让创作者能够控制场景布局和运动,同时为外观、光照和视觉效果留下想象空间。然而,合适的代理并非唯一定义,这使得配对代理-视频数据难以大规模自动构建。我们提出Proxy2World,一种可控世界模型,它从普通的带姿态的RGBD视频中学习这些互补能力,而无需在作者创作的代理-视频对上进行训练。该模型通过跨模态流匹配联合学习深度条件下的RGB生成和联合RGBD生成。学习这两个任务使得推理时的代理-相机混合去噪能够遵循代理结构,同时产生自然、详细的视觉效果。我们进一步引入ProxyBench,在多样化的场景、相机轨迹和主体运动集合上评估这一能力。在ProxyBench上的实验表明,Proxy2World在结构遵循和视觉质量之间取得了比相机控制和几何条件方法更好的平衡,这得到了定量指标、VLM评估、人类评估和多样化定性结果的支持。
英文摘要:
Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2World, a controllable world model that learns these complementary capabilities from ordinary posed RGBD videos, without training on authored proxy-video pairs. The model jointly learns depth-conditioned RGB generation and joint RGBD generation through cross-modal flow matching. Learning both tasks enables proxy-camera hybrid denoising at inference to follow the proxy structure while producing natural, detailed visuals. We further introduce ProxyBench to evaluate this capability across a diverse set of scenes, camera trajectories, and subject motions. Experiments on ProxyBench show that Proxy2World achieves a better balance between structural adherence and visual quality than camera-controlled and geometry-conditioned methods, supported by quantitative metrics, VLM assessments, human evaluations and diverse qualitative results.