arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21712cs.CV

ZYT-World:用于闭环自动驾驶仿真的实时可控世界模型

ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation

Boni Hu, Xiong Wei, Haoming Huang, Yong Huang, Chenbo Wang, Yi Yang, Jiancheng Wang, Ruicheng Zhu, Zhimin Yang, Guanglai Liu, Qiaowan Jin, Dongzhuo Wang, Haiwei… 展开作者

Boni Hu, Xiong Wei, Haoming Huang, Yong Huang, Chenbo Wang, Yi Yang, Jiancheng Wang, Ruicheng Zhu, Zhimin Yang, Guanglai Liu, Qiaowan Jin, Dongzhuo Wang, Haiwei Kuang, Jiajun Fan, Yue Wu, Jiaxin Wei, Hao Sun, Feihong Yan, Yuyao Zhou, Wei Bi, Kaixuan Wang, Zichao Guo, Xiaozhi Chen

首次发表
浏览论文内容

中文总结 AI 辅助

ZYT-World提出单一架构原生生成四鱼眼三针孔视图,通过Plucker适配器、自适应层归一化、布局条件及蒸馏实现实时可控闭环仿真,单步模型速度提升107.7倍并保留90%以上图像质量。

中文摘要 AI 辅助

生成式世界模型为端到端和视觉-语言-动作驾驶策略提供了可控且可重复的闭环仿真,但生产部署暴露了三个未解决的需求:以原生分辨率忠实再现混合鱼眼-针孔相机装置;协调因果的、逐时间步的交互与长时程稳定性和低延迟;以及在重新访问某个地点时保持场景身份。我们提出了ZYT-World,这是一种单一架构,原生生成视场角大于180°的四个鱼眼视图和三个针孔视图。投影特定的Plucker适配器编码相机几何,自车运动自适应层归一化提供全局运动控制,轻量级像素对齐布局通过实例级边界框、朝向和颜色来条件化交通参与者与信号。异构训练结合了全装置几何覆盖与高分辨率细节。教师强制、因果一致性蒸馏、自滚动分布匹配蒸馏以及RigCritic将40步双向教师模型转化为单步、逐潜变量流式生成器,其中RigCritic联合评估七视图装置。一个19M参数的变分自编码器解码器(TinyVAE)、W8A8量化以及我们的推理引擎分别降低了解码、主干和增量执行成本。最后,从真实采集数据中导出的跨轨迹对训练了一个即插即用的隐式记忆模块,以保留地点特定证据。在内部多视图测试集上,单步模型保留了教师模型超过90%的PSNR和SSIM,而FID、FVD和LPIPS保持在教师模型的11%以内。在图2的仅生成器计时下,它比40步双向教师模型快107.7倍。TinyVAE解码比Wan快59.8倍。30秒滚动和跨轨迹重访展示了预期的长时程和记忆行为。

英文摘要

Generative world models offer controllable and repeatable closed-loop simulation for end-to-end and vision-language-action driving policies, but production deployment exposes three unresolved requirements: faithfully reproducing a mixed fisheye-pinhole rig at native resolutions; reconciling causal, per-timestep interaction with long-horizon stability and low latency; and preserving scene identity when a location is revisited. We present ZYT-World, a single architecture that natively generates four fisheye views with field of view > 180° and three pinhole views. Projection-specific Plucker adapters encode camera geometry, ego-motion adaptive layer normalization provides global motion control, and a lightweight pixel-aligned layout conditions traffic participants and signals through instance-level boxes, headings and colors. Heterogeneous training combines full-rig geometric coverage with high-resolution detail. Teacher forcing, causal consistency distillation, self-rollout distribution matching distillation, and RigCritic transform a 40-step bidirectional teacher into a one-step, per-latent streaming generator, with RigCritic evaluating the seven-view rig jointly. A 19M-parameter variational autoencoder decoder (TinyVAE), W8A8 quantization, and our inference engine reduce decoding, backbone, and incremental-execution costs, respectively. Finally, cross-trajectory pairs derived from real captures train a plug-in implicit-memory module that preserves place-specific evidence. On the internal multi-view test set, the one-step model retains more than 90% of the teacher's PSNR and SSIM, while FID, FVD, and LPIPS stay within 11% of the teacher. Under the generator-only timing in Figure 2, it is 107.7 times faster than the 40-step bidirectional teacher. TinyVAE decodes 59.8 times faster than Wan. 30s rollouts and cross-trajectory revisits show the intended long-horizon and memory behavior.

发表机构

  • ZYT AI Team(ZYT AI团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑