arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ABot-World-0:在单台桌面GPU上进行无限交互式世界展开

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo

arXiv 2607.19191首次发表:更新:

AI 中文总结

介绍ABot-World-0这一用于实时长视野闭环交互的动作条件视频世界模型,利用多源数据学习世界动态。通过多种技术提炼模型,设计控制界面与部署堆栈,在单台桌面GPU上实现高效视频流传输,实验验证其有竞争力的可控性和世界演变能力。

AI 中文摘要

我们展示了ABot-World-0,这是一种用于实时、长视野闭环交互的动作条件视频世界模型,由跨越AAA游戏、模拟引擎和互联网视频的多源数据基础设施支持,以学习可控的世界动态。WorldExplorer在训练反馈的引导下执行智能体驱动的收集,同时统一管道应用14种确定性质量检查、基于VLM的评估以及同步动作和文本注释。我们通过教师强制和ODE蒸馏将双向动作条件教师逐步提炼为因果学生,并引入LongForcing以使长学生自展开与扩展视野教师对齐,减轻累积分布偏移和自回归漂移。原始键盘动作提供用于场景漫游和第三人称角色交互的统一控制界面,而参考角色记忆在第三人称展开期间提供用于身份一致性的持久外观线索。对于部署,我们与轻量级VAE解码器、高效注意力、内存感知调度和低位DiT推理共同设计了一个流推理堆栈。在优化的低位配置下,ABot-World-0在单台NVIDIA RTX 5090桌面GPU上以高达16 FPS的速度流式传输720P视频,动作到第一帧的延迟为1.2秒,峰值VRAM约为19GiB。在WorldRoamBench和扩展交互式展开上的实验证明了具有竞争力的可控性和连贯的长视野世界演变。

英文摘要

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑