arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28491cs.AIcs.RO

AcrossVAM1.0:面向文本辅助机器人视频预测的粒子世界建模

AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

Yafei Zhang, Nan Wu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出AcrossVAM1.0模型,通过分解为粒子运动与外观,在文本辅助下实现机器人视频预测,在VRS基准上显著降低轨迹误差、提升PSNR/SSIM等指标,但仍存在LPIPS未超越基线等局限。

中文摘要 AI 辅助

预测机器人视频需要精确的运动推理和高频外观的保留,但整体像素模型会将这些目标纠缠在一起,且常将其进展掩盖在强大的最后一帧基线之后。我们提出AcrossVAM1.0,一种轻量的文本辅助视频动作模型,它将未来预测分解为以对象为中心的运动和密集外观。一个冻结的SAM3-DLP编解码器将4帧上下文分解为机器人、机械臂、夹爪的语义粒子,以及一个背景潜变量。一个参数规模为0.28M的时空Transformer对齐粒子身份、将其状态向前滚动,并通过FiLM由冻结的OpenCLIP指令嵌入进行调制。一个因果双流解码器将粒子渲染的运动与仅从最后一帧观测编码的外观相结合;一个残差细化器和学习到的交付掩码生成5个未来帧,且不访问未来外观。在我们由多样化真实机器人轨迹构建的VRS基准上,粒子动力学相比持久性将轨迹误差降低了21.0%。在3个交付掩码种子上,AcrossVAM1.0将未来帧的PSNR/SSIM从19.97/0.796提升至20.573/0.8004,而原始粒子生成将运动区域的PSNR从11.89提升至13.23。交付模型在LPIPS上尚未超越持久性,且正确与打乱的语言将轨迹误差仅改变2.8-3.1%。我们报告了这些限制,以及神谕、阴性对照、多种子和按机器人的分析。结果表明,显式粒子动力学是机器人视频预测的一种有前景的低维接口,而鲁棒的语言接地和外观交付仍是主要的开放挑战。

英文摘要

Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle identities, rolls their states forward, and is modulated by a frozen OpenCLIP instruction embedding through FiLM. A causal dual-stream decoder combines particle-rendered motion with appearance encoded exclusively from the last observed frame; a residual refiner and learned delivery mask produce five future frames without access to future appearance. On our VRS benchmark constructed from diverse real-robot trajectories, particle dynamics reduce trajectory error by 21.0\% over persistence. Across three delivery-mask seeds, AcrossVAM1.0 improves future-frame PSNR/SSIM from 19.97/0.796 to 20.573/0.8004, while raw particle generation improves motion-region PSNR from 11.89 to 13.23. The delivered model does not yet beat persistence in LPIPS, and correct-versus- shuffled language changes trajectory error by only 2.8--3.1%. We report these limitations alongside oracle, negative-control, multi-seed, and per-robot analyses. The results show that explicit particle dynamics are a promising low-dimensional interface for robot video prediction, while robust language grounding and appearance delivery remain the principal open challenges.

发表机构

  • Across Physical AI(跨越物理人工智能公司)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑