arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21228cs.ROcs.AIcs.CV

FOCAL-VLA:面向视觉-语言-动作模型的子任务引导几何蒸馏与隐式世界建模

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

Zhiyuan Gao, Di Wen, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Schäfer, Kunyu Peng, Michael Beetz

首次发表
浏览论文内容

中文总结 AI 辅助

FOCAL-VLA通过子任务引导的几何蒸馏和隐式世界建模,增强VLA模型的空间与时间理解,在仿真和真实操作任务中超越基线。

中文摘要 AI 辅助

基于预训练视觉-语言模型构建的视觉-语言-动作(VLA)模型在多种机器人操作任务中展现出强大的性能。然而,直接将当前2D观测映射到动作的VLA模型往往缺乏足够的空间和时间理解,限制了其在精确和长时程操作中的表现。近期方法通过几何监督和整个场景的未来状态预测来增强VLA模型,但这些方法可能受到冗余场景信息的干扰,使模型难以学习与当前交互相关的几何和动态。为解决此问题,我们提出FOCAL-VLA,一个结合子任务引导的几何蒸馏与隐式世界建模的框架,以学习当前空间结构和未来交互动态的表征。为使几何学习聚焦于当前子任务,我们通过将几何潜在变量与子任务相关图像区域的特征对齐,将几何知识从VGGT迁移到VLA模型。为捕捉当前交互的未来3D演化,我们利用当前和未来演示帧的Track4World特征进行隐式世界建模。这两种互补的表征共同指导动作生成,而无需在推理时运行VGGT或Track4World。实验表明,FOCAL-VLA在仿真基准和真实世界操作任务上均优于基线方法。项目网站:此https URL。

英文摘要

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: https://zhiyuan-gao.github.io/FOCAL-VLA/.

发表机构

  • University of Bremen(不来梅大学)
  • Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院)
  • Robotics Institute Germany (RIG)(德国机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑