arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HABILIS Brain 0:视觉-语言-动作的几何变化监督与残差流恢复

HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery

Jinu Pahk, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Byoung-Tak Zhang

arXiv 2609.25558首次发表:更新:

发表机构

Tommoro Robotics(Tommoro Robotics)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出GC-VLA,通过预测多视角几何变化标记实现视觉-语言-动作策略的几何监督,并引入GCRF残差流,在LIBERO上达到99.55%成功率。

AI 中文摘要

视觉-语言-动作策略受益于几何监督,但仅当前帧的几何信息无法明确描述与操作相关的变化。该设计旨在学习一种与具身无关的视觉接口,该接口可在机器人视频和第一人称视频上进行预训练,之后再进行特定于机器人的动作对齐。我们提出了几何变化视觉-语言-动作模型(GC-VLA),该模型从当前观测中学习预测多视角的未来-当前几何变化标记。离线帧对定义了标称的0.5秒预测范围;未来观测仅用于构建训练目标。阶段1训练一个几何变化视觉-语言模型(GC-VLM)。阶段2引入了一个连续的ActionExpert,并在视觉-语言模型接口处停止动作流梯度,同时将其与机器人动作对齐。阶段3允许这些梯度与ActionExpert共同更新可训练的视觉-语言模型组件。阶段4冻结GC-VLA,并应用几何条件残差流(GCRF),使用二元干预路由器和从闭环反馈中学习的单一有界残差速度策略。GC-VLA在LIBERO上达到了95.20%的成功率,而带有GCRF的GC-VLA达到了99.55%。推理时仅使用当前观测和学习的GC表示,不执行离线的目标编码器。

英文摘要

Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associated with manipulation. This design is motivated by the goal of learning an embodiment-agnostic visual interface that can be pretrained across robot and egocentric video before robot-specific action alignment. We introduce Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations. Offline frame pairs define a nominal 0.5-second prediction horizon; future observations are used only to construct training targets. Stage 1 trains a geometry-change vision-language model (GC-VLM). Stage 2 introduces a continuous ActionExpert and aligns it with robot actions while stopping action-flow gradients at the VLM interface. Stage 3 enables these gradients to update the trainable VLM components jointly with the ActionExpert. Stage 4 freezes GC-VLA and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback. GC-VLA achieves 95.20% success on LIBERO, and GC-VLA with GCRF achieves 99.55%. Inference uses current observations and the learned GC representation without executing the offline target encoders.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑