arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

力,永不嫌迟:通过反应式力注入加速视觉-语言-动作模型的训练后优化

Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection

Yi Wang, Wendi Chen, Zimo Wen, Han Xue, Xueqi Li, Wenye Yu, Zhijie Chen, Hao Yang, Jun Lv, Chuan Wen, Cewu Lu

arXiv 2607.14236首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Innovation Institute; Southern University of Science and Technology; Noematrix Ltd.(上海交通大学; 上海创新研究院; 南方科技大学; 诺玛矩阵有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对预训练VLA策略在接触状态下表现不佳的问题,提出LIFT框架,通过嫁接反应式动作专家、注入力及结合在线DAgger循环,提升其在富含接触操作任务中的性能,且证明相关组件对稳健操作很重要。

AI 中文摘要

预训练的视觉-语言-动作(VLA)策略提供了强大的语言条件操作知识,但在操作进入接触状态时,如场景被遮挡、深度模糊或小的力误差使执行偏离离线演示分布时,仍主要受视觉驱动且表现不佳。我们提出了LIFT(用于VLA训练后优化的后期反应式力注入),这是一个力感知训练后框架,在保留预训练VLA策略的一般操作知识的同时,为其添加接触反应性。LIFT在原始动作专家旁边嫁接一个反应式动作专家,从预训练动作权重初始化,并通过因果力记忆和零初始化交叉注意力注入最近的6D末端执行器力,使动作在执行期间得以刷新。为解决接触反馈的策略依赖分布转移问题,LIFT进一步将反应式力注入与在线DAgger循环相结合,该循环在离线任务对齐数据和人工校正的在线展开的混合数据上进行训练。在毛巾折叠、书本插入和河内环放置任务中,LIFT比仅基于视觉的训练后优化学习更快且性能更高,同时消融实验表明反应式力记忆和在线校正数据对于强大的富含接触的操作都很重要。我们的代码和数据将公开可用。

英文摘要

Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is occluded, depth is ambiguous, or small force errors push execution off the offline demonstration distribution. We present LIFT (Late Reactive Injection of Force for VLA Post-Training), a force-aware post-training framework that adds contact reactivity to a pretrained VLA policy while preserving its general manipulation knowledge. LIFT grafts a reactive action expert beside the original action expert, initializes it from pretrained action weights, and injects recent 6D end-effector force through causal force memory and zero-initialized cross attention, enabling actions to be refreshed during execution. To address the policy-dependent distribution shift of contact feedback, LIFT further couples reactive force injection with an online DAgger loop that trains on a mixture of offline task-alignment data and human-corrected online rollouts. Across towel folding, book insertion, and Hanoi ring placement, LIFT learns faster and reaches higher performance than vision-only post-training, while ablations show that reactive force memory and online corrective data are both important for robust contact-rich manipulation. Our code is publicly available at https://github.com/y-wng/lift.

CommentsAccepted to CoRL 2026.Project page: https://lift-policy.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑