arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦视觉、语言与动作以实现高效的多任务机器人策略

Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

Xiatao Sun, Chen Liang, Ziyao Zeng, Qian Wang, Haoyang Zhang, Yue Sun, Qiucheng Li, Daniel Rakita

arXiv 2609.18374首次发表:更新:

发表机构

Yale University; Peking University; Digients(耶鲁大学; 北京大学; Digients)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型高计算成本问题,提出解耦具身模型DEM,用独立视觉和语言编码器加MeanFlow动作头,在模拟和真实任务中达到相当成功率,同时推理更快、能耗更低。

AI 中文摘要

视觉-语言-动作(VLA)模型将动作模块附加到具有数十亿参数的视觉-语言模型(VLM)上,并在每个控制步骤中为该骨干网络付出计算代价。对于低级操作策略而言,这种代价可能是不必要的:VLM提供视觉和语言嵌入,而近年来的独立视觉编码器和仅编码器语言模型在视觉嵌入和语言理解基准上现已达到或超过大型VLM。我们通过一项受控实验研究这一问题。在保持演示数据、训练预算、任务和测量平台固定的情况下,我们改变解耦策略的视觉编码器、语言编码器和动作头,并与七个VLA基线进行比较。该研究产生了解耦具身模型(DEM),它将微调的DINOv3编码器和冻结的NeoBERT编码器与MeanFlow头配对,该头在单次前向传播中生成每个动作块。在18个模拟操作任务(包含保留的语言改写和随机化场景)以及三个真实机器人任务上,DEM在我们的评估协议下实现了与最先进的VLM骨干策略相当的成功率,同时其推理频率是后者的八到十七倍,每次推理的能耗低六到十五倍。在此任务范围内,现代解耦组件提供了更好的成功率-延迟-能耗权衡。

英文摘要

Vision-Language-Action (VLA) policies commonly run Vision-Language Model (VLM) backbones with billions of parameters at every policy inference, which costs latency and energy. We revisit a decoupled alternative for multi-task manipulation: separate vision and language encoders whose representations condition a compact action head. We run a standardized comparison that varies the vision encoder, the language encoder, and the action head while holding the demonstrations, the training-step budget, the tasks, the evaluation protocol, and the measurement platform fixed, against seven VLA baselines. The resulting Decoupled Embodiment Model (DEM) combines a fine-tuned DINOv3 vision encoder, a frozen NeoBERT language encoder, and a MeanFlow head that generates an action chunk in one forward pass. On 18 RoboCasa tasks evaluated with held-out instruction paraphrases and randomized scenes, DEM reaches 55.6\% mean success against 56.9\% for GR00T N1.7 and 54.6\% for $π_{0.5}$, and on three real-robot tasks it reaches 66.0\% against 68.0\% for GR00T N1.7. On the same workstation, DEM needs 6.1\,ms per policy forward pass, a maximum throughput of 162.7 policy calls per second, and draws an estimated 2.07\,J of GPU energy per call, eight to seventeen times the throughput and six to fifteen times less energy than these VLM-backbone policies. Within this trained-task regime, DEM sits on the observed success--latency--energy frontier and provides a strong, efficient baseline for language-conditioned robot skills.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑