arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLIFT:通过非侵入式闭环迭代微调将Gemini机器人端侧模型转化为人形机器人专家

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka, Peng Xu, Jinyu Xie, Thomas Tian

arXiv 2607.29172首次发表:更新:

发表机构

University of California, Berkeley; Google DeepMind; NVIDIA Research(加州大学伯克利分校; 谷歌DeepMind; 英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对闭源机器人模型的托管式SFT API范式,提出CLIFT方法,在不访问模型内部机制的情况下,将Gemini机器人端侧模型经两次飞轮循环提升至接近完美的敏捷接触任务成功率。

AI 中文摘要

尽管机器人基础模型的能力不断提升,但最强的模型通常基于专有数据训练且保持闭源,限制了下游用户将其适配到新任务、实体和部署场景的能力。参照大语言模型社区,针对闭源权重机器人基础模型的新兴访问范式是托管式监督微调(SFT)API,用户提交训练数据后可获得微调策略,但无法访问模型权重、梯度或训练内部机制。这类API让下游用户能利用强大的专有基础模型,却将策略改进限制为纯模仿,排除了依赖内部训练信号的强化学习及其他闭环方法。这一限制在敏捷、接触丰富的人形机器人操作中尤为突出,由于新状态、动作跟踪动态、延迟和控制器特定故障模式,策略输出与部署行为的差距较大。我们研究该托管API范式对人形机器人适配的有效性,以及如何在其中实现闭环改进以推动策略达到任务精通。我们在真实人形机器人上开展首批托管API适配的实证研究,实例化为Gemini机器人端侧模型(GROD)。我们发现,通过API直接进行SFT的性能显著优于在相同演示数据上训练的领先开源视觉语言动作模型(VLA),但在敏捷、接触丰富的任务上仍未达到部署级精通。为缩小这一差距,我们提出CLIFT:闭环迭代微调,它将部署时的奖励反馈转化为API兼容的监督数据,无需访问权重、梯度、似然或损失即可实现闭环策略改进——在两次飞轮循环后,推动GROD达到接近完美的成功率,全程无需“打开模型黑盒”。

英文摘要

While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradients, or training internals. While such APIs let downstream users leverage powerful proprietary foundation models, they restrict policy improvement to pure imitation, ruling out reinforcement learning and other closed-loop methods that rely on internal training signals. This limitation is particularly acute for agile, contact-rich humanoid manipulation, where the gap between policy outputs and deployed behavior is large due to novel states, action tracking dynamics, latency, and controller-specific failure modes. We study how effective this managed-API regime is for humanoid adaptation, and how closed-loop improvement can be realized within it to push policies toward task mastery. We conduct one of the first empirical studies of managed-API adaptation on a real humanoid, instantiated on Gemini Robotics On-Device (GROD). We find that direct SFT through the API substantially outperforms a leading open-weight VLA trained on the same demonstrations, yet still falls short of deployment-level mastery on agile, contact-rich tasks. To close this gap, we introduce CLIFT: Closed-Loop Iterative Fine-Tuning, which turns deployment-time reward feedback into API-compatible supervised data and enables closed-loop policy improvement without accessing weights, gradients, likelihoods, or losses-pushing GROD to near-perfect success after two flywheel cycles, all without "opening the model box."

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑