发表机构
Eastern Institute of Technology; Jiangnan University; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology (Guangzhou); National University of Singapore(东方理工大学; 江南大学; 上海交通大学; 香港科技大学(广州); 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ForceRFT提出力引导残差强化学习框架,利用腕部反馈和人类监督优化VLA策略,在插头插入等任务中提升自主成功率。
AI 中文摘要
力条件化的视觉-语言-动作(VLA)策略能够响应接触,但当仅基于示范进行训练时,其恢复行为可能受限于示范覆盖范围,并且无法从部署结果中学习。人类纠正性模仿提供了额外的恢复示例,但其目标匹配局部动作目标,而未显式优化任务回报。我们提出了ForceRFT,一种力引导的残差强化学习框架,从人类监督和自主任务结果中学习接触相关的修正。一个冻结的、基于示范训练的SmolVLA先验生成力条件化的动作块,而一个轻量级残差actor利用在动作块执行期间收集的腕部反馈来细化单个末端执行器位姿指令。决策时的腕部力矩、其时间变化以及选定的基础运动条件同时调节残差修正和值估计。人类修正监督残差actor,而经过验证的自主转换训练双critic,并支持对同一actor进行值引导更新。自举仅限于自主片段,防止TD信用跨越人类干预边界。在插头插入、环上钉装配和白板擦拭任务上的真实机器人实验显示,其自主成功率高于所评估的基于示范训练和残差模仿的基线。与残差模仿的比较支持值引导的残差优化,而插头插入的消融实验表明直接使用执行时腕部反馈的益处。
英文摘要
Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective matches local action targets without explicitly optimizing task return. We present ForceRFT, a force-guided residual reinforcement learning framework that learns contact-dependent corrections from human supervision and autonomous task outcomes. A frozen, demonstration-trained SmolVLA-based prior generates force-conditioned action chunks, while a lightweight residual actor refines individual end-effector pose commands using wrist feedback acquired during chunk execution. The decision-time wrist wrench, its temporal change, and the selected base motion condition both residual correction and value estimation. Human corrections supervise the residual actor, while verified autonomous transitions train the twin critics and support value-guided updates to the same actor. Bootstrapping is restricted to autonomous segments, preventing TD credit from crossing human-intervention boundaries. Real-robot experiments on plug insertion, ring-on-peg assembly, and whiteboard wiping show higher autonomous success rates than the evaluated demonstration-trained and residual-imitation baselines. Comparisons with residual imitation support value-guided residual optimization, while plug-insertion ablations indicate the benefit of direct execution-time wrist feedback.
Comments8 pages, 4 figures, 3 tables