Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
视觉-语言-动作模型的流匹配策略的强化微调
机构 * Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences, Beijing, China(脑启发认知人工智能实验室,自动化研究所,中国科学院,北京,中国) ; Beijing Institute of AI Safety and Governance, China(北京人工智能安全与治理研究院,中国) ; State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术国家重点实验室) ; Beijing Key Laboratory of Safe AI and Superalignment, China(北京安全人工智能与超对齐重点实验室,中国) ; University of Chinese Academy of Sciences (UCAS), Beijing, China(中国科学院大学(UCAS),北京,中国) ; Long-term AI,Beijing,China(长期人工智能,北京,中国)
AI总结 针对流匹配模型强化微调中重要性采样计算困难的问题,提出流策略优化算法,通过条件流匹配目标、结构感知信用分配等技术实现稳定在线微调,在LIBERO和ALOHA任务上超越基线。
Comments Accepted to ICRA 2026