在并行仿真中扩展视觉语言模型奖励学习用于机器人操作
Scaling Vision-Language Reward Learning for Robot Manipulation in Parallel Simulation
浏览论文内容
中文总结 AI 辅助
提出RAPID系统,通过GPU并行回放、单请求偏好标注和自适应更新,在IsaacLab中实现机器人操作奖励学习8倍加速和95.5%API节省,成功率提升至98.7%。
中文摘要 AI 辅助
视觉语言模型(VLMs)可以替代人类标注者进行基于偏好的奖励学习,但顺序API请求和单环境数据收集使得训练缓慢且成本高昂。我们提出了RAPID(具有自适应并行图像多样性的奖励学习),该系统将GPU并行回放与数据感知策略更新、单请求偏好标注、自动奖励稳定和代表性图像采样相结合。我们在IsaacLab中的五个Franka Panda操作任务上评估了这些组件。并行回放和自适应更新首次大幅缩短了训练时间:在匹配的两阶段提示下,平均运行时间从9.18小时降至3.13小时。启用所有RAPID组件后,训练在1.15小时内完成,每次运行使用896次而非19,840次API调用,总体最终成功率从86.3%提升至98.7%。这代表了8.0倍的端到端加速和95.5%的API使用量减少。使用Gemma 3 12B和GPT-4.1 mini进行的离线评估表明,单请求提示在两种模型上均降低了标注延迟和成本。代码可在以下网址获取:此HTTPS URL。
英文摘要
Vision-language models (VLMs) can replace human annotators in preference-based reward learning, but sequential API requests and single-environment data collection make training slow and costly. We present RAPID (Reward learning with Adaptive Parallel Image Diversity), a system that couples GPU-parallel rollout with data-aware policy updates, single-request preference labeling, automatic reward stabilization, and representative image sampling. We evaluate these components on five Franka Panda manipulation tasks in IsaacLab. Parallel rollout and adaptive updates provide the first substantial reduction in training time: under matched two-stage prompting, mean runtime falls from 9.18 to 3.13 hours. With all RAPID components enabled, training completes in 1.15 hours using 896 rather than 19,840 API calls per run, and aggregate final success rises from 86.3\% to 98.7\%. This represents an 8.0$\times$ end-to-end speedup and a 95.5\% reduction in API usage. An offline evaluation with Gemma~3 12B and GPT-4.1 mini demonstrates that single-request prompting reduces labeling latency and cost across both models. Code is available at: https://github.com/rapid-vlm/rapid-vlm-rl.
发表机构
- KU Leuven(鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。