教导强化学习智能体更好地行动:VLM作为在线强化学习的行动顾问
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出VARL框架,利用视觉-语言模型为在线强化学习智能体提供行动建议,在保证最优性和收敛性的同时提升样本效率,尤其在稀疏奖励任务中表现显著,且计算开销低。
AI中文摘要:
在复杂任务中的在线强化学习是耗时的,因为学习最优Q函数需要大量的交互步骤。视觉-语言行动(VLA)策略是解决多样化任务的一个有前景的方向;然而,它们在低级控制上的性能仍然有限,且有效部署通常需要特定于任务的专家演示进行微调。在本文中,我们提出了VARL(VLM作为在线强化学习的行动顾问),这是一个利用视觉-语言模型(VLM)的领域知识为强化学习智能体提供行动建议的框架。与先前方法不同,VARL提供行动建议而非设计启发式奖励,从而保证最优性和收敛性不变。建议的行动增加了样本多样性,并最终提高了样本效率,尤其是在稀疏奖励任务中。为了验证VARL的有效性,我们在多种环境和智能体设置下对其进行了评估。结果表明,VARL大幅提高了样本效率,且未引入显著的计算开销。这些优势使VARL成为在线强化学习的通用框架,并使得在真实环境中直接从零开始应用强化学习成为可能。
英文摘要:
Online reinforcement learning in complex tasks is time-consuming, as massive interaction steps are needed to learn the optimal Q-function.Vision-language action (VLA) policies represent a promising direction for solving diverse tasks; however, their performance on low-level control remains limited, and effective deployment often requires task-specific expert demonstrations for fine-tuning. In this paper, we propose \textbf{VARL} (\textbf{V}LM as \textbf{A}ction advisor for online \textbf{R}einforcement \textbf{L}earning), a framework that leverages the domain knowledge of vision-language models (VLMs) to provide action suggestions for reinforcement learning agents. Unlike previous methods, VARL provides action suggestions rather than designing heuristic rewards, thereby guaranteeing unchanged optimality and convergence. The suggested actions increase sample diversity and ultimately improve sample efficiency, especially in sparse-reward tasks. To validate the effectiveness of VARL, we evaluate it across diverse environments and agent settings. Results show that VARL greatly improves sample efficiency without introducing significant computational overhead. These advantages make VARL a general framework for online reinforcement learning and make it feasible to directly apply reinforcement learning from scratch in real-world environments.