发表机构
Massachusetts Institute of Technology; MIT-IBM Computing Research Lab; Improbable AI Lab(麻省理工学院; 麻省理工学院-IBM计算研究实验室; 英普罗巴布尔人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对强化学习中探索难的问题,提出提示驱动的探索策略(PDE),利用视觉-语言模型对轨迹视频推理,从弱策略轨迹中优化提示,实现后验采样,能让强化学习从零奖励起步学习成功策略并提升样本效率。
AI 中文摘要
探索对于强化学习至关重要,因为策略无法通过重复采样其偏好行为来改进。标准方法在动作空间中注入随机性,但这种抖动只能产生接近原始的轨迹。摆脱弱策略通常需要全局扰动,而动作噪声无法产生。大语言模型和视觉-语言-动作模型提供了一条途径:它们根据自然语言提示来调整策略,修改提示会引发全局变化。挑战在于找到能引发有用全局变化的提示。对于很少成功的弱策略,奖励过于稀疏难以选择。我们的想法是从轨迹本身优化提示:视觉-语言模型对轨迹视频进行推理,诊断策略的响应并重写提示以引出更好的行为。此过程在提示层面实现了后验采样,这是一个经典的强化学习探索框架。我们将此策略称为提示驱动的探索(PDE)。在操纵和推理任务中,PDE使强化学习即使从零奖励开始也能学习到成功的策略,并更广泛地提高样本效率。
英文摘要
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt to elicit better behavior next time. This procedure resembles posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintains an implicit distribution over useful prompts and updates it from observed rollouts. We call this strategy Prompt-Driven Exploration (PDE). Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at https://xinyunsunshine.github.io/prompt-rl.