StructRL:面向基于流的视觉-语言-动作模型的结构化动作空间探索
StructRL: Structured Action-Space Exploration for Flow-Based VLAs
浏览论文内容
中文总结 AI 辅助
StructRL通过将策略随机性重定位至动作空间的三个耦合设计,解决基于流的VLA模型的结构化噪声稀释问题,在模拟与真实任务上提升了探索效率及OOD性能。
中文摘要 AI 辅助
基于流的视觉-语言-动作(Vision-Language-Action, VLA)模型目前被广泛应用于连续机器人操纵任务,而在线强化学习(Reinforcement Learning, RL)正成为使这类模型适配新任务的关键技术。现有RL方法通常在去噪链内部注入随机性,且多采用各向同性或时间独立的噪声。然而,有效的机器人探索需要结构化噪声:这类噪声需具备时间平滑性,且在不同动作组上的缩放比例不同。我们发现,仅将链内噪声切换为结构化形式并不足够:在中间流时间注入的噪声会在执行前的剩余去噪步骤中被削弱,我们将这一现象称为「结构化噪声稀释」。为此,我们提出StructRL,该方法通过三个耦合选择将策略随机性重定位至动作空间,从而避免噪声稀释:(i)确定性常微分方程(Ordinary Differential Equation, ODE)解码器;(ii)直接注入动作空间的结构化噪声;(iii)最后一步重放,即策略梯度更新避免为中间去噪状态分配似然。这一设计使结构化探索与执行动作绑定,同时为流解码器提供易处理的训练信号。在多个模拟操纵基准任务和两个真实世界任务上,针对三种基于流的VLA模型的实验显示,StructRL相比先前的链内基线方法,提升了探索效率与分布外(Out-of-Distribution, OOD)性能,证明了结构化动作空间探索在通过RL适配基于流的VLA模型方面的有效性。项目页面:this https URL
英文摘要
Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for adapting them to new tasks. Existing RL methods typically inject stochasticity inside the denoising chain, often through isotropic or temporally independent noise. However, effective robot exploration calls for structured noise: temporally smooth and scaled differently across action groups. We show that simply switching the in-chain noise to a structured form does not suffice: noise added at an intermediate flow time can be weakened by the remaining denoising steps before execution, a phenomenon we call \emph{Structured Noise Dilution}. We propose \textbf{StructRL}, which avoids dilution by relocating policy stochasticity to the action space via three coupled choices: (i) a deterministic ODE decoder, (ii) structured noise injected directly in the action space, and (iii) last-step replay, where policy-gradient updates avoid assigning likelihoods to intermediate denoising states. This keeps structured exploration tied to the executed action while providing a tractable training signal for the flow decoder. Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL. \textbf{Project page:} https://flyfaerss.github.io/structrl/
发表机构
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Fudan University(复旦大学)
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。