发表机构
The University of Texas at Austin; University of Central Florida(德克萨斯大学奥斯汀分校; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VLA-Feedback双时间尺度架构,将扩散规划与高频视觉反馈结合,实现实时动作修正,在动态任务中成功率大幅提升。
AI 中文摘要
视觉-语言-动作(VLA)模型通过将预训练视觉-语言模型的语义知识与富有表现力的动作生成策略相结合,在机器人操作中展现出强大的泛化能力。基于扩散的动作生成器在建模时间上连贯的动作块方面尤为有效,但这些动作块通常在推理后以开环方式执行。当物体移动、接触变化或场景在执行过程中演变时,这限制了响应性。我们提出了VLA-Feedback,一种双时间尺度架构,将低频扩散规划与高频视觉反馈相结合。VLA-Feedback并非在执行前对动作块进行完全去噪,而是将其最终去噪步骤保留为轻量级反馈接口,使得每个动作在执行前都能利用最新观测进行修正。这种设计保留了扩散规划器的表现力,同时无需重新运行完整的视觉-语言扩散模型即可实现实时动作修正。VLA-Feedback在静态LIBERO任务上与GR00T表现相当,同时将动态模拟任务的平均成功率从27.5%提升至85.0%。在真实机器人任务中,平均成功率从51%提升至73%。更多材料可在我们的项目页面找到:此https URL。
英文摘要
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but these chunks are typically executed open-loop after inference. This limits responsiveness when objects move, contacts change, or the scene evolves during execution. We propose VLA-Feedback, a two-timescale architecture that combines low-frequency diffusion planning with high-frequency visual feedback. Rather than fully denoising an action chunk before execution, VLA-Feedback retains its final denoising step as a lightweight feedback interface, allowing each action to be corrected using the latest observation before it is executed. This design preserves the expressiveness of the diffusion planner while enabling real-time action correction without rerunning the full vision-language diffusion model. VLA-Feedback matched GR00T on static LIBERO tasks while improving average success on dynamic simulation tasks from 27.5% to 85.0%. On real-robot tasks, it improved average success from 51% to 73%. Additional materials can be found on our project page: https://vla-feedback.github.io.
Comments10th Conference on Robot Learning (CoRL 2026), Austin TX, USA