arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时切换:视觉-语言-动作模型的可靠动作块扩展

When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek, Byung-kwan Lee, Suha Kwak, Jeany Son

arXiv 2610.05719首次发表:更新:

发表机构

KAIST; GIST; NVIDIA; POSTECH(韩国科学技术院; 光州科学技术院; 英伟达; 浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型推理昂贵导致机器人走走停停的问题,提出RACE框架,通过预测子技能转换时机并条件化动作生成,实现更长动作块的可靠执行,在模拟和真实机器人上均提升成功率并减少空闲时间。

AI 中文摘要

视觉-语言-动作(VLA)模型作为机器人操作中的统一策略,但其昂贵的推理迫使机器人在策略调用之间暂停,导致走走停停的执行,中断了平滑运动并延长了任务完成时间。扩展动作块减少了策略调用次数,从而减少了这些暂停,但预测更远的未来使得长块执行变得不可靠。为了理解这种不可靠性的来源,我们分析了长块内的动作误差,发现它们集中在操作子技能之间的转换处,并随着块长度的增加而急剧增长。这表明转换时机的重要性,即块内何时切换子技能。受此观察启发,我们引入了RACE(可靠动作块扩展),一个从辅助的一步去噪过程中预测转换时机并据此条件化动作生成的框架。通过学习和条件化转换时机,RACE减少了子技能转换处的误差,并实现了更长动作块的可靠执行。在模拟基准测试中,RACE在相同块长度下优于微调;在2倍更长的块下,它在成功率上超越了最近的最先进和高效的VLA;在4倍更长的块下,它仍保持竞争力。在真实机器人上,RACE使用4倍更长的块,将走走停停执行造成的空闲时间减少了约5倍,同时实现了比相同块长度微调更高的成功率。代码和真实机器人演示可在https URL获取。

英文摘要

Vision-Language-Action (VLA) models serve as unified policies for robotic manipulation, yet their expensive inference forces robots to pause between policy calls, resulting in stop-and-go execution that interrupts smooth motion and prolongs task completion. Extending the action chunk reduces policy calls and hence these pauses, but predicting farther into the future makes long-chunk execution unreliable. To understand where this unreliability arises, we analyze action errors within long chunks and find that they concentrate around transitions between manipulation subskills, growing sharply with chunk length. This suggests the importance of transition timing, i.e., when to switch subskills within a chunk. Motivated by this observation, we introduce RACE (Reliable Action-Chunk Extension), a framework that predicts the transition timing from an auxiliary one-step denoising pass and conditions action generation on it. By learning and conditioning on transition timing, RACE reduces errors at subskill transitions and enables reliable execution of longer action chunks. Across simulation benchmarks, RACE outperforms fine-tuning at the same chunk length; with 2x longer chunks, it surpasses recent state-of-the-art and efficient VLAs in success rate, and with 4x longer chunks, it remains competitive. On a real robot, RACE uses 4x longer chunks, which reduces the idle time caused by stop-and-go execution by about 5x, while achieving a higher success rate than fine-tuning with the same chunk length. Code and a real-robot demo are available at https://github.com/Seonghoon-Yu/RACE-VLA

CommentsPre-print

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑