TempoBridge:面向视觉-语言-动作策略的语言引导速度控制
TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
浏览论文内容
中文总结 AI 辅助
TempoBridge是一个轻量级框架,利用冻结的VLA表示和因果阶段路由器,根据语言速度提示调节动作,无需额外演示或微调,将LIBERO速度成功率从52.6%提升至89.7%,并泛化到未见速度表达。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型能有效理解要执行的任务,但对执行方式(如快速或慢速移动)的控制能力有限。我们提出TempoBridge,一个轻量级框架,利用冻结的VLA表示,根据指令中每个任务阶段的速度提示来调节动作,无需额外的速度条件机器人演示或速度特定的基础策略微调。TempoBridge从上下文VLM表示中提取速度提示,通过因果阶段路由器将其与任务进度对齐,并在执行过程中调节名义运动指令。在LIBERO任务中,TempoBridge在规范速度指令下将速度成功率从52.6%提升至89.7%,同时保持较高的任务成功率。在没有速度提示时,它也保持接近基线的性能,并能无需额外训练泛化到未见过的速度表达。在物理机器人上的实验进一步展示了真实世界操作中语言条件速度调节的有效性。
英文摘要
Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly. We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. Across LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under canonical tempo instructions while retaining high task success. It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training. Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.
发表机构
- Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院(KAIST))
机构由 AI 辅助整理,请以论文原文为准。