频率条件流匹配用于视觉-语言-动作模型
Frequency-Conditioned Flow Matching for Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
FreqFM提出频率条件流匹配框架,将动作频率作为显式条件维度,在DCT坐标中构建频谱匹配源分布并平衡目标,无需改变VLA主干,在LIBERO等基准及真实机器人任务上显著提升性能。
中文摘要 AI 辅助
机器人动作是时间上相关的轨迹,其频率分量以高度非均匀的能量分布编码不同尺度的运动。然而,基于流匹配的视觉-语言-动作(VLA)模型通常在时间坐标中生成动作,没有显式建模或系统利用这种频率异质性。我们提出了FreqFM,一种用于VLA模型的频率条件流匹配框架。它将动作频率从隐式的轨迹属性提升为贯穿整个生成流程的显式条件维度。具体而言,在DCT频率坐标中,FreqFM构建了频谱匹配的源分布,跨频率自适应平衡目标,并使用相应的参考传输尺度约束每个频率的引导残差。FreqFM无需改变VLA主干即可集成到现有流匹配动作专家中。在LIBERO、LIBERO-Plus和VLA-Arena上,FreqFM持续提升性能,包括在LIBERO-Plus上获得9.3个百分点的提升,并在六个真实机器人任务上进一步证明了其有效性。
英文摘要
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.
发表机构
- AGIBOT
机构由 AI 辅助整理,请以论文原文为准。