通过延迟感知引导融合实现异步多模态扩散策略组合
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
- Shanghai Jiao Tong University(上海交通大学)
- Noematrix(无矩阵)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对多模态扩散策略在模态差异下的融合问题,提出LAG - Fusion框架,通过延迟感知引导融合实现异步策略组合,推导参考帧重定位规则,在接触丰富操纵实验中提升了策略响应性和任务性能。
AI中文摘要:
扩散策略在机器人模仿学习中显示出强大潜力,近期扩展引入其他模态以提升操纵性能。但模态在信息内容、传感速率和推理延迟方面存在差异。现有多模态扩散策略通常依赖同步融合或手动设计的多频率架构。我们提出LAG - Fusion,一种用于异步多模态扩散策略组合的延迟感知引导融合框架。它允许特定模态策略以其原生推理速率运行并在可用时提供去噪引导。为使异步组合一致,我们推导了相对动作表示下扩散变量的参考帧重定位规则。通过将低频视觉策略与高频力策略组合在接触丰富的操纵中实例化LAG - Fusion。实验表明,LAG - Fusion在异构模态延迟下比同步融合和专门设计的力感知基线提高了策略响应性和任务性能。
英文摘要:
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to improve manipulation performance. However, these modalities often differ not only in information content but also in sensing rates and inference latencies. Existing multimodal diffusion policies typically rely on synchronous fusion or manually designed multi-frequency architectures, which either slow down high-frequency feedback or limit extensibility to new modality combinations. We propose LAG-Fusion, a latency-aware guidance fusion framework for asynchronous multimodal diffusion policy composition. LAG-Fusion allows modality-specific policies to operate at their native inference rates and contribute denoising guidance whenever available. To make asynchronous composition consistent, we derive a reference-frame rebasing rule for diffusion variables under relative action representations, enabling delayed guidance to be aligned before fusion. We instantiate LAG-Fusion in contact-rich manipulation by composing a low-frequency vision policy with a high-frequency force policy. Experiments under heterogeneous modality latencies show that LAG-Fusion improves policy responsiveness and task performance over synchronous fusion and specially designed force-aware baselines.