arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemCorr-DP:由参考引导的扩散策略的反事实对应条件约束

MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference

Tan Su, Haoxiang Yang, Ruxin Wang, Binghui Xie

arXiv 2609.06615首次发表:更新:

发表机构

Southern University of Science and Technology; Sun Yat-sen University(南方科技大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对物体位置与视角同时变化时行为克隆策略失效的问题,提出MemCorr-DP扩散策略,利用显式3D参考关系与反事实配对目标,在组合偏移下实现96.67%的闭环成功率。

AI 中文摘要

行为克隆的视觉运动策略在其训练分布附近可以保持准确性,但当物体位置和相机视角同时改变时,这些策略可能会失效。一条成功的参考轨迹包含了迁移相同交互所需的几何信息,但策略必须将该几何信息与当前场景对齐,并在去噪过程中对其保持敏感。为了应对这些挑战,我们提出了MemCorr-DP,这是一种扩散策略,它将冻结的RoMa v2匹配提升为当前场景与参考轨迹之间的显式3D关系。一个反事实配对目标为相反的(两种)行为分配相同的物理状态和带噪声的动作,同时保留特定于参考的去噪目标。随后,混合条件微调使策略从真实几何信息适应到测量到的对应误差。我们最强的评估将门放置在训练支持范围之外的最外侧位置带,并将查询相机改变±15°。在这种组合偏移下,MemCorr-DP实现了96.67%的闭环成功率,而具有相同动作架构的视觉Transformer的成功率为88.00%。目标消融和参考干预实验表明,行为会响应所选的参考,而匹配的对照组则更倾向于完整的(3D)关系集,而非仅基于未来运动或质心几何。这些结果支持将显式3D参考关系作为一种鲁棒的条件接口,用于在评估任务中空间和视角变化叠加的场景。

英文摘要

Behavior-cloned visuomotor policies can remain accurate near their training distribution yet fail when object position and camera viewpoint change together. A successful reference trajectory contains the geometry needed to transfer the same interaction, but the policy must align that geometry with the current scene and remain sensitive to it during denoising. To address these challenges, we present MemCorr-DP, a diffusion policy that lifts frozen RoMa v2 matches into explicit 3D relations between the current scene and the reference trajectory. A counterfactual paired objective assigns opposite behaviors the same physical state and noisy action while retaining reference-specific denoising targets. Mixed-condition fine-tuning then adapts the policy from ground-truth geometry to measured correspondence errors. Our strongest evaluation places the Door in the outermost position bands beyond the training support and changes the query camera by $\pm15^\circ$. Under this combined shift, MemCorr-DP achieves 96.67% closed-loop success, compared with 88.00% for a visual Transformer with the same action architecture. Objective ablations and reference interventions show that behavior responds to the selected reference, while matched controls favor the complete relation set over future motion or centroid geometry alone. These results support explicit 3D reference relations as a robust conditioning interface when spatial and viewpoint changes are compounded in the evaluated task.

Comments12 pages, 6 figures. Tan Su, Haoxiang Yang, and Ruxin Wang contributed equally. Corresponding author: Binghui Xie

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑