IMPACT:注意力是用于可扩展交互感知世界模型训练的交互图
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
浏览论文内容
中文总结 AI 辅助
该研究针对世界模型交互建模的监督分配不匹配问题,提出IMPACT框架,通过先验引导注意力校准构建交互图重加权监督,在机械臂等任务上优于MSE训练基线。
中文摘要 AI 辅助
世界模型在具身智能体的动作条件未来预测方面已取得显著进展,但仍难以对物理上合理的交互进行建模。现有方法通过用编码运动、几何或语义的外部表示约束生成过程来解决这一局限。获取这些时空密集表示通常需要辅助估计器或人工标注,限制了训练的可扩展性。我们转而重新审视训练目标,发现全局平均均方误差(MSE)去噪目标下存在监督分配不匹配的问题:占主导地位的静态内容控制了优化信号,而对交互生成至关重要的稀疏动态对象区域则受到不成比例的欠监督。受此观察启发,我们引入了IMPACT,这是一种具有先验引导注意力校准与定位的可扩展交互感知模型训练框架。IMPACT将与被操作对象标记关联的交叉注意力用作动作条件变化的内部时空先验,从该先验中采样候选区域,并用分离的局部预测误差对其进行校准以构建交互图,再用该图对去噪监督进行重加权,无需外部表示或推理时修改。在机械臂和人手操作任务上开展的广泛实验,涵盖多种控制模态和DiT骨干网络,结果显示IMPACT始终优于对应的MSE训练基线,提升了交互保真度、物理合理性和视觉质量。
英文摘要
World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Zhongguancun Academy(中关村学院)
- Tsinghua University(清华大学)
- Manifold AI
机构由 AI 辅助整理,请以论文原文为准。