解耦视觉-语言-动作模型中的虚假相关性:通过预测域不变潜在前瞻
Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead
浏览论文内容
中文总结 AI 辅助
提出域不变潜在前瞻(DILL)框架,通过预测域不变未来潜在变量解耦任务相关结构与域特定变化,缓解VLA模型中的捷径学习,提升视觉鲁棒性,在LIBERO-Plus上平均成功率69.1%,超最强基线11.4个百分点。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型在视觉分布偏移下仍然脆弱,往往依赖于与特定域因素相关的虚假相关性,而非任务相关结构。我们提出域不变潜在前瞻(DILL),一种表示学习框架,用于缓解VLA策略中的捷径学习。我们的关键思想是利用从域变换轨迹数据中学习到的域不变未来潜在变量来监督策略。任务-域编码器通过对比目标和高斯解耦正则化进行训练,以分离任务相关结构与特定域的视觉变化。学习到的编码器随后通过前瞻预测和域解耦为VLA策略学习提供未来潜在变量,鼓励策略关注任务相关结构而非偶然的视觉因素。反事实任务视图评估表明,DILL减少了捷径依赖,而LIBERO-Plus评估展示了改进的视觉鲁棒性,平均成功率为69.1%,比最强基线高出11.4个百分点。真实世界操作实验进一步支持DILL在受控模拟之外的应用性。补充的潜在空间诊断表明,这些行为增益伴随着更好地保留任务一致结构同时抑制特定域变化的表示。我们的项目页面可在此https URL获取。
英文摘要
Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.
发表机构
- OpenMind
- Seoul National University(首尔大学)
- Hyundai Motors(现代汽车)
- Ajou University(亚洲大学)
- Tommoro Robotics
机构由 AI 辅助整理,请以论文原文为准。