arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04396cs.CV

CofactVLA:通过反事实干预消除视觉-语言-动作模型的混淆

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

Yan Zhang, Yinan Wu, Haoran Duan, Jungong Han

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型的视觉优先与因果混淆问题,提出CofactVLA框架,通过OPG和CCR机制消除视觉混淆,在仿真基准和真实机器人实验中均实现性能提升。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在机器人操控领域取得了显著进展,但它们普遍存在视觉优先现象。由于密集视觉流与稀疏语言指令之间存在严重的模态不平衡,VLA模型常陷入因果混淆:策略不将语言作为主要因果驱动因素,而是通过过拟合虚假视觉混淆因素(如突出物体或熟悉布局)完全绕过原始指令。为系统性缓解该偏差,我们将动作生成过程形式化为双路径去混淆图(DDG),并提出新型因果干预框架CofactVLA。该框架在单次前向传播中动态构建语言掩码的反事实分支,通过两种协同机制隔离并中和视觉混淆因素:一是动作级正交投影引导(OPG),在连续流匹配中从反事实视觉偏差几何投影事实速度场,提取纯语义意图;二是特征级反事实协方差缩减(CCR),通过惩罚协方差差的正特征空间数学去混淆潜在表示,明确抑制主导视觉捷径同时保留因果语言意图。大量实验表明,CofactVLA在多种仿真基准上达到新的最优性能;在仿真之外的真实机器人实验中,该方法在缩小泛化差距方面展现出因果效能,在分布外场景下取得52.3%的绝对成功率提升。

英文摘要

Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3\% absolute success rate gain under out-of-distribution scenarios.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑