IMPACT-VLA:面向视觉-语言-动作策略的基于反事实轨迹的交互感知多模态传播归因
IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies
浏览论文内容
中文总结 AI 辅助
针对VLA策略中多模态贡献阶段不明的问题,提出IMPACT-VLA方法,通过构建阶段-模态块并执行闭环反事实重执行,在30个LIBERO任务上揭示了模态贡献的阶段性变化与条件耦合。
中文摘要 AI 辅助
视觉-语言-动作(VLA)策略利用视觉观测、本体感觉状态和语言指令等多模态输入执行机器人操作任务。然而,目前尚不清楚每个模态在哪些执行阶段对最终任务成功做出贡献,以及输入干预如何通过后续状态、观测和动作进行传播。现有的归因方法主要衡量局部敏感性或时间聚合重要性,限制了其捕获阶段依赖贡献和跨阶段依赖的能力。我们提出了面向视觉-语言-动作策略的基于反事实轨迹的交互感知多模态传播归因(IMPACT-VLA)。IMPACT-VLA从一次成功的参考轨迹中的动作转换构建行为阶段,将其与策略查询边界对齐,并将阶段-模态块定义为归因单元。然后,它执行闭环反事实重执行,以量化每个块对最终任务成功的贡献。我们进一步分析了跨阶段的非加性交互和轨迹传播,同时区分了行为恢复与功能恢复。在30个使用OpenVLA-OFT的LIBERO机器人操作任务中,主导模态转换发生在25个任务中(83.3%),闭环归因比静态动作扰动更忠实地识别了任务关键信息。在早期阶段输入替换下,负交互对的后期块边际增益增加了约3.3倍,而功能恢复可能在没有行为恢复的情况下发生。这些结果揭示了多模态输入何时支持任务成功,以及它们的贡献如何在闭环执行期间变得有条件耦合。
英文摘要
Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, and language instructions. However, it remains unclear at which execution stages each modality contributes to final task success and how input interventions propagate through subsequent states, observations, and actions. Existing attribution approaches primarily measure local sensitivity or temporally aggregated importance, limiting their ability to capture phase-dependent contributions and cross-phase dependencies. We propose Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies (IMPACT-VLA). IMPACT-VLA constructs behavioral phases from action transitions in a successful reference rollout, aligns them with policy query boundaries, and defines phase-modality blocks as attribution units. It then performs closed-loop counterfactual re-execution to quantify each block's contribution to final task success. We further analyze cross-phase non-additive interactions and trajectory propagation while distinguishing behavioral from functional recovery. Across 30 LIBERO robot manipulation tasks using OpenVLA-OFT, dominant-modality transitions occurred in 25 tasks (83.3%), and closed-loop attribution identified task-critical information more faithfully than Static Action Perturbation. Later-block marginal gains for negatively interacting pairs increased by approximately 3.3x under early-phase input replacement, while functional recovery could occur without behavioral recovery. These results reveal when multimodal inputs support task success and how their contributions become conditionally coupled during closed-loop execution.
发表机构
- Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。