arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视觉语言动作模型测试时模态适应的因果感知推理-诊断-细化框架

A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

Haoyu Zhang, Yuwei Wu, Jin Chen, Gao Zhi, Zhenxin Diao, Mingyang Gao, Kun Wu, Yongchun Liu, Fan Li

arXiv 2607.25516首次发表:更新:

发表机构

Beijing Institute of Technology; AInnovation Co. Ltd.; Beijing Innovation Center of Humanoid Robotics(北京理工大学; A创新有限公司; 北京人形机器人创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型模态融合问题,提出因果感知推理-诊断-细化框架,通过视觉观察推理、因果效应诊断及动作预测细化实现模态适应,设计因果感知动作细化器,实验验证了该框架在多VLA主干上提升整体性能的有效性。

AI 中文摘要

视觉语言动作(VLA)模型根据视觉观察和本体感觉状态预测顺序动作以执行语言指令指定的任务。但由于机器人操作涉及动态阶段,视觉观察重要性随时间变化,VLA模型中模态融合仍是开放问题。本文提出推理-诊断-细化(IDR)框架,可与多种VLA架构集成在测试时细化动作预测。IDR先在视觉观察的事实和反事实场景下推理动作,诊断视觉观察因果效应作为动态重要性估计,用于无训练方式细化动作预测。还设计了因果感知动作细化器实现IDR框架,包括零填充干预、基于规范的量化和门控残差融合。仿真基准和现实任务实验表明该框架提升了多VLA主干的整体性能,证明了在测试时动态调整视觉重要性的有效性。

英文摘要

Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time. In this paper, we propose an infer-diagnose-refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑