感知何种模态起作用:用于鲁棒视觉-语言-动作(VLA)策略的证据门控正则化
Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies
浏览论文内容
中文总结 AI 辅助
针对VLA策略的模态纠缠问题,提出无推理开销的证据门控正则化(EGR),在仿真与真实机器人设置中显著提升了应对传感器损坏、干扰物等情况的任务成功率。
中文摘要 AI 辅助
视觉-语言-动作(VLA)策略融合多模态感官输入,但在有限且同质化的机器人演示上训练会催生虚假的传感器间相关性,而非与任务相关的信号,我们将这种失效现象称为模态纠缠。在现实世界的遮挡和干扰物下,这表现为对非信息传感器的损坏产生不必要的敏感性,以及仅存一个信息传感器时出现的单模态不足。我们提出证据门控正则化(EGR),这是一种与模态无关的训练目标,无推理时开销。EGR推导每帧每传感器的任务相关信号,以门控两个状态条件一致性目标:低证据传感器的不变性,以及高证据传感器的单传感器充分性。我们引入基于BEHAVIOR-1K的基准,包含快速推理专用诊断套件和47个基于 rollout 的技能,用于应对模态纠缠。我们在该基准和两个具有完全不同 embodiment 的真实机器人设置上验证EGR:带有两个Kinova机械臂和三个RGB相机的双臂设置,以及结合视觉和GelSight触觉传感器的单臂MELFA ASSISTA设置。EGR在全模态下将模拟成功率(SR)从12.5%提升至16.4%(+31%),在非信息传感器损坏下从9.4%提升至16.5%(+75%),在单传感器 fallback下从2.8%提升至6.1%(+120%);在物理物体干扰物下,EGR使双臂设置的SR从30%提升至85%(+183%),触觉设置从55%提升至70%(+27%)。
英文摘要
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).
发表机构
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
- Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室)
机构由 AI 辅助整理,请以论文原文为准。