发表机构
University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DIVA提出双空间意图感知视觉衰减模块,通过锚定-衰减机制调节视觉令牌影响,提升VLA策略性能,在LIBERO上平均成功率提升至98.0%,并增强真实世界鲁棒性。
AI 中文摘要
视觉-语言-动作(VLA)策略通常将密集的视觉补丁令牌馈送到语言-动作主干网络中,以保留场景上下文,但缺乏明确的机制来调节不同视觉令牌对策略计算的影响强度。我们提出DIVA,一种采用“先锚定后衰减”设计的双空间意图感知视觉衰减模块。DIVA将高层任务意图与低层视觉证据相结合,估计逐补丁的相关性锚点,然后在两个互补空间中应用这些锚点:在主干网络入口前对投影后的视觉令牌重新加权,并在主干网络内部持续衰减低相关性的视觉状态。DIVA保留完整的视觉令牌序列,且无需外部接地监督。在LIBERO基准上,DIVA将OpenVLA-OFT的平均成功率从96.6%提升至98.0%,并将其零样本LIBERO-Plus得分从69.6提高到72.6。真实世界实验进一步表明,在任务无关的视觉扰动下,DIVA持续取得增益,支持了意图感知视觉衰减在仿真环境之外的鲁棒性。
英文摘要
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.