超越外观变化:面向VLA模型的任务语义动作校准
Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models
浏览论文内容
中文总结 AI 辅助
针对VLA模型在任务语义变化下动作漂移或响应不足的问题,提出BAS-VLA框架,通过破坏中心校准与证据门控辅助,在保持干净性能的同时实现任务语义分离。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型在具身操作任务中取得了强劲性能,但仍缺乏一种明确的机制来平衡行为稳定性与任务语义敏感性。我们识别出两种互补的失败模式。在任务保持变化下,任务语义保持不变但场景外观发生变化(如风格、光照、杂乱或释义),策略常常表现出不必要的动作漂移。相反,在语义破坏变化下,关键任务语义(如目标物体或约束)被改变,策略往往无法产生足够不同的行为,而是遵循原始轨迹。为解决这一差距,我们提出BAS-VLA,一种构建在冻结基础VLA之上的任务语义动作校准框架。BAS-VLA采用以破坏为中心的校准核心作为默认路径,并引入一种选择性证据门控的保持辅助模块,该模块仅在检测到干扰变化且任务语义保持一致时激活。在OpenPI-pi0.5 / LIBERO-Object Milk-Swap基准上,BAS-VLA在干净条件(98.0%)和语义保持条件(97.5%)下保持高成功率,同时在故意目标物体交换下将干净标准成功率降至0.0%,展示了强大的陈旧任务抑制和任务语义分离能力。在验证的风格保持变化中,它将成功率从42%提升至70%,且不降低干净性能。这些结果强调,可靠的VLA行为需要超越外观鲁棒性,转向明确的任务语义动作校准。
英文摘要
Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学国家计算机动画中心)
- Shanghai Jiao Tong University(上海交通大学)
- Eastern Institute of Technology, Ningbo(宁波东方理工大学)
机构由 AI 辅助整理,请以论文原文为准。