发表机构
Shanghai Normal University; East China Normal University; Northwest A&F University; Xidian University(上海师范大学; 华东师范大学; 西北农林科技大学; 西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言-动作模型存在的位置盲区问题,提出两阶段黑箱框架,通过网格划分与单侧对数似然比检验定位盲区,再经LoRA微调缓解,在五个模型上验证了方法的有效性。
AI 中文摘要
近期的视觉-语言-动作(Vision-Language-Action, VLA)模型在机器人操纵任务中取得了可观的性能,通常通过预定义物体配置下的成功率来衡量,这种评估方法隐含假设模型在工作空间内具备空间均匀的能力。然而,该假设并不成立:即使指令和其他所有场景因素保持固定,仅将与任务无关的干扰项重新放置,也可能在局部、空间连贯的区域内大幅提升失败概率,我们将这些区域称为位置盲区(Positional Blind Spots, PBS)。本文提出一种两阶段黑箱框架以揭示和缓解PBS:在揭示阶段,我们将工作空间网格化,并应用单侧对数似然比检验来定位具有显著升高风险的PBS单元格;在缓解阶段,我们通过LoRA对从这些PBS区域收集的演示数据进行策略微调,提升该区域的能力,同时基本保持工作空间其余部分的性能。我们在两个基准测试中对五种最先进的VLA策略评估了该框架,发现PBS在所有策略中普遍存在且空间集中,失败率最高达0.58。我们的搜索策略平均F1分数为0.678,分别比随机搜索和自适应采样基线高出0.268和0.178;在发现区域的引导下,针对性微调将整体失败率降低了40.00%至85.19%。
英文摘要
Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over predefined object configurations, an evaluation that implicitly assumes spatially uniform competence across the workspace. However, this assumption does not hold: even with the instruction and every other scene factor held fixed, merely relocating a task-irrelevant distractor can sharply raise the failure probability within localized, spatially coherent regions, which we term Positional Blind Spots (PBS). In this paper, we propose a two-stage black-box framework to uncover and mitigate PBS. During the uncovering stage, we grid the workspace and apply a one-sided log-likelihood-ratio test to localize PBS cells with significantly elevated risk. During the mitigation stage, we fine-tune the policy via LoRA on demonstrations collected from these PBS regions, improving competence there while largely preserving performance across the rest of the workspace. We evaluate our framework on five state-of-the-art VLA policies across two benchmarks, and find that PBS are pervasive and spatially concentrated in all of them, with failure rates up to 0.58. Our search strategy achieves an average F1-score of 0.678, outperforming random search and adaptive sampling baselines by 0.268 and 0.178, respectively. Guided by the discovered regions, targeted fine-tuning reduces the overall failure rate by 40.00%--85.19%.