arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏特征策略遗忘缓解视觉-语言-动作模型中的状态幻觉

Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models

Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim

arXiv 2610.09496首次发表:更新:

发表机构

Chung-Ang University; Georgia Institute of Technology; NAVER AI Lab(中央大学; 佐治亚理工学院; NAVER AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型中的状态幻觉问题,通过稀疏自编码器识别相关特征,提出SOUL方法选择性遗忘幻觉策略知识,在模拟和真实环境中显著减少失败并提升任务成功率。

AI 中文摘要

视觉-语言-动作(VLA)模型通过利用预训练视觉-语言模型的丰富表示,在机器人操作中展现出强大的泛化能力。然而,它们在真实环境中的部署仍受到反复出现的不可靠行为的限制。在本工作中,我们研究了状态幻觉,这是一种反复出现的失败模式,即VLA继续行动,仿佛未实现的机器人-物体状态已经达成。我们的分析发现,状态幻觉与对任务相关视觉区域的注意力减弱同时发生,通过稀疏自编码器的机制解释揭示,与幻觉相关的稀疏特征在这些失败发生时被激活。基于这一分析,我们提出了SOUL(稀疏特征策略遗忘),它选择性地遗忘与状态幻觉行为相关的策略知识,其中从幻觉失败和成功行为中识别出的稀疏特征分别作为明确的遗忘目标和保留目标。在模拟和真实环境中跨VLA架构的实验表明,我们的方法大幅减少了幻觉失败并提高了任务成功率,而不会显著损害现有的操作能力。这些结果表明,可解释的特征分析为选择性修改机器人策略中的不良知识提供了实用基础。

英文摘要

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑