arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理是一把双刃剑:视觉-语言-动作模型中的架构与跨阶段鲁棒性

Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models

Tuan Duong Trinh, Naveed Akhtar, Basim Azam

arXiv 2607.17786首次发表:更新:

发表机构

University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉-语言-动作模型中添加推理步骤对模型鲁棒性的影响,通过在三个模型上进行实验,发现潜在迭代模型鲁棒性最差,推理输出监测在公平测试下失败,计划-行动一致性探测器在自适应攻击下效果不佳,为可行防御设定了前提条件。

AI 中文摘要

添加推理步骤是否会使视觉-语言-动作(VLA)模型对扰动更具鲁棒性?直观上,在行动前进行推理的策略应比直接将观察映射到行动的策略更好地吸收扰动输入。我们在跨越推理范围的三个模型(无推理、文本思维链和潜在迭代循环)上直接测试这一前提,在LIBERO和SimplerEnv的视觉、推理和行动阶段对每个模型进行扰动。研究围绕两个问题展开:推理设计是否会改变鲁棒性,以及推理在运行时能否作为安全信号被反馈回来?我们发现潜在迭代模型的鲁棒性最差:在随机噪声和白盒扰动下其任务成功率都会崩溃,而其他两个模型则保持稳定。这种脆弱性是结构性的而非累积性的:在推理时改变推理深度几乎不会改变它。推理输出原则上可以被监测,但监测器在公平测试下会失败。一个在简单评估下看似近乎完美的计划-行动一致性探测器在自适应攻击下却降至随机水平。在匹配误报率校准下,将其与动作异常探测器融合也无法使防御后的成功率高于未防御时。在白盒视觉阶段攻击下,限于这些输出级行为探测器,这个上限是任何可行防御必须首先满足的前提条件。

英文摘要

Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons before acting should absorb a perturbed input better than one that maps observations directly to actions. We test this premise head-on across three models that span the reasoning spectrum (no reasoning, a text chain-of-thought, and a latent iterative loop), perturbing each at the vision, reasoning, and action stages on LIBERO and SimplerEnv. Two questions organize the study: does the reasoning design shift robustness, and can the reasoning be read back at runtime as a safety signal? We find that the latent-iterative model is by far the least robust: under both stochastic noise and white-box perturbation its task success collapses, while the other two hold. This fragility is structural rather than cumulative: varying the reasoning depth at inference barely moves it. Reasoning outputs can in principle be monitored, but the monitors fail under fair tests. A plan--action consistency probe that looks near-perfect under naive evaluation falls to chance under adaptive attack. Under matched-FPR calibration, fusing it with an action-anomaly probe never lifts defended success above undefended. Scoped to these output-level behavioral probes under white-box vision-stage attack, this ceiling is a precondition that any viable defense must first satisfy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑