arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弯曲的拒绝:测量与预测具身视觉语言规划器中的任务可塑性

Refusals That Bend: Measuring and Predicting Task Malleability in Embodied VLM Planners

Leo Y. Lin, Mikhail Kuznetsov, Muslum Ozgur Ozmen, Z. Berkay Celik

arXiv 2609.38971首次发表:更新:

发表机构

Purdue University; Amazon; Arizona State University(普渡大学; 亚马逊; 亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过向环境中放置日常物品,测试具身VLM规划器拒绝的可塑性,发现20.2%的任务可被绕过,并提出基于开源VLM信号预测任务可塑性的方法,建议部署前评估。

AI 中文摘要

具身视觉语言模型(VLMs)正越来越多地被部署为机器人的高层规划器,因为它们在多样环境中具有泛化能力。然而,这要求其安全对齐在未见环境中同样有效。现有的红队测试假设对手优化提示词、像素或环境中的文本,现有基准则询问规划器是否在固定场景中识别或缓解危险。两者均未询问规划器已给出的拒绝是否能在环境的普通变化下存续。我们通过将一个日常物品放入环境中提出该问题,且像素、梯度或提示词均不受对抗控制。在宪法守护的规划器最初拒绝的846项任务中,我们发现20.2%的任务可通过一个或多个物品翻转为合规,且物品数量因任务而异。此外,物品无需针对任务选择,即从固定列表中抽取的物品,在不知环境或指令的情况下,其绕过安全的频率与针对特定任务提出的物品相当。我们定性对比了最常被绕过和最不常被绕过的任务,发现区别在于指令和环境中的危险显著性。因此,对安全绕过的敏感性是任务的一种属性,我们称之为“可塑性”,并表明可在目标被查询之前对其进行预测。从一个小型开源VLM读取的复合信号识别可塑性任务的频率是随机选择的2.4倍。日常物品,无论是由对手放置还是由环境的普通重排引入,都足以推翻拒绝。由于敏感性由任务的指定方式决定,我们建议在部署前对每个任务评估可塑性。

英文摘要

Embodied vision-language models (VLMs) are increasingly deployed as high-level planners for robots because they generalize across diverse environments. However, this requires their safety alignment to also hold in unseen environments. Existing red-teaming assumes an adversary who optimizes the prompt, the pixels, or text in the environment, and existing benchmarks ask whether a planner recognizes or mitigates a hazard in a fixed scene. Neither asks whether a refusal the planner has already given survives an ordinary change to the environment. We ask that question by placing a single everyday object into the environment, with no pixel, gradient, or prompt under adversarial control. On $846$ tasks that a constitution-guarded planner initially refuses, we find $20.2\%$ of tasks can be flipped to compliance by one or more objects, and the number of objects differs from one task to another. In addition, the object need not be chosen for the task, i.e., items drawn from a fixed list, with no knowledge of the environment or the instruction, bypass safety about as often as items proposed for the specific task. We qualitatively contrast the tasks bypassed most and least often and find that the distinction lies in how conspicuous the hazard is in the instruction and environment. Susceptibility to safety bypass is therefore a property of the task, which we call its \emph{malleability}, and we show that it can be predicted before the target is ever queried. A composite of signals read from a small open-source VLM identifies malleable tasks $2.4\times$ as often as picking at random. Everyday objects, whether placed by an adversary or introduced by ordinary rearrangement of the environment, are thus sufficient to overturn a refusal. Because susceptibility is determined by how a task is specified, we recommend assessing malleability per task prior to deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑