发表机构
Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CLAMP 通过将场景证据转化为解码时约束,利用硬掩码和 HMM 前瞻模块,提升视觉-语言模型在具身规划中的可执行性与安全性。
AI 中文摘要
具身规划日益依赖视觉-语言模型(VLM)将指令和视觉观察转化为可执行的动作序列。然而,流畅的计划并不总是可执行的。VLM 可能引用未在视觉上观察到的物体,选择所需可供性不可用的动作,或违反语法和动作约束。我们提出 CLAMP,一个多模态约束锚定框架,将场景证据转化为冻结的 VLM 规划器在解码时的约束。CLAMP 利用初始观察将物体引用限制为场景支持的物体,同时提供的符号动作模型指定状态转换和目标。在解码过程中,硬掩码消除无效的下一个词元候选,而基于隐马尔可夫模型(HMM)的世界状态前瞻模块根据动作前置条件和目标可达性重新加权剩余可行候选的概率。这使得规划器能够保留 VLM 的语言先验,同时防止视觉上不支持、不安全或不可行的候选进入计划。对于未见过的任务和环境,CLAMP 在测试时使用从冻结的 VLM 采样的无标签续写来调整 HMM。在 VLABench、SafeAgentBench 和 TaPA 上的实验表明,场景锚定的约束改善了物体锚定和安全性,而大多数剩余失败源于感知错误或约束规范不匹配。
英文摘要
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executable. A VLM may refer to objects that are not visually observed, select actions whose required affordances are unavailable, or violate syntax and action constraints. We introduce CLAMP, a multimodal constraint-grounding framework that turns scene evidence into decoding-time constraints for a frozen VLM planner. CLAMP uses the initial observation to restrict object references to those supported by the scene, while a provided symbolic action model specifies state transitions and goals. During decoding, hard masks eliminate invalid next-token candidates, while a Hidden Markov Model (HMM)-based world-state lookahead module reweights the probabilities of the remaining feasible candidates based on action preconditions and goal reachability. This allows the planner to retain the VLM's language prior while preventing visually unsupported, unsafe, or infeasible candidates from entering the plan. For unseen tasks and environments, CLAMP adapts the HMM at test time using label-free continuations sampled from the frozen VLM. Experiments on VLABench, SafeAgentBench, and TaPA show that scene-grounded constraints improve object grounding and safety, while most remaining failures stem from perception errors or misaligned constraint specifications.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026