arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

告诉机器人不要做什么:一种否定理解视角

Tell Robot What Not to Do: A Negation Understanding Perspective

Fazeng Li, Gan Sun, Hao Cheng, Weihong Ren, Yang Cong

arXiv 2610.11952首次发表:更新:

发表机构

South China University of Technology; Harbin Institute of Technology(华南理工大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出NegaAlign框架,通过引入否定变换层和教师引导对齐机制,结合NegaBench基准,提升了VLAs遵循否定指令的成功率,且不影响肯定指令性能。

AI 中文摘要

指令跟随使机器人能够执行自然语言指定的各类任务,是人机交互的基础能力。除了传达期望的结果,用户还需要明确不应执行的约束条件。本文研究如何让视觉-语言-动作模型(VLAs)遵循否定指令,即机器人必须在满足明确排除条件的同时完成任务目标。为此,我们提出NegaAlign,这是一种参数高效的即插即用框架,仅通过图像-语言监督即可扩展预训练VLAs以遵循否定指令。具体而言,我们在视觉-语言骨干网络的选定层中引入否定变换层,以重塑中间指令表示;同时设计教师引导的对齐机制,使与指令相关的视觉标记对齐,从满足否定约束的指令中迁移与动作相关的 grounding。训练阶段使用从现有演示构建的监督数据,仅更新插入的层,保持包括动作生成器在内的所有预训练参数冻结。我们还推出NegaBench,这是一个涵盖五个领域10种场景的仿真基准,用于系统评估否定约束下的操纵任务。在GR00T、π₀和π₀.5上的实验表明,NegaAlign在否定指令跟随任务中取得了一致提升。NegaAlign拥有1160万可训练参数,在NegaBench上将π₀.5的否定指令成功率从2.60%提升至88.45%,在真实世界任务中从12.4%提升至88.8%,同时保留了对肯定指令的性能。

英文摘要

Instruction following enables robots to perform diverse tasks specified in natural language, making it a fundamental capability for human-robot interaction. Beyond communicating desired outcomes, users also need to specify constraints on what not to do. We investigate how to enable vision-language-action models (VLAs) to follow negated instructions, where robots must accomplish task goals while respecting explicit exclusions. To this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone. Specifically, we introduce Negation Transformation Layers into selected layers of the vision-language backbone to reshape intermediate instruction representations. Meanwhile, a teacher-guided alignment mechanism is designed to align instruction-relevant visual tokens, transferring action-relevant grounding from instructions that satisfy the negated constraint. The training phase uses supervision constructed from existing demonstrations and updates only the inserted layers, keeping all pretrained parameters frozen, including the action generator. We further introduce NegaBench, a simulation benchmark spanning 10 scenarios across five domains for systematically evaluating manipulation under negated constraints. Experiments across GR00T, $π_0$, and $π_{0.5}$ demonstrate consistent improvements in negated instruction following. With 11.6M trainable parameters, NegaAlign increases the negated-instruction success rate of $π_{0.5}$ from 2.60% to 88.45% on NegaBench and from 12.4% to 88.8% on real-world tasks, while retaining performance on affirmative instructions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑