发表机构
NYU Shanghai(上海纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究具身智能体中指令语义与动作执行之间的迁移鸿沟,提出SAT-Bench基准和VISA接口,证明语义可恢复但动作表达弱,VISA能大幅降低无效指令的盲目执行。
AI 中文摘要
具身语言落地(embodied language grounding)需要的不仅仅是识别指令的指代对象:恢复的语义还必须控制智能体所展现的动作。我们将这一缺失环节研究为语义-动作鸿沟(semantic-action gap),即指令语义可恢复但在原生连续动作中表达较弱。我们引入了SAT-Bench,一个固定观测的反事实基准,它在仅改变指令语义的同时保持视觉场景和智能体状态不变。在LIBERO目标名称和像素级关系交换上,目标恢复率达到100.0%和95.8%,而OpenVLA的动作敏感性仅为6.8%和7.7%。这一鸿沟在额外的1000个组合性和时间/程序性反事实中持续存在,整体动作敏感性为6.1%。隐藏状态、无阈值、跨策略和展开诊断进一步支持了这种语义-动作迁移失败。我们引入了VISA,一个轻量级的执行时接口,将恢复的语义转换为允许(ALLOW)、延迟(DEFER)、目标一致性和验证选择决策。VISA将无效指令的盲目执行从92.7%降至2.8%,同时保留了94.0%的正常命令,验证选择进一步改善了目标一致的动作暴露,而无需更新底层策略。总体而言,具身语言评估应衡量语义-动作迁移,而不仅仅是语义解析。
英文摘要
Embodied language grounding requires more than identifying the referent of an instruction: recovered semantics must also control the action an agent exposes. We study this missing link as a semantic-action gap, where instruction semantics are recoverable but weakly expressed in native continuous actions. We introduce SAT-Bench, a fixed-observation counterfactual benchmark that holds the visual scene and agent state fixed while changing only instruction semantics. On LIBERO target-name and pixel-grounded relation swaps, target recovery reaches 100.0% and 95.8%, whereas OpenVLA action sensitivity remains only 6.8% and 7.7%. The gap persists across 1,000 additional compositional and temporal/procedural counterfactuals, with overall action sensitivity of 6.1%. Hidden-state, threshold-free, cross-policy, and rollout diagnostics further support this semantic-action transfer failure. We introduce VISA, a lightweight execution-time interface that converts recovered semantics into ALLOW, DEFER, target-consistency, and verified-selection decisions. VISA reduces invalid-instruction blind execution from 92.7% to 2.8% while preserving 94.0% of normal commands, and verified selection further improves target-consistent action exposure without updating the underlying policy. Overall, embodied language evaluation should measure semantic-action transfer, not semantic parsing alone.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 24 pages, 7 figures