arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STeP:用于视觉语言模型动作生成精确规范的信号时序逻辑

STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

Kasra Torshizi, Anukriti Singh, Sidharth Mathur, Khuzema Habib, Leo Du, Pratap Tokekar

arXiv 2607.18580首次发表:更新:

发表机构

University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言动作模型缺乏可解释性及难以遵循精确自然语言指令的问题,提出用信号时序逻辑(STL)连接高级语言理解与低级机器人执行的分层框架,经实验验证该框架能提高语言条件下机器人规划的精度、可靠性和可解释性。

AI 中文摘要

视觉语言动作(VLA)模型虽有出色泛化能力,但缺乏可解释性,难以遵循编码空间、时间和逻辑要求的精确自然语言指令。我们提出一个分层框架,使用信号时序逻辑(STL)作为连接高级语言理解与低级机器人执行的共享表示。高级策略利用VLM将语言指令分解为高级子任务,为每个子任务生成STL规范并选择低级策略执行子任务。STL规范将语言意图转化为精确约束,低级策略选择决定是通过STL引导的模型预测控制直接执行约束,还是在执行感知复杂或接触丰富行为的学习策略时进行监控。通过将STL集成到计划验证、低级策略、子任务监控和重新规划中,我们的框架使基于语言的计划在运行时能够使用通用形式结构进行检查、优化和修订。我们在真实桌面领域评估了该方法,并展示了形式规范如何提高语言条件下机器人规划的精度、可靠性和可解释性。

英文摘要

Natural-language robot instructions often specify more than a coarse task goal: they may impose spatial, temporal, and logical requirements that must remain satisfied throughout execution. We present STeP, a specification-based agentic framework that uses Signal Temporal Logic (STL) as an explicit interface between high-level language reasoning and low-level robot execution. Rather than encoding such requirements implicitly in a learned policy, STeP formalizes them as task specifications that can be decomposed across multi-stage manipulation, enforced during execution, monitored online, and used as structured feedback for replanning. We evaluate STeP on standard LIBERO and LIBERO-PRO, and introduce LIBERO-Constrained, a new benchmark for manipulation tasks with spatial, temporal, and logical requirements, together with real-world tabletop experiments. On LIBERO-PRO, STeP retains 54%-89% success across five of six evaluated perturbation settings, substantially outperforming VLA and code-as-policy baselines under distribution shift. On LIBERO-Constrained, STeP achieves 80% safe success across 49 task-constraint instances; on real-world tasks, it improves safe success over a specification-free model-based baseline across all four task categories, with gains of up to 45 percentage points. These results support explicit formal specifications as a practical interface between foundation-model reasoning and reliable robot execution.

Comments9 Pages, 5 Figures, 3 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑