arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video2STL:将VLM生成的时间规范落地用于机器人学习

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh

arXiv 2609.37519首次发表:更新:

发表机构

University of Southern California; University of Florida(南加州大学; 佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Video2STL将视频转换为参数化STL规范,分离长短时标奖励,实现跨形态机器人学习,在操作和四足任务中超越基线。

AI 中文摘要

基于视频的策略学习特别有前景,因为它无需动作标注或与具体形态匹配的演示即可展示目标行为。一个核心挑战是决定从视频中向机器人传递哪些信息。现有方法通常将视觉观察转换为标量相似度或价值信号,或要求基础模型直接生成奖励代码。这些方法可能使任务的时间结构难以检查、落地和重用。我们提出Video2STL,一个将仅含观察的视频转换为参数化信号时序逻辑(STL)规范,并利用该形式化表示进行机器人学习的框架。视觉语言模型提取与形态无关的语义事件轨迹,并构建一个符号化时间规范库。模型确定任务结构,而数值谓词阈值和时间边界则从成功的机器人轨迹中落地。对于策略学习,我们分离短时标和长时标的时间信息:短时域规范通过滚动窗口的定量鲁棒性提供密集奖励,而基于保留的长时域规范的因果监控器为有效的时间前缀提供一次性进度奖励。同一表示支持从人类或动物视频到机器人控制的跨形态迁移。在四个操作任务中,Video2STL实现了平均85.8%的一次成功率和67.0%的最终成功率,而原生密集PPO为81.5%/59.5%,Text2Reward为65.0%/42.3%;在四足机器人 locomotion 中,基于Qwen-3.8和GPT-5.6的Video2STL策略在0.3至2.1米/秒的速度范围内实现了100%的成功率,同时在高速能效方面保持竞争力。项目网页:https://video2stl。

英文摘要

Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑