arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vision-TL-Action:基于视觉观测和时序逻辑的神经符号轨迹生成

Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic

Zezhi Liu, Zhiwei Zheng, Hanqian Luo, Deyun Qin, Shizhen Wu, Yongchun Fang

arXiv 2607.26770首次发表:更新:

AI 中文总结

本文提出Vision-TL-Action模型,通过融合视觉与TL节点标记生成动作轨迹,在Panda、AntMaze任务中优于或接近基线,实现无需物体几何信息的视觉到TL目标的轨迹生成。

AI 中文摘要

时序逻辑(TL)为制定长时程机器人任务提供了一种组合式语言,但现有的TL条件轨迹生成器可通过在任务图中编码精确的物体几何信息来规避感知到符号的绑定问题。本文提出Vision-TL-Action模型,其可从多视角图像、无坐标的TL语法图以及机器人初始状态生成动作轨迹。TL节点标记与空间视觉标记通过双向交叉注意力进行融合,融合后的表示用于条件化流匹配轨迹生成器;视觉标记仅补充归一化图像平面位置和相机视角标识符,而仅用于训练的谓词到区域目标则鼓励模型对参考物体进行接地。与该领域先前工作一致,本研究采用Success@K(即K条采样轨迹中至少有一条满足TL规范的任务占比)评估模型:在Panda任务中,本模型的Success@1024达到67.45%,而神谕状态基线为59.11%;在AntMaze任务中,本模型的Success@256达到96.35%,与神谕结果96.88%相当。分辨率和干预研究表明,空间细节依赖于语义接地,且谓词身份会影响注意力和性能。这些结果表明,本模型可实现从视觉观测和结构化TL目标到动作轨迹的直接映射,无需在推理阶段使用物体几何信息,代码可在指定URL获取。

英文摘要

Temporal logic (TL) provides a compositional language for the formulation of long horizon robotic tasks, but existing TL-conditioned trajectory generators can sidestep perception-to-symbol binding by encoding exact object geometry in the task graph. We introduce \emph{Vision-TL-Action}, which generates action trajectories from multi-view images, a coordinate-free TL syntax graph, and the robot initial state. TL-node tokens and spatial visual tokens are fused through bidirectional cross-attention, and the resulting representation conditions a flow-matching trajectory generator. Visual tokens are augmented only with normalized image-plane locations and camera-view identifiers, while a training-only predicate-to-region objective encourages grounding to referenced objects. Consistent with prior work in this domain, we evaluate the model using Success@$K$, the fraction of tasks for which at least one of K sampled trajectories satisfies the TL specification. On Panda task, our model achieves 67.45% Success@1024, compared with 59.11% for the oracle-state baseline. On AntMaze task, it achieves 96.35% Success@256, comparable to the oracle result of 96.88%. Resolution and intervention studies show that spatial detail depends on semantic grounding and predicate identity affects both attention and performance. These results demonstrate a direct mapping from visual observations and structured TL goals to action trajectories without requiring object geometry at inference. Code is available at https://github.com/AricLau07/vision-tl-action.

Comments17 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑