arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

技能级世界模型的流策略作为动作:面向长时程规划的学习与符号抽象

Flow Policies as Actions of Skill-Level World Models: Learned and Symbolic Abstractions for Long-Horizon Planning

Andreu Matoses Gimenez, Andrei-Carlo Papuc, Chris Pek, Javier Alonso-Mora

arXiv 2610.04767首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将流匹配策略的输入作为技能级动作,构建四种抽象(压缩种子、两种学习代码、符号标签),在块重排任务上验证,符号标签在六技能内成功率超90%,并分析搜索限制。

AI 中文摘要

潜在世界模型使机器人能够通过预测动作的后果来进行规划。使用控制率动作规划长任务需要许多预测步骤,这扩大了搜索空间并累积误差。技能级动作缩短了这些序列,但符号技能词汇需要领域知识和带标签的演示。我们从在分割为完整技能的演示上训练的流匹配策略的输入构建技能级动作。该策略将噪声种子和观测(可选地带有代码或标签)映射到完整的技能执行,因此一次执行即一次世界模型转移。在此机制上,我们提出了四种具有递增任务知识的动作抽象:压缩种子、从演示中学习到的两种离散代码,以及一种符号标签。我们使用共同的世界模型训练程序和规划框架,在需要多达14个连续技能的模拟块重排任务上评估它们。符号标签在需要多达六个技能的任务中成功率超过90%,超出后性能下降。在没有标签的情况下,一种以对象为中心的学习代码在单技能任务上与之匹配,并在两到五个技能的任务中保持其成功率的二分之一到四分之三。消融研究将标签的大部分优势归因于其规划器知道哪些动作适用,而非标签本身。超过六个技能后,限制成功的是搜索,而非世界模型。项目页面:此https URL

英文摘要

Latent world models enable robots to plan by predicting the consequences of actions. Planning long tasks with control-rate actions requires many prediction steps, which enlarges the search space and accumulates error. Skill-level actions shorten these sequences, but a symbolic skill vocabulary requires domain knowledge and labeled demonstrations. We construct skill-level actions from the inputs of a flow-matching policy trained on demonstrations segmented into complete skills. The policy maps a noise seed and an observation, optionally with a code or label, to a complete skill execution, so one execution is one world-model transition. On this mechanism we propose four action abstractions with increasing task knowledge: a compressed seed, two discrete codes learned from the demonstrations, and a symbolic label. We evaluate them with a common world-model training procedure and planning framework on simulated block rearrangement tasks that require up to 14 sequential skills. The symbolic label succeeds in over 90% of the tasks that require up to six skills and degrades beyond. Without any label, an object-centric learned code matches it on single-skill tasks and retains half to three quarters of its success on tasks of two to five skills. Ablations attribute much of the label's advantage to its planner knowing which actions are applicable, rather than to the label itself. Beyond six skills the search, not the world model, limits success. Project page: https://andreumatoses.github.io/research/flow-skill-wm

Comments9 pages, 7 figures, 1 table. Submitted to IEEE ICRA 2027. Project page: https://andreumatoses.github.io/research/flow-skill-wm

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑