arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldGuide:面向流程任务执行的目标导向视频世界模型

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan

arXiv 2610.12459首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出WorldGuide模型,通过联合训练规划器与执行器实现闭环流程任务执行,在WorldGuide-Bench和VideoCraft-Bench上的任务成功率均优于对比模型MiniMax-H3,验证了规划与执行耦合的重要性。

AI 中文摘要

视频生成器和基于视频的世界模型可合成合理的视觉轨迹,但长程流程任务要求生成过程适配实际已生成的内容。模型必须从自身生成的状态中确定下一个动作,执行该动作,并识别任务何时完成。开环生成无法适配执行结果,而现有闭环系统常依赖预训练执行器或间接验证,这在动作决策与成功实现之间留下了缺口。我们将流程视频生成形式化为「视觉世界空间中的闭环任务执行」,并提出WorldGuide。仅给定初始图像和任务目标,WorldGuide会预测一个原子动作,生成其对应的视频片段,并利用生成结果选择下一个动作或终止。规划器(Planner)和执行器(Executor)在相同的步骤级流程演示上进行训练:规划器学习从视觉进度预测下一个原子动作或任务完成,执行器则直接被训练以实现预测的动作。分层视觉记忆以受限的历史 token 成本维持长程执行的状态。由于缺乏用于联合规划器-执行器训练的步骤级动作-视频监督,我们引入WorldGuide Bench:涵盖245个任务、27个流程类别的约5.9万个带步骤标注的视频。WorldGuide在WorldGuide-Bench上实现了33.33%的任务成功率,而强大的近期视频模型MiniMax-H3(尽管获得了参考动作规划)的成功率为29.90%;在仅以目标为条件的情况下,WorldGuide在VideoCraft-Bench上的成功率为47.69%,MiniMax-H3为32.73%。这些结果证明了将规划与学习到的执行相结合对目标导向流程视频生成的重要性。

英文摘要

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33\% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90\% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69\% on \textbf{VideoCraft-Bench} compared with 32.73\% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.

Comments34 pages, 14 figures, 15 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑