arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.01477cs.RO

动作代理:代理视频生成与流约束扩散的结合

Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

  • Intelligent Space Robotics Laboratory, Skoltech(斯克里普切尔技术大学智能空间机器人实验室)

机构由 AI 辅助整理,请以论文原文为准。

Jeffrin Sam, Nguyen Khang, Yara Mahmoud, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

更新

中文总结 AI 辅助

本文提出Action Agent框架,结合代理导航视频生成与流约束扩散控制,通过两阶段方法提升多体机器人导航性能,实现从语言和图像输入生成物理合理的第一人称导航视频,并在真实和仿真环境中取得高成功率。

中文摘要 AI 辅助

我们提出了Action Agent,一种两阶段框架,将代理导航视频生成与流约束扩散控制结合,用于多体机器人导航。在第一阶段,大型语言模型(LLM)作为协调模块,选择视频扩散模型,通过迭代验证优化提示,并积累跨任务记忆,从语言和图像输入生成物理合理的第一人称导航视频,将视频生成成功率从35%(单次生成)提高到50个导航任务中的86%。在第二阶段,我们引入FlowDiT,一种流约束扩散Transformer,将优化的目标视频和语言指令转换为连续速度命令,使用动作空间去噪扩散。FlowDiT整合DINOv2视觉特征、学习的光流用于自身运动表示,以及CLIP语言嵌入用于语义停止。我们预训练在RECON户外导航数据集上,并在Isaac Sim中对203个Unitree G1人形机器人episode进行微调以校准速度动态。一个单个4300万参数检查点在仿真中实现73.2%的导航成功率,在未见过的室内环境中,真实Unitree G1机器人上实现64.7%的任务完成率,同时以40-47Hz的速度运行。我们评估了Action Agent在三种形态上:Unitree G1人形机器人(真实硬件)、无人机和轮式移动机器人(Isaac Sim),证明将轨迹想象与执行解耦可以产生一个可扩展且具有形态意识的语言引导导航范式。

英文摘要

We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In Stage I, a large language model (LLM) acts as an orchestration module that selects video diffusion models, refines prompts through iterative validation, and accumulates cross-task memory to synthesize physically plausible first-person navigation videos from language and image inputs. This increases video generation success from 35% (single-shot) to 86% across 50 navigation tasks. In Stage II, we introduce FlowDiT, a Flow-Constrained Diffusion Transformer that converts optimized goal videos and language instructions into continuous velocity commands using action-space denoising diffusion. FlowDiT integrates DINOv2 visual features, learned optical flow for ego-motion representation, and CLIP language embeddings for semantic stopping. We pretrain on the RECON outdoor navigation dataset and fine-tune on 203 Unitree G1 humanoid episodes collected in Isaac Sim to calibrate velocity dynamics. A single 43M-parameter checkpoint achieves 73.2% navigation success in simulation and 64.7% task completion on a real Unitree G1 in unseen indoor environments under open-loop execution, while operating at 40--47 Hz. We evaluate Action Agent across three embodiments: a Unitree G1 humanoid (real hardware), a drone, and a wheeled mobile robot (Isaac Sim), demonstrating that decoupling trajectory imagination from execution yields a scalable and embodiment-aware paradigm for language-guided navigation.

补充信息

↑