arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoFL-S: 用于局部语言条件导航的空间可查询扇区流场

CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa, Zicen Xiong, Jinjie Li, Moju Zhao

arXiv 2607.02222首次发表:更新:

发表机构

Dragon Lab, Department of Mechanical Engineering, The University of Tokyo(龙实验室,机械工程系,东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CoFL-S,一种低层级视觉-语言-动作框架,通过预测机器人局部可见扇区上的语言条件流场并滚动生成连续轨迹,在连续时间Habitat基准测试中优于动作令牌和动作块基线,并实现零样本真实世界部署。

AI 中文摘要

视觉-语言导航越来越强调高层指令推理、记忆、全局地图构建和指令分解,而低层动作表示相对未被充分探索。我们提出CoFL-S,一种低层级视觉-语言-动作框架,它预测机器人局部可见扇区上的语言条件流场,并通过滚动预测的场生成连续轨迹。为了训练这种低层表示,我们将每个VLN-CE片段(原本是整个片段的指令与动作序列配对)转换为帧级局部监督,带有对齐的子指令和匹配的动作、轨迹及密集流场目标。为了评估,我们引入了一个连续时间Habitat基准,它将低层动作接口与指令分解分离,并通过共享的速度-命令控制器执行所有方法,从而实现跨不同规划器频率的与分解无关的闭环比较,而不是VLN-CE中固定的离散前进-转向转换。在匹配的编码器和训练设置下,CoFL-S在连续时间Habitat基准测试中,在所有规划器频率下均一致优于动作令牌和动作块基线,而零样本真实世界闭环部署进一步展示了其在模拟之外的相对于两个基线的优势。

英文摘要

Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the predicted field. To train this low-level representation, we convert each VLN-CE episode, originally a whole-episode instruction paired with an action sequence, into frame-level local supervision with aligned sub-instructions and matched action, trajectory, and dense flow-field targets. For evaluation, we introduce a continuous-time Habitat benchmark that isolates low-level action interfaces from instruction decomposition and executes all methods through a shared velocity-command controller, enabling decomposition-independent closed-loop comparison across different planner frequencies rather than fixed discrete forward-and-turn transitions in VLN-CE. Under matched encoders and training settings, CoFL-S consistently outperforms baselines across planner frequencies in the continuous-time Habitat benchmark, and zero-shot real-world closed-loop deployment further shows its advantage over the evaluated baselines beyond simulation. See the project website at https://github.com/ut-dragon-lab/CoFL

Comments29 pages, 13 figures

Journal refConference on Robot Learning, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑