arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16476cs.RO

通过从野外视频进行可扩展学习揭示具身城市导航中的长尾分布

Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos

Bingyi Xia, Han Bao, Zhewei Chen, Hanjing Ye, Jingwen Yu, Yuhan Pang, Wenjun Xu, Jiankun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出可扩展框架,从网络规模野外 egocentric 视频学习点目标城市导航策略,自动注释视频并训练视觉-语言-动作策略,表征并诊断导航长尾分布与失败模式,验证了知识迁移效果并揭示长尾结构。

中文摘要 AI 辅助

从真实世界数据中学习具身城市导航策略受到任务特定数据收集成本以及罕见但对安全至关重要场景覆盖有限的限制。为应对这些挑战,我们提出一种可扩展框架,用于从网络规模的野外 egocentric 视频中学习点目标城市导航,同时系统地揭示其长尾分布。该框架用度量轨迹和结构化导航语义自动注释未经筛选的网络视频,随后用于训练用于可解释导航规划的视觉-语言-动作策略。我们基于模型性能和感知-运动模式的分布来表征长尾分布,并采用基于反思的分析来诊断反复出现的失败模式。在网络视频数据和真实世界城市导航任务上的实验表明,从无约束视频中实现了有效的知识迁移,且在整体导航性能之外揭示了连贯的长尾结构。

英文摘要

Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learning point-goal urban navigation from web-scale in-the-wild egocentric videos while systematically exposing its long tail. The framework automatically annotates uncurated web videos with metric trajectories and structured navigation semantics, which are then used to train a vision-language-action policy for interpretable navigation planning. We characterize the long tail based on model performance and the distribution of perception-motion patterns, and employ reflection-based analysis to diagnose recurring failure modes. Experiments on web-video data and real-world urban navigation tasks demonstrate effective knowledge transfer from unconstrained videos and reveal coherent long-tail structures beyond aggregate navigation performance.

↑