面向嵌入式微型移动导航的轻量级机器学习驱动单目人行道路径提取
Lightweight Machine Learning-Driven Monocular Sidewalk Path Extraction for Embedded Micromobility Navigation
浏览论文内容
中文总结 AI 辅助
该研究针对嵌入式微型移动导航的人行道路径提取问题,提出历经三次迭代的单目视觉流水线,采用轻量级SegFormer-B0模型,通过对比规划方法,实现低耗时、高精度的路径提取,适配嵌入式设备。
中文摘要 AI 辅助
人行道尺度的路径提取要求在杂乱、地图稀疏的环境中,于紧凑、低功耗硬件上可靠运行感知与规划。我们提出一种用于微型移动系统的单目视觉人行道路径提取流水线,该流水线历经三次设计迭代:从骨架图基线,到距离变换走廊规划,再到轻量级图像空间架构,并系统比较了鸟瞰图(BEV)与图像空间域的五种路径规划方法。采用半监督师生框架训练的紧凑SegFormer-B0学生模型,使用OneFormer Swin-L伪标签训练,达到手动标注的交并比(IoU)为0.946,每帧处理时间11.7毫秒,优于基线检查点(IoU 0.758,18.9毫秒)。在32个手动标注帧的受控规划器比较中,图像空间中点规划实现最低横向中心误差(14.3像素),耗时2.2毫秒,较BEV距离变换规划(926.8毫秒,65.0像素中心误差)提速421倍,同时保持相当的掩码路径对齐度(98.5%对比98.6%)。在六个校园序列(22679帧)的全视频回放证实,改进后的分割将时间不稳定性从1.46%降至0.33%,并将模板路径可用性从73.7%提升至79.3%。我们进一步表明,仅BEV路径提取在单目场景中脆弱:在一次分析运行中,99.3%的帧未产生有效BEV路径。最终推荐架构为:图像空间中点规划为主,图像空间距离变换规划为后备,BEV仅用于可视化,该架构在CPU上运行完整感知到路径的堆栈每帧耗时不足50毫秒,适用于嵌入式行人速度的微型移动系统。
英文摘要
Sidewalk-scale path extraction demands perception and planning that run reliably on compact, low-power hardware in cluttered, map-sparse environments. We present a monocular vision pipeline for sidewalk path extraction in micromobility systems that progresses through three design iterations, from a skeleton-graph baseline through distance-transform corridor planning to a lightweight image-space architecture, and provides a systematic comparison of five path-planning methods across both bird's-eye-view (BEV) and image-space domains. A compact SegFormer-B0 student model, trained with a semi-supervised teacher-student framework using OneFormer Swin-L pseudo-labels, achieves a hand-annotated IoU of 0.946 at 11.7 ms per frame, improving over the baseline checkpoint (IoU 0.758, 18.9 ms). In a controlled planner comparison on 32 hand-labeled frames, image-space midpoint planning achieves the lowest lateral center error (14.3 px) at 2.2 ms, a 421x speedup over BEV distance-transform planning (926.8 ms, 65.0 px center error), while maintaining comparable mask-path alignment (98.5% versus 98.6%). A full-video replay across six campus sequences (22,679 frames) confirms that the improved segmentation reduces temporal instability from 1.46% to 0.33% and increases template-path availability from 73.7% to 79.3%. We further show that BEV-only path extraction is fragile in monocular settings: in one profiled run, 99.3% of frames produced no valid BEV path. The final recommended architecture, image-space midpoint primary, image-space distance-transform fallback, and BEV reserved for visualization, runs the full perception-to-path stack in under 50 ms per frame on CPU, making it suitable for embedded pedestrian-speed micromobility systems.
发表机构
- University of Oklahoma(俄克拉荷马大学)
- School of Electrical and Computer Engineering(电气与计算机工程学院)
- School of Aerospace and Mechanical Engineering(航空航天与机械工程学院)
机构由 AI 辅助整理,请以论文原文为准。