发表机构
Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学模式识别实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对智能手机辅助导航中区分人行道与不安全区域的需求,提出安全导向语义分割框架,引入数据集并评估多种分割架构,通过多指标评估模型,结果显示应综合评估辅助人行道感知,有助于选择平衡多方面因素的模型。
AI 中文摘要
对于盲人和视力受损行人(BVIP)来说,独立的人行道移动性至关重要,而基于智能手机的辅助导航需要能区分可步行人行道和相邻不安全区域的感知模型。本研究提出了一个面向未来移动引导的安全导向语义分割框架。引入了SENSATION-DS数据集,评估了五种分割架构。通过多种指标评估模型,合成增强通常提高分割精度,SAM2伪标签更持续减少错误。UPerNet-MobileNetV3离线mIoU最高,DeepLabV3Plus-MobileNetV3道路误判率最低且安卓运行时性能最高。结果表明辅助人行道感知应综合评估,现实效益需经BVIP用户验证,该评估有助于选择平衡多种因素的模型。
英文摘要
Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perception models that distinguish walkable sidewalks from adjacent unsafe regions. This study presents a safety-oriented semantic segmentation framework for future mobile guidance. We introduce SENSATION-DS, a chest-height pedestrian-view dataset with 2,752 image-mask pairs and nine-class navigation-relevant taxonomy. External urban and sidewalk datasets were harmonized to this label space, and five segmentation architectures were evaluated using staged target-domain adaptation with mask-conditioned synthetic images and Segment Anything Model 2 (SAM2) pseudo-labels. Models were assessed using mean Intersection over Union (mIoU), road- and sidewalk-specific metrics, Road-as-Sidewalk Error Rate as a proxy false-safe measure, and Android Open Neural Network Exchange benchmarking. Synthetic augmentation generally improved segmentation accuracy, whereas SAM2 pseudo-labels more consistently reduced Road-as-Sidewalk errors. UPerNet-MobileNetV3 achieved the highest offline mIoU (0.715 +/- 0.006), while DeepLabV3Plus-MobileNetV3 achieved the lowest Road-as-Sidewalk Error Rate (0.079) and highest Android runtime at 512x384 (7.383 FPS). These results show that assistive sidewalk perception should be evaluated jointly by segmentation accuracy, proxy false-safe behavior, and smartphone deployment feasibility, while real-world benefit requires validation with BVIP users. This evaluation supports selecting models that balance accurate perception, conservative error behavior, and practical runtime.
Comments17 pages, 4 figures, 3 tables. Submitted to Assistive Technology