AI 中文总结
该研究针对视障人士出行的信息过载问题,提出前瞻式导航框架ForeSightGuide,结合语义理解与危险预测,在新数据集及公共基准上实现最优性能,冗余警报与漏险率表现优异。
AI 中文摘要
电子助行设备对视障人士的独立出行至关重要。视觉语言模型(VLMs)能提供丰富的环境理解,但在动态场景中常出现过多误报,导致认知过载。为解决该问题,我们提出ForeSightGuide,一款将语义场景理解与预测性危险评估相结合的前瞻式辅助导航框架。与反应式系统不同,ForeSightGuide利用VLMs的推理能力预判障碍物运动,有效过滤非威胁性物体,提供简洁、可执行的导航指引。为验证我们的方法,我们引入了一个在复杂动态真实交通场景中采集的新型数据集,用于评估预测能力。在公共基准及我们提出的数据集上开展的大量实验表明,ForeSightGuide达到了当前最优性能。值得注意的是,它将每条导航指引的冗余警报降至0.299,同时保持0.112的低漏险率,显著减轻了信息过载,证明了其在安全步行辅助方面的有效性。
英文摘要
Electronic travel aids are pivotal for the independent mobility of the visually impaired. While Vision-Language Models (VLMs) offer rich environmental understanding, they often suffer from excessive false positives in dynamic scenarios, leading to cognitive overload. To address this, we present ForeSightGuide, an anticipatory assistive guidance framework that couples semantic scene understanding with predictive hazard assessment. Unlike reactive systems, ForeSightGuide leverages the reasoning capabilities of VLMs to anticipate obstacle motion, effectively filtering out non-threatening objects to provide concise, actionable guidance. To validate our approach, we introduce a novel dataset captured in complex, dynamic real-world traffic scenes, designed to benchmark predictive capabilities. Extensive experiments on both public benchmarks and our proposed dataset demonstrate that ForeSightGuide achieves state-of-the-art performance. Notably, it significantly mitigates information overload by reducing redundant alerts to 0.299 per guidance output while maintaining a low missed-hazard rate of 0.112, proving its efficacy for safe walking assistance.