arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21828cs.HCcs.AI

Touvigation:面向盲人和低视力用户在陌生室内环境中的具身自适应物体获取

Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments

George Xi Wang, Xiangyu Li, Shaoyue Wen, Jiaqian Hu, Junan Xie, Yupeng Wang, Ziyue Shi, Qijun Chen, Maaike Bouwmeester, Yuhua Jin, Jing Qian

首次发表
浏览论文内容

中文总结 AI 辅助

针对盲人及低视力用户在陌生室内环境中的物体获取难题,提出结合视觉语言理解与持久空间建模的Touvigation系统,通过多阶段自适应引导,实现100%任务成功率,显著优于现有方法。

中文摘要 AI 辅助

盲人和低视力用户在陌生室内环境中定位并实际获取物体时常常面临挑战。现有的基于视觉语言模型的辅助系统能够提供语义描述,但可能引入延迟、幻觉以及与具身行动契合度不佳的引导。我们提出了Touvigation,一种免手持物体获取系统,它将视觉语言理解与持久局部空间建模相结合,以提供低延迟、基于身体相对位置的引导。基于对八名盲人和低视力参与者的形成性访谈,我们设计了一个多阶段引导框架,该框架在用户从定向、行走、到伸手及触觉验证的过渡过程中自适应调整空间参考。我们与多模态大语言模型辅助系统和无辅助搜索进行了对比,对12名盲人和低视力参与者评估了Touvigation。Touvigation实现了100%的任务成功率,而多模态辅助系统为58%,无辅助搜索为85%,同时减少了完成时间和认知负荷。我们的研究结果表明,持久空间接地和自适应具身引导能够改善盲人和低视力用户的物体获取。

英文摘要

Blind and low-vision users often face challenges when locating and physically acquiring objects in unfamiliar indoor environments. Existing vision-language-model-based assistants can provide semantic descriptions but may introduce latency, hallucinations, and guidance that is poorly aligned with embodied action. We present Touvigation, a hands-free object acquisition system that combines vision-language understanding with persistent local spatial modeling to provide low-latency, body-relative guidance. Drawing on formative interviews with eight blind and low-vision participants, we design a multi-stage guidance framework that adapts spatial references as users transition from orienting, to walking, to reaching and tactile verification. We evaluated Touvigation with 12 blind and low-vision participants against a multimodal large-language-model assistant and unassisted search. Touvigation achieved 100% task success, compared with 58% for the multimodal assistant and 85% for unassisted search, while reducing completion time and cognitive workload. Our findings demonstrate how persistent spatial grounding and adaptive embodied guidance can improve object acquisition for blind and low-vision users.

发表机构

  • Stony Brook University(石溪大学)
  • New York University(纽约大学)
  • Brown University(布朗大学)
  • Imperial College London(帝国理工学院)
  • Middlebury Institute of International Studies at Monterey(明德大学蒙特雷国际研究学院)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Tongji University(同济大学)
  • Shanghai Qibao Dwight High School(上海七宝德怀特高级中学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑