VIABench:一个从视障人士收集的用于视障辅助的综合视频基准测试
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
- Nanjing University(南京大学)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视障人士视觉信息获取难及多模态大语言模型在视障辅助中实用价值待探索的问题,引入VIABench视频基准测试,定义三个核心任务,提出严格测试流程,实验表明当前模型对视障人士支持不足,有望推动定制模型开发以改善其体验。
AI中文摘要:
视障人士因获取视觉信息有限而面临重大日常挑战。尽管多模态大语言模型在一般视觉和语言任务上取得了显著成果,但在实际视障辅助中的实用价值仍未充分探索。为填补这一空白,我们引入了VIABench,这是一个专门设计的综合视频基准测试,使用视障人士自己录制或分享的第一人称视频来评估多模态大语言模型在视障辅助场景中的表现。VIABench定义了三个核心任务,每个任务针对视觉辅助中的不同需求。主动提醒任务评估模型解释正在进行的视频内容并主动预测和口头描述即将到来的关键导航事件的能力;视觉问答任务评估模型回答用户关于视频中环境或物体问题的能力;视觉引导交互任务测试情境感知推理以完成用户与环境之间的有意交互。为确保进行稳健和公平的评估,我们提出了一个严格的基准测试流程,支持在线(实时)和离线设置。我们的实验表明,当前的多模态大语言模型仍难以对视障人士提供全面支持,尤其是在主动提醒任务中,该任务需要准确的预测和实时响应能力。我们希望VIABench能推动未来研究开发针对实际辅助的定制多模态大语言模型,最终改善视障人士的导航和交互体验。代码和数据将在这个https网址发布。
英文摘要:
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.