arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28218cs.CV

聚焦关键之处:面向低视力辅助的显著性驱动视觉-语言模型

Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance

Jiazhao Liang, Hao Huang, Shuaihang Yuan, Congcong Wen, Geeta Chandra Raju Bethala, Giles Hamilton-Fletcher, Yu Hao, John-Ross Rizzo, Mengyu Wang, Anthony Tzes, Yi Fang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有视觉-语言模型未建模人类感知优先级的问题,提出显著性驱动的Salience-LLaVA模型,构建三个显著性数据集并引入SCMI评估,部署于辅助眼镜实现低视力辅助。

中文摘要 AI 辅助

视觉-语言模型(VLMs)发展迅速,为支持盲人和低视力人群的辅助技术提供了极具前景的能力。然而,现有VLMs主要针对通用图像描述任务设计,未明确建模人类感知优先级,从而限制了其强调场景中最相关信息的能力。为解决这一差距,我们提出了一种显著性驱动的图像描述框架,该框架根据以人类为中心的辅助需求对场景元素进行优先级排序。我们整理了三个具有显著性感知的数据集,即Salience COCO、Salience Flickr和Salience VizWiz,这些数据集包含物体级显著性标注,旨在反映不同环境中与低视力用户最相关的视觉信息。基于这些数据集,我们引入了Salience-LLaVA,这是一种具有显著性感知的VLM,它整合显著性线索以生成描述,其中重要元素按重要性顺序提及。我们的工作有四个主要贡献:构建了经低视力参与者验证的显著性感知数据集;提出了Salience-LLaVA,用于按重要性顺序描述物体;引入SCMI以评估排序准确性;并将该系统部署在辅助眼镜上,以展示其在现实世界中的实用性。代码和数据集可在以下网址获取:this https URL

英文摘要

Vision-language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing VLMs are primarily designed for general-purpose captioning and do not explicitly model human perceptual priorities, thereby limiting their ability to emphasize the most relevant information in a scene. To address this gap, we propose a salience-driven captioning framework that prioritizes scene elements according to their importance for human-centered assistance. We curate three salience-aware datasets, namely, Salience COCO, Salience Flickr, and Salience VizWiz, with object-level salience annotations designed to reflect the visual information most relevant to low vision users across different environments. Building on these datasets, we introduce Salience-LLaVA, a salience-aware VLM that incorporates salience cues to generate captions in which important elements are mentioned in the order of importance. Our work makes four main contributions. We build salience-aware datasets verified by low vision participants, propose Salience-LLaVA to describe objects in the order of importance, introduce SCMI to evaluate ordering accuracy, and deploy the system on assistive glasses to demonstrate real-world practicality. Code and datasets are available at: https://github.com/topo-focus/Topofocus

发表机构

  • New York University Tandon School of Engineering(纽约大学坦登工程学院)
  • New York University Abu Dhabi(纽约大学阿布扎比分校)
  • NYU Grossman School of Medicine(纽约大学格罗斯曼医学院)
  • NYU Langone Health(纽约大学兰贡医疗中心)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

↑