发表机构
Indian Institute of Technology Mandi; Mohamed bin Zayed University of Artificial Intelligence(印度理工学院曼迪分校; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对盲人和低视力用户,提出紧凑型VLM Smol-VL-BLV,通过蒸馏和GRPO后训练增强空间感知,在多个基准上显著提升性能,并可在移动设备离线运行。
AI 中文摘要
据估计,全球约有10亿人存在视力障碍,然而当前的视觉语言模型(VLMs)生成的描述过于模糊,无法为盲人和低视力(BLV)用户提供安全导航。大型VLM能够生成符合音频描述规范的高质量叙述,但无法在移动设备上运行;小型VLM具有竞争力的延迟,但缺乏空间细节、方向提示和危险意识,难以用于导航辅助。我们提出了Smol-VL-BLV,一个面向盲人和低视力用户的紧凑型VLM,通过500M解码器Transformer模型和两种后训练机制弥补了这一差距:(1)师生蒸馏和(2)组相对策略优化(GRPO),并使用针对方向性语言、度量距离和危险检测的复合BLV奖励。由于多阶段后训练可能引发灾难性遗忘,我们在最后一阶段GRPO微调之后增加了一个轻量级微调阶段,以恢复通用描述质量,同时保留BLV特定的空间基础。我们的最佳模型在各种基准测试中显著优于基线,包括VQA、BLV描述生成、OCR和延迟等任务。与基线相比,相对改进方面,空间分数提高了19.3%,社交分数提高了14.8%。此外,OCR-Bench提升了101.5%,TextVQA准确率提高了44.2%。这些结果表明,针对BLV的后训练既改善了特定于可访问性的空间基础,也提升了通用视觉-文本推理能力。通过混合精度量化部署在中端Android智能手机上,模型大小约为450 MB,完全在设备端运行,离线且无需网络依赖,生成的描述延迟取决于主机硬件能力。我们的模型、数据集和代码已在此https URL公开发布。
英文摘要
An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/
Comments14 pages, Accepted in EMNLP 2026