AI 中文总结
提出GapSight框架,利用损失间隙监督让VLMs选择性重读图像局部区域,在多个基准测试中显著提升了LLaVA、InternVL2.5等VLMs的细节类任务性能。
AI 中文摘要
视觉语言模型(VLMs)在以细节为中心的问题上表现不佳,具体原因是:答案虽在图像中可见,但在图像被压缩为低分辨率全局视图后丢失。为每个查询分配更多视觉标记可改善部分OCR和文档案例,但会不加区分地消耗计算资源,还可能干扰依赖全局上下文的任务。我们提出GapSight,一个用于学习视觉重读的框架:VLM先进行全局扫视,当问题需要局部证据时,再选择性返回自由格式区域。监督信号来自目标模型自身的失败信号。离线阶段,我们比较仅全局视图与候选裁剪增强视图下的答案损失或多项选择题选项边际;能提升目标答案的裁剪会成为特定模型的重读标签。轻量级自由格式裁剪路由将这些标签提炼为单次推理策略,可从全局状态预测是否需要重读、预期效用以及连续裁剪框。在LLaVA-1.5-7B、InternVL2.5-8B和Qwen2-VL-2B-Instruct上,GapSight在涵盖OCR、文档、图表、信息图、VStarBench和MME-RealWorld-Lite的六个基准测试中,相较于无缩放基线均有提升。在InternVL2.5-8B上,GapSight将六个基准的平均得分从52.25提升至64.29,优于CropVLM(57.16)、ViCrop(55.84)和ZoomRefine(54.43)。机制分析显示,该路由可修正具体错误答案、根据任务调整行动率,并形成有利的标记-性能曲线。这些结果表明,损失间隙监督是指导VLMs何时何地重新观察的实用途径。
英文摘要
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolution global view. Allocating more visual tokens to every query improves some OCR and document cases, but it spends computation indiscriminately and can disturb tasks that rely on global context. We propose GapSight, a framework for learning visual re-reading: a VLM first takes a global glance, then selectively returns to a free-form region when the question calls for local evidence. The supervision comes from the target model's own failure signal. Offline, we compare answer loss or multiple-choice option margin under a global-only view and candidate crop-augmented views; crops that improve the target answer become model-specific review labels. A lightweight free-form crop router distills these labels into a one-shot inference policy that predicts whether to review, expected utility, and a continuous crop box from the global state. Across LLaVA-1.5-7B, InternVL2.5-8B, and Qwen2-VL-2B-Instruct, GapSight improves the Base no-zoom baseline on six benchmarks spanning OCR, documents, charts, infographics, VStarBench, and MME-RealWorld-Lite. On InternVL2.5-8B, GapSight raises the six-benchmark average from 52.25 to 64.29, above CropVLM (57.16), ViCrop (55.84), and ZoomRefine (54.43). Mechanism analyses show that the router rescues concrete wrong answers, adapts its action rate by task, and forms a favorable token-performance profile. These results position loss-gap supervision as a practical route to teaching VLMs when and where to look again.