arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当提示词成为像素:面向多模态推理的提示词-区域定位

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang

arXiv 2608.04726首次发表:更新:

AI 中文总结

该研究针对多模态大语言模型在视觉化任务语义场景下的准确率下降问题,提出提示词-区域定位方法,有效提升了基准测试准确率且无需OCR或区域元数据。

AI 中文摘要

多模态大语言模型(MLLM)越来越多地对截图和文档进行推理,其中任务本身可能以像素形式呈现。然而,基准测试通常将问题设置为文本形式,因此尚不清楚模型在不同渠道是否同样能很好地使用相同指令。我们引入可视化任务语义(Visualized Task Semantics,VTS),这是一种受控干预方法,它在保持源问题和答案固定的情况下,将问题移入图像中。在六个多模态大语言模型和四个基准测试中,所有24个模型-任务对的准确率均出现下降,平均下降17.8个百分点。模型通常能正确转录视觉问题,但无法使用它,这暴露出超出光学字符识别(OCR)的语义渠道差距。为缩小这一差距,我们提出提示词-区域定位,其核心设计是将问题区域与类型化语义对齐,并从掩码视图中恢复其清晰表示。在匹配的训练成本下,我们的方法将四个基准测试的VTS准确率从58.0提升至66.3,同时保持原始界面的准确率,且推理时不需要OCR或区域元数据。读取承载任务的文本并将其定位为推理指令是不同的能力。

英文摘要

Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text, leaving it unclear whether models use the same instruction equally well across channels. We introduce Visualized Task Semantics (VTS), a controlled intervention that moves the question into the image while keeping the source problem and answer fixed. Across six MLLMs and four benchmarks, accuracy drops in all 24 model-task pairs, by 17.8 points on average. Models often transcribe the visual question correctly yet fail to use it, exposing a semantic channel gap beyond OCR. To reduce this gap, we present prompt-region grounding, whose core design aligns the question region with typed semantics and recovers its clean representation from a masked view. At matched training cost, our method raises four-benchmark VTS accuracy from 58.0 to 66.3 while preserving accuracy on the original interface, and requires no OCR or region metadata at inference. Reading task-bearing text and grounding it as an instruction for reasoning are distinct capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑