Q-CueGraph:用于多模态推理的查询条件化视觉证据图
Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning
浏览论文内容
中文总结 AI 辅助
Q-CueGraph是一种用于多模态推理的查询条件化视觉证据图,它为冻结读取器生成受预算约束的坐标级观测,在多个基准测试中显著提升了推理性能,尤其适用于证据可定位、问题能区分位置且分辨率受限的场景。
中文摘要 AI 辅助
高分辨率像素及裁剪或缩放工具赋予多模态大语言模型检查图像的能力,但这些工具并未提供可靠的、针对任务条件的检查位置决策策略。Q-CueGraph明确做出该决策,它将问题与图像表征映射为面向冻结读取器的、受预算约束的坐标级观测结果。文本丰富的图像使用可复用的OCR/布局图;自然图像搜索则在相同的选择、组合及预算约束接口后实例化查询条件化视觉节点。可选的效用优化模块从训练答案的正确性中学习冻结读取器可使用哪些候选裁剪,无需区域框监督。采用冻结的Qwen2.5-VL-7B读取器,Q-CueGraph在V*Bench上达到0.833的准确率,而使用19%图像面积预算的全图像推理准确率为0.696;在InfographicVQA上,Q-CueGraph仅使用约一半图像面积就达到全图像ANLS的92%。在六个基准测试中,当证据可定位、问题能区分证据位置且分辨率限制全图像读取时,显式观测最具价值。
英文摘要
Multimodal large language models (MLLMs) can miss fine details in a full image that they recognize in a closer view. Recovering this evidence requires deciding where to look and how much surrounding context to retain. We present Q-CueGraph, a query-conditioned evidence acquisition method for frozen MLLMs. For text-rich images, it builds a reusable graph of OCR lines and layout relations. Each question activates anchors, expands them into contextual regions, and selects candidates for a single observation window. Query-conditioned object detections support natural-image search through the same region-selection and composition interface. A lightweight candidate scorer further learns which observations support correct answers from frozen-reader feedback and training answers, without evidence-box supervision. Across six benchmarks, we examine the roles of query conditioning, evidence composition, and learned answerability. With Qwen2.5-VL-7B, Q-CueGraph raises V*Bench accuracy from 0.696 to 0.832 using 19.1% of source-image area, and retains 92% of full-image ANLS on InfographicVQA using about half the image area. The analyses show that useful evidence depends on both its relevance to the question and the context available to the reader. Q-CueGraph makes these choices explicit before answer generation.