AI 中文总结
该研究针对多模态大语言模型视觉问答中存在的物体幻觉问题,提出区域感知模型ReVA,通过双桥接结构融合整图与区域特征,在POPE数据集上将平均F1提升至82.85%,有效改善了视觉接地效果。
AI 中文摘要
多模态大语言模型(MLLM)在视觉问答(VQA)领域已取得显著进展,但在需要精确空间推理和细粒度视觉理解的问题上仍存在不足。这些不足常表现为物体、属性和空间层面的幻觉,即模型因缺乏区域级和细粒度视觉接地,生成自信但与视觉内容不符的响应。为应对这一挑战,本文提出ReVA,一种区域感知VQA模型,采用冻结的CLIP ViT-L/14视觉Transformer(ViT)和Qwen2.5-7B-Instruct大语言模型(LLM),通过双桥接结构连接两者,将整图和区域级表示与LLM的嵌入空间对齐。图像桥接将ViT最终Transformer块的特征映射为图像令牌;区域桥接则利用RAM++(识别任意内容模型)、spaCy和Grounding DINO构成的检测器栈,为每个边界框生成K个区域令牌,该检测器栈提供与问题无关和问题相关的自动零样本边界框,且区域桥接从ViT各层丰富的中间特征中裁剪特征,使早期纹理和后期物体线索更明显。图像令牌与区域令牌拼接为LLM提示前缀,用于在回答问题时联合编码场景级上下文和细粒度区域证据。在VQAv2、MMBench、POPE和SEED-Bench上评估,ReVA在POPE上的平均F1达82.85%,而无区域令牌的图像令牌基线为81.14%。这些结果表明,显式的区域感知视觉表示可减少物体幻觉,提升MLLM的事实接地能力。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as object, attribute, and spatial hallucinations, where models generate confident but visually unsupported responses due to insufficient region-level and fine-grained visual grounding. To address this challenge, we propose ReVA, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space. The image bridge maps final transformer block features into image tokens. The region bridge maps cropped features from enriched intermediate features across ViT blocks so early texture and later object cues are more evident, into K region tokens for every bounding box. ReVA uses a detector stack that supplies automatic zero-shot bounding boxes that are both question-agnostic and question-dependent, using RAM++ (Recognize Anything Model), spaCy, and Grounding DINO. The image tokens and region tokens are concatenated as an LLM prompt prefix to jointly encode scene-level context and fine-grained regional evidence when answering questions. Evaluated on VQAv2, MMBench, POPE, and SEED-Bench, ReVA achieves 82.85% mean F1 on POPE, compared with 81.14% for an image-token baseline without region tokens. These results demonstrate that explicit region-aware visual representations reduce object hallucination and improve the factual grounding of MLLMs.
Comments11 pages. Code: https://github.com/anoop675/reva