arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoRe:视觉语言模型中跨图像比较推理的综合框架

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

Lin Peng, Cong Wan, Zeyu Guo, SongLin Dong, Yihong Gong

arXiv 2607.12786首次发表:更新:

发表机构

Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言模型跨图像比较推理难题,提出CoRe框架,含自动构建的训练集CoRe-20K、结构化奖励框架TriSR及基准CoRe-Bench,实验显示其在CoRe-Bench上大幅超越现有模型,在标准基准上也有竞争力。

AI 中文摘要

跨图像比较推理对视觉语言模型(VLM)来说仍然具有挑战性,特别是在正确预测需要细粒度属性基础和全局一致推理时。我们提出了CoRe,一个针对此问题的统一框架。CoRe包括:通过多专家协作管道从结构化视觉元数据自动构建的大规模基于三元组的训练集CoRe-20K;在GRPO优化下联合监督属性基础、判断对齐和三元组一致性的结构化奖励框架TriSR;以及首个专门用于细粒度跨图像比较推理的基准CoRe-Bench。实验表明,CoRe在CoRe-Bench上显著优于现有VLM,在标准多模态基准上也具有竞争力,部分准确率比最强基线提高了28.2个百分点。

英文摘要

Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-20K, a large-scale triplet-based training set automatically constructed from structured visual metadata through a multi-expert collaborative pipeline, covering counting, depth, distance, and spatial relations; (ii) TriSR, a structured reward framework that jointly supervises attribute grounding, judgment alignment, and triplet consistency under GRPO optimization; and (iii) CoRe-Bench, the first benchmark dedicated to fine-grained cross-image comparative reasoning. Experiments show that CoRe substantially outperforms existing VLMs on CoRe-Bench while remaining competitive on standard multimodal benchmarks, achieving a 28.2-point gain in partial accuracy over the strongest baseline.

CommentsAccepted by ACMMM2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑