VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
机构 * Hong Kong University of Science and Technology(香港科技大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 视觉推理 :vision-language model(abstract)
Comments Project Page: https://vlm2-bench.github.io/ Camera Ready version