arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-27 至 2026-01-27 共收录 14 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 14 篇

2601.18356 2026-01-27 cs.LG 85%

Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning

通过检索增强的跨模态推理使医疗视觉-语言模型在多模态中实现因果推理

Weiqin Yang, Haowen Xue, Qingyi Peng, Hexuan Hu, Qian Huang, Tingbo Zhang

机构 * University of Adelaide(阿德莱德大学) Hohai University(河海大学) Amap

专题命中 视觉推理 :vision-language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.LG

AI总结 本文提出多模态因果检索增强生成框架,通过整合因果推理原理与多模态检索,提升医疗VLMs在诊断预测和视觉问答中的事实准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL 84%

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11020 2026-01-27 cs.CV cs.AI 81%

GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation

GeoVLMath: 通过跨模态奖励增强视觉-语言模型中的几何推理以辅助线创建

Shasha Guo, Liang Pang, Xi Wang, Yanling Wang, Huawei Shen, Jing Zhang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Renmin University of China(中国人民大学) Zhipu AI(智谱AI)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 GeoVLMath通过跨模态奖励模型提升视觉-语言模型在复杂立体几何问题中的几何推理能力。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18386 2026-01-27 cs.CV 70%

ARMOR: Agentic Reasoning for Methods Orchestration and Reparameterization for Robust Adversarial Attacks

ARMOR: 为对抗攻击的鲁棒性进行方法编排与重参数化中的代理推理

Gabriel Lee Jun Rong, Christos Korgialas, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.CV

AI总结 ARMOR通过视觉语言模型引导的代理协作,提升对抗攻击的鲁棒性和跨架构迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18619 2026-01-27 cs.AI 70%

Visual Attention Reasoning via Hierarchical Search and Self-Verification

通过分层搜索与自验证的视觉注意力推理

Wei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang, Ming Zhu, Haichuan Tang, Xuelong Li

专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出通过分层搜索与自验证的视觉注意力推理框架,有效提升多模态大语言模型的视觉定位和推理能力,显著降低幻觉发生率。

Comments The paper is withdrawn by the authors after discovering a flaw in the theoretical derivation presented in the Method section. This incorrect step leads to conclusions that are not supported by the corrected derivation. The authors plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17197 2026-01-27 cs.CL cs.LG 70%

Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding

超越字面:跨风格多模态推理用于隐喻语言理解

Seyyed Saeid Cheshmi, Hahnemann Ortiz, James Mooney, Dongyeop Kang

机构 * University of Minnesota(明尼苏达大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.LG

AI总结 本文提出一种三步框架,通过跨风格多模态推理提升隐喻语言理解能力,实验显示推理轨迹和跨风格训练能显著提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17123 2026-01-27 cs.HC cs.CV cs.RO 70%

Acoustic Field Video for Multimodal Scene Understanding

用于多模态场景理解的声学场视频

Daehwa Kim, Chris Harrison

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出声学场视频作为多模态场景理解的新输入方式,通过整合空间声学数据显著提升视觉-语言模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24330 2026-01-27 cs.CV 70%

SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

SenseNova-MARS: 通过强化学习赋能多模态代理推理与搜索

Yong Xien Chng, Tao Hu, Wenwen Tong, Xueheng Li, Jiandong Chen, Haojia Yu, Jiefan Lu, Hewei Guo, Hanming Deng, Chengjun Xie, Gao Huang, Dahua Lin, Lewei Lu

机构 * SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV

AI总结 SenseNova-MARS通过强化学习赋能多模态代理推理与搜索,提升视觉-语言模型在复杂视觉任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00555 2026-01-27 cs.LG cs.AI cs.CL cs.CV 67%

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

MMedAgent-RL: 优化多智能体协作以实现多模态医疗推理

Peng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng, Zihan Dong, Xian Wu, Xiangru Tang, Hongtu Zhu, Yun Li, Linjun Zhang, Shujie Liu, Yan Lu, Huaxiu Yao

机构 * UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Microsoft Research(微软研究院) CMU(卡内基梅隆大学) Rutgers University(罗格斯大学) Yale University(耶鲁大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MMedAgent-RL通过强化学习优化多智能体协作,提升多模态医疗推理性能,实现23.6%的性能提升。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15304 2026-01-27 cs.IR 67%

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

MLLMRec: 基于图细化的多模态推荐偏好推理范式

Yuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan, Wanyu Chen, Fei Cai, Honghui Chen

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract)

AI总结 MLLMRec通过图细化和多模态大语言模型提升多模态推荐的用户偏好推理与物品表示学习准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06835 2026-01-27 cs.CV 57%

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

X-LeBench:一个用于极长第一人称视频理解的基准数据集

Wenqi Zhou, Kai Cao, Hao Zheng, Yunze Liu, Xinyi Zheng, Miao Liu, Per Ola Kristensson, Walterio Mayol-Cuevas, Fan Zhang, Weizhe Lin, Junxiao Shen

机构 * University of Bristol(布里斯托大学) University of Manchester(曼彻斯特大学) University of Cambridge(剑桥大学) College of AI, Tsinghua University(清华大学人工智能学院) Meta Memories.ai Research(Memories.ai研究)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

AI总结 X-LeBench通过模拟真实日常生活场景,构建了首个极长第一人称视频理解基准数据集,揭示了长视频理解中的关键挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17706 2026-01-27 cs.CL cs.CV 57%

A Computational Approach to Visual Metonymy

一种视觉隐喻的计算方法

Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang

机构 * University of Cincinnati(辛辛那提大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种基于符号学理论的计算方法,通过构建ViMET数据集评估多模态语言模型在理解视觉隐喻方面的认知推理能力,并揭示了机器在处理间接视觉参考上的局限性。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17223 2026-01-27 cs.CL cs.AI 57%

Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning

超越结果验证:用于结构化推理的可验证过程奖励模型

Massimiliano Pronesti, Anya Belz, Yufang Hou

机构 * IBM Research Europe - Ireland(IBM欧洲研究院-爱尔兰) Dublin City University(都柏林城市大学) IT:U Interdisciplinary Transformation University Austria(IT:U跨学科转型大学奥地利)

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

AI总结 本文提出可验证过程奖励模型,用于提升结构化推理的连贯性和准确性,实验显示其在医学证据评估中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25975 2026-01-27 cs.CL cs.PL 50%

SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation

SymCode:通过可验证代码生成实现数学推理的神经符号方法

Sina Bagheri Nezhad, Yao Li, Ameeta Agrawal

机构 * Portland State University(波特兰州立大学) ElastixAI

专题命中 视觉推理 :grounding(abstract)

AI总结 SymCode通过可验证代码生成实现数学推理,显著提升准确性并增强模型的透明度和可靠性。

Comments camera-ready EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏