arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4524 篇

2509.11303 2025-09-30 cs.CL 50%

Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context

Dasol Choi, Jungwhan Kim, Guijin Son

专题命中 视觉推理 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06820 2025-09-30 cs.CL cs.SD eess.AS 50%

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

Wenyu Zhang, Yingxu He, Geyu Lin, Zhuohan Liu, Shuo Sun, Bin Wang, Xunlong Zou, Jeremy H. M. Wong, Qiongqiong Wang, Hardik B. Sailor, Nancy F. Chen, Ai Ti Aw

机构 * Institute for Infocomm Research (I 2 {}^{\text{2}} R), Agency for Science, Technology and Research (A*STAR)(信息与通信研究机构(I2R),科技研究局(A*STAR)) Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(前沿人工智能研究中心(CFAR),科技研究局(A*STAR))

专题命中 视觉推理 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13888 2025-09-30 cs.RO 50%

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

Ji Zhang, Shihan Wu, Xu Luo, Hao Wu, Lianli Gao, Heng Tao Shen, Jingkuan Song

机构 * Southwest Jiaotong University(西南交通大学) University of Electronic Science and Technology of China(电子科技大学) Tongji University(同济大学)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23698 2025-09-30 cs.CL 50%

VIVA+: Human-Centered Situational Decision-Making

Zhe Hu, Yixiao Ren, Guanzhong Liu, Jing Li, Yu Yin

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)

专题命中 视觉推理 :multimodal large language model(abstract)

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04059 2025-09-29 cs.CL 50%

Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning

Zhilin Wang, Zhe Yang, Yun Luo, Yafu Li, Xiaoye Qu, Ziqian Qiao, Haoran Zhang, Runzhe Zhan, Derek F. Wong, Jizhe Zhou, Yu Cheng

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) Sichuan University(四川大学) Shanghai Jiao Tong University(上海交通大学) University of Macau(澳门大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉推理 :multimodal large language model(abstract)

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19274 2025-09-24 cs.CL cs.MM 50%

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha

机构 * Indian Institute of Technology Patna(印度理工学院帕纳巴分校) Banasthali Vidyapeeth University(班纳萨利大学) Pandit Deendayal Energy University(德英德能源大学) Manipal University Jaipur(马哈拉施特拉邦大学贾伊普尔分校) Dwarkadas J. Sanghvi College of Engineering(德瓦尔卡斯J.桑格维工程学院)

专题命中 视觉推理 :vision-language model(abstract)

Comments EMNLP MAINS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15717 2025-09-22 cs.RO 50%

Imagination at Inference: Synthesizing In-Hand Views for Robust Visuomotor Policy Inference

Haoran Ding, Anqing Duan, Zezhou Sun, Dezhen Song, Yoshihiko Nakamura

机构 * Department of Robotics, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(机器人系,Mohamed bin Zayed人工智能大学)

专题命中 视觉推理 :visual reasoning(abstract)

Comments Submitted to IEEE for possible publication, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07917 2025-09-19 cs.RO 50%

MolmoAct: Action Reasoning Models that can Reason in Space

Jason Lee, Jiafei Duan, Haoquan Fang, Yuquan Deng, Shuo Liu, Boyang Li, Bohan Fang, Jieyu Zhang, Yi Ru Wang, Sangho Lee, Winson Han, Wilbert Pumacay, Angelica Wu, Rose Hendrix, Karen Farley, Eli VanderBilt, Ali Farhadi, Dieter Fox, Ranjay Krishna

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Washington(华盛顿大学)

专题命中 视觉推理 :grounding(abstract)

Comments Updated GR00T result to N1.5

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14581 2025-09-11 cs.MM eess.IV 50%

Memory-Anchored Multimodal Reasoning for Explainable Video Forensics

Chen Chen, Runze Li, Zejun Zhang, Pukun Zhao, Fanqing Zhou, Longxiang Wang, Haojian Huang

专题命中 视觉推理 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06661 2025-09-03 cs.RO 50%

Domain-Conditioned Scene Graphs for State-Grounded Task Planning

Jonas Herzog, Jiangpin Liu, Yue Wang

机构 * Zhejiang University(浙江大学)

专题命中 视觉推理 :grounding(abstract)

Comments Accepted for IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19283 2025-08-28 cs.CR 50%

Rethinking Denial-of-Service: A Conditional Taxonomy Unifying Availability and Sustainability Threats

Mark Dorsett, Scott Man, Tim Koussas

专题命中 视觉推理 :grounding(abstract)

Comments 7 pages, 3 figures, 3 tables,

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15854 2025-08-25 cs.CL 50%

QU-NLP at QIAS 2025 Shared Task: A Two-Phase LLM Fine-Tuning and Retrieval-Augmented Generation Approach for Islamic Inheritance Reasoning

Mohammad AL-Smadi

机构 * Qatar University(卡塔尔大学)

专题命中 视觉推理 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15164 2025-08-22 cs.CL 50%

ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following

Seungmin Han, Haeun Kwon, Ji-jun Park, Taeyang Yoon

机构 * Dongguk University(东国大学)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14941 2025-08-22 cs.MM cs.CL 50%

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视觉推理 :grounding(abstract)

Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14425 2025-08-19 cs.CL 50%

From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning

Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

机构 * Computational Linguistics, Department of Linguistics University of Potsdam(乌特雷赫特大学语言学系计算语言学部) German Research Center for Artificial Intelligence (DFKI), Berlin(德国人工智能研究中心(DFKI)柏林)

专题命中 视觉推理 :grounding(abstract)

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09804 2025-08-14 cs.CL 50%

BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning

Ahmed Masry, Abhay Puri, Masoud Hashemi, Juan A. Rodriguez, Megh Thakkar, Khyati Mahajan, Vikas Yadav, Sathwik Tejaswi Madhusudhan, Alexandre Piché, Dzmitry Bahdanau, Christopher Pal, David Vazquez, Enamul Hoque, Perouz Taslakian, Sai Rajeswar, Spandana Gella

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09641 2025-08-14 cs.CE 50%

VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

Zhaowei Liu, Xin Guo, Haotian Xia, Lingfeng Zeng, Fangqi Lou, Jinyi Niu, Mengping Li, Qi Qi, Jiahuan Li, Wei Zhang, Yinglong Wang, Weige Cai, Weining Shen, Liwen Zhang

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04088 2025-08-08 cs.CL 50%

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

Jianghangfan Zhang, Yibo Yan, Kening Zheng, Xin Zou, Song Dai, Xuming Hu

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01715 2025-08-05 cs.RO 50%

Towards Zero-Shot Terrain Traversability Estimation: Challenges and Opportunities

Ida Germann, Mark O. Mints, Peer Neubert

机构 * Intelligent Autonomous Systems Group, Institute of Computational Visualisitics, University of Koblenz(智能自主系统组,计算可视化研究所,科隆大学)

专题命中 视觉推理 :vision-language model(abstract)

Comments Accepted and presented at the 1st German Robotics Conference (GRC); March 13-15, 2025, Nuremberg, Germany https://ras.papercept.net/conferences/conferences/GRC25/program/GRC25_ContentListWeb_3.html#sada_48

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23135 2025-08-01 cs.CL 50%

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Ananya Sadana, Yash Kumar Lal, Jiawei Zhou

机构 * Stony Brook University(石溪大学)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19885 2025-07-29 cs.CL 50%

Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam

Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages, Adriano Jose Pereira, Uri Adrian Prync Flato, Eduardo Pontes dos Reis, Joaquim Edson Vieira, Paulo Sergio Panse Silveira, Edson Amaro Junior

机构 * Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨) Hospital Israelita Albert Einstein(埃因斯坦医院) Departamento de Pacientes Graves(重症患者部门) Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心) Departmento de Cirurgia(外科部门) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 50%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08224 2025-07-22 cs.RO 50%

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

Chan Young Park, Jillian Fisher, Marius Memmel, Dipika Khullar, Seoho Yun, Abhishek Gupta, Yejin Choi

专题命中 视觉推理 :vision-language model(abstract)

Comments Code Available: https://github.com/chan0park/SelfReVision

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13396 2025-07-21 cs.IR cs.CL 50%

DyG-RAG: Dynamic Graph Retrieval-Augmented Generation with Event-Centric Reasoning

Qingyun Sun, Jiaqi Yuan, Shan He, Xiao Guan, Haonan Yuan, Xingcheng Fu, Jianxin Li, Philip S. Yu

机构 * Beihang University(北航大学) Guangxi Normal University(广西师范大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 视觉推理 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07562 2025-07-11 cs.CL 50%

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Jierun Chen, Tiezheng Yu, Haoli Bai, Lewei Yao, Jiannan Wu, Kaican Li, Fei Mi, Chaofan Tao, Lei Zhu, Manyi Zhang, Xiaohui Li, Lu Hou, Lifeng Shang, Qun Liu

机构 * Huawei Technologies(华为技术有限公司) HKUST(香港科技大学)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12084 2025-07-03 cs.CL 50%

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Jianshu Zhang, Dongyu Yao, Renjie Pi, Paul Pu Liang, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港科技大学) Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉推理 :vision-language model(abstract)

Comments Project Page: https://vlm2-bench.github.io/ Camera Ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21762 2025-06-30 cs.HC 50%

ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues

Oliver Huang, Carolina Nobre

专题命中 视觉推理 :vision-language model(abstract)

Comments VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16856 2025-06-30 cs.CL 50%

MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers

Yang Tian, Zheng Lu, Mingqi Gao, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) SCUT SEU BAAI(北京人工智能研究院)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18102 2025-06-24 cs.CL 50%

InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating

Fuyu Wang, Jiangtong Li, Kun Zhu, Changjun Jiang

机构 * Key Laboratory of Embedded System and Service Computing, Ministry of Education, Tongji University(嵌入式系统与服务计算重点实验室,教育部,同济大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

专题命中 视觉推理 :grounding(abstract)

Comments 20 pages; Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11080 2025-06-16 cs.CL 50%

MANBench: Is Your Multimodal Model Smarter than Human?

Han Zhou, Qitong Xu, Yiheng Dong, Xin Yang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通讯学院,华中科技大学) School of Basic Medicine, Huazhong University of Science and Technology(基础医学院,华中科技大学)

专题命中 视觉推理 :multimodal large language model(abstract)

Comments Multimodal Benchmark, Project Url: https://github.com/micdz/MANBench, ACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏