arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-07 至 2025-08-07 共收录 45 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2508.04059 2025-08-07 cs.CV 83%

Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

Zhaochen Liu, Kaiwen Gao, Shuyi Liang, Bin Xiao, Limeng Qiao, Lin Ma, Tingting Jiang

专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04197 2025-08-07 cs.CV cs.AI 62%

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li, Yu Zhou, Can Ma, Xiaojun Bi

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China School of Cyber Science Engineering, Nanjing University of Science VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China Key Laboratory of Ethnic Language Intelligent Analysis Security Governance of MOE, Minzu University of China Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Security Governance of MOE, Minzu University of China

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by 2025 ACM MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04271 2025-08-07 cs.DC 50%

S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge

JinYi Yoon, JiHo Lee, Ting He, Nakjung Choi, Bo Ji

专题命中 视觉问答 :visual question answering(abstract)

Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 8 篇

2503.10905 2025-08-07 cs.AI cs.CV cs.LG 88%

Learning to Inference Adaptively for Multimodal Large Language Models

Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi, Somali Chaterji, Yingyu Liang, Yin Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) The University of Hong Kong(香港大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);LLaVA(abstract);visual reasoning(abstract);MLLM(abstract)

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03733 2025-08-07 cs.LG cs.AI cs.CL cs.CV 82%

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang

机构 * College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院) Shanghai Innovation Institute(上海创新研究院) Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心) School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院) Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科) MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室) Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院) Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系) Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院) X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室) Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科) Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心) Ningbo Institute of Technology, Beihang University(北航宁波理工学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03737 2025-08-07 cs.CL cs.AI 79%

GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models

Ashutosh Bandooni, Brindha Subburaj

专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI

Comments 6 pages, 3 figures. Accepted, Presented and Published as part of Proceedings of the 6th International Conference on Recent Advantages in Information Technology (RAIT) 2025

Journal ref 2025 6th International Conference on Recent Advances in Information Technology (RAIT), Dhanbad, India, 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04625 2025-08-07 cs.CV cs.CE 70%

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

Zichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang, Rongjin Li, Zihua Rong, Haoyang He, Zhuodi Hao, Xinyang Hu, Kun Ji, Ziyan Ma, Mengyuan Ji, Jun Zhang, Chenghao Ma, Qianhe Zheng, Yang Liu, Yiling Huang, Xinyi Hu, Qing Huang, Zijian Xie, Shiyao Peng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV 2025. arXiv admin note: text overlap with arXiv:2311.06602 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04453 2025-08-07 cs.CV 70%

Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion

Qingguo Hu, Ante Wang, Jia Song, Delai Qiu, Qingsong Liu, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Xiamen Unisound Intelligence Technology Co., Ltd(厦门Unisound智能科技有限公司) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),中华人民共和国文化和旅游部,中国)

专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03722 2025-08-07 cs.CV cs.AI 62%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 视觉推理 :MLLM(abstract);分类 cs.CV、cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03958 2025-08-07 cs.IR cs.AI cs.CL cs.LG 62%

A Comparative Study of Specialized LLMs as Dense Retrievers

Hengran Zhang, Keping Bi, Jiafeng Guo

机构 * Key Laboratory of Network Data Science and Technology(网络数据科学与技术重点实验室) Institute of Computing Technology(计算技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of AI Safety(人工智能安全国家重点实验室) University of Chinese Academy of Science(中国科学院大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI、cs.LG

Comments Accepted by CCIR25 and published by Springer LNCS or LNAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04043 2025-08-07 cs.CV 57%

VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning

Yuheng Ji, Yipu Wang, Yuyang Liu, Xiaoshuai Hao, Yue Liu, Yuting Zhao, Huaihai Lyu, Xiaolong Zheng

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 15 篇

2404.06798 2025-08-07 cs.CV 85%

Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding

Ke Zou, Yang Bai, Bo Liu, Yidi Chen, Zhihao Chen, Yang Zhou, Xuedong Yuan, Meng Wang, Xiaojing Shen, Xiaochun Cao, Yih Chung Tham, Huazhu Fu

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Department of Radiology, West China Hospital, Sichuan University(四川大学华西医院放射科) College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院) Department of Mathematics, Sichuan University(四川大学数学系) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院) Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore and the Singapore Eye Research Institute, Singapore National Eye Centre(新加坡国立大学 Yong Loo Lin 医学院眼科系及新加坡眼科研所、新加坡国家眼科中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17125 2025-08-07 cs.CV cs.AI 84%

DOGR: Towards Versatile Visual Document Grounding and Referring

Yinan Zhou, Yuxin Chen, Haokun Lin, Yichen Wu, Shuyu Yang, Zhongang Qi, Chen Ma, Li Zhu, Ying Shan

机构 * Xi’an Jiaotong University(西安交通大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) City University of Hongkong(香港城市大学) Institute of Automation, CAS(中国科学院自动化研究所) Harvard University(哈佛大学) vivo Mobile Communication Co.(vivo移动通信公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 22 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04572 2025-08-07 cs.CV 83%

Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding

Jun Li, Che Liu, Wenjia Bai, Mingxuan Liu, Rossella Arcucci, Cosmin I. Bercea, Julia A. Schnabel

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Imperial College London(伦敦帝国学院) University of Trento(特伦托大学) Helmholtz AI and Helmholtz Munich(海德堡人工智能与海德堡慕尼黑) King’s College London(伦敦国王学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04389 2025-08-07 cs.AI 83%

GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning

Weitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Minnesota(明尼苏达大学) Cisco Research(思科研究)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04546 2025-08-07 cs.CV 79%

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04299 2025-08-07 cs.CV 79%

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun, Jiawei Wang, Hongmin Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04175 2025-08-07 cs.CV 77%

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

Jingyi Liao, Yongyi Su, Rong-Cheng Tu, Zhao Jin, Wenhao Sun, Yiting Li, Dacheng Tao, Xun Xu, Xulei Yang

专题命中 视觉定位与Grounding :vision-language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03967 2025-08-07 cs.CV cs.CR cs.IR 70%

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid

机构 * Laboratory of IEMN, Univ. Polytechnique Hauts-de-France(IEMN实验室,法国高等技术法国大学) KU 6G Research Center, Khalifa University(KU 6G研究中心,哈利法大学) Sorbonne Center for Artificial Intelligence, Sorbonne University(人工智能研究中心,索邦大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04418 2025-08-07 cs.MM cs.CV cs.MA cs.SD eess.AS 57%

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

Jinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang, Xiaojun Chang, Hisham Cholakkal, Rao Muhammad Anwer

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Project page: https://github.com/jasongief/TGS-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04120 2025-08-07 cs.CV 57%

CLIPVehicle: A Unified Framework for Vision-based Vehicle Search

Likai Wang, Ruize Han, Xiangqun Zhang, Wei Feng

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Shenzhen University of Advanced Technology(深圳大学先进技术学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00099 2025-08-07 cs.CY cs.MA physics.soc-ph 50%

Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior

Egil Diau

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Position paper on LLM-agent simulation of financial structures. This update clarifies setup and adds a reciprocity-based table. Builds on arXiv:2505.02945 and 2505.08319

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04598 2025-08-07 cs.RO 50%

$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything

Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Xinyu Zheng, Pengwei Wang, Zhongyuan Wang, Wenbo Ding, Shanghang Zhang

专题命中 视觉定位与Grounding :VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04504 2025-08-07 cs.CY 50%

Moving beyond harm. A critical review of how NLP research approaches discrimination

Katrin Schulz, Marjolein Lanzing, Giulia Martinez Brenner

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04183 2025-08-07 cs.CL 50%

Characterizing Deep Research: A Benchmark and Formal Definition

Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi, Aaron Halfaker, Amit Deshpande, Navin Goyal, Ankur Gupta, Nagarajan Natarajan, Amit Sharma

机构 * Microsoft Research(微软研究院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments First three authors contributed equally (ordered alphabetically)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03830 2025-08-07 cs.PL 50%

If-T: A Benchmark for Type Narrowing

Hanwen Guo, Ben Greenman

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 2, Article 17

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2507.21167 2025-08-07 cs.CV cs.AI 62%

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * independent researcher(独立研究者)

专题命中 文档图表理解 :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 2 篇

2508.04280 2025-08-07 cs.LG cs.AI 84%

Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success

George Bredis, Stanislav Dereka, Viacheslav Sinii, Ruslan Rakhimov, Daniil Gavrilov

机构 * George Bredis(无) Stanislav Dereka(无) Viacheslav Sinii(无) Ruslan Rakhimov(无) Daniil Gavrilov(无)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04482 2025-08-07 cs.AI cs.CL cs.CV cs.LG 82%

OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) OPPO AI Center(OPPO AI中心) University of Chinese Academy of Sciences(中国科学院大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 GUI与屏幕智能体 :MLLM(title);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ACL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 5 篇

2411.05273 2025-08-07 cs.RO cs.AI cs.LG 84%

Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Sreyas Venkataraman, Yufei Wang, Ziyu Wang, Navin Sriram Ravie, Zackory Erickson, David Held

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈格浦尔分校) IIIS, Tsinghua University(清华大学人工智能研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

Comments 7 pages. Accepted at the LangRob Workshop 2024 @ CoRL, 2024. Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏