arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-13 至 2025-10-13 共收录 42 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2506.11034 2025-10-13 cs.LG cs.AI cs.CL 84%

CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models

Aneesh Komanduri, Karuna Bhaila, Xintao Wu

机构 * Department of Electrical Engineering and Computer Science University of Arkansas(电气工程与计算机科学系 奎萨克大学)

专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract);分类 cs.AI、cs.LG

Comments Accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025 Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08791 2025-10-13 cs.CV 79%

Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering

Yuanhao Zou, Zhaozheng Yin

机构 * University of Michigan(密歇根大学) Stony Brook University(石溪大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments CVPR2025 Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09507 2025-10-13 cs.CV cs.RO 70%

PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

Zixin Zhang, Kanghao Chen, Xingwang Lin, Lutao Jiang, Xu Zheng, Yuanhuiyi Lyu, Litao Guo, Yinchuan Li, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Beihang University(北航大学) Knowin

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08608 2025-10-13 cs.CL cs.AI 70%

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Agency for Science, Technology and Research, Singapore(新加坡科技研究局) Indian Institute of Technology Delhi(印度理工学院德里分校) Alibaba DAMO Academy(阿里巴巴达摩院) Microsoft Research Asia(微软亚洲研究院) Shanghai University of Finance and Economics(上海财经大学) Inner Mongolia University(内蒙古大学) Kyoto University(京都大学) Jiangxi Normal University(江西师范大学) Korea University(韩国大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09266 2025-10-13 cs.CL 50%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 视觉问答 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 5 篇

2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09361 2025-10-13 cs.CV 70%

BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

Junyan Ye, Dongzhi Jiang, Jun He, Baichuan Zhou, Zilong Huang, Zhiyuan Yan, Hongsheng Li, Conghui He, Weijia Li

机构 * Sun Yat-sen University(中山大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) CUHK MMLab(香港中文大学多模态实验室) Peking University(北京大学)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06944 2025-10-13 cs.LG cs.AI cs.CL cs.CV 67%

AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance

Lixuan He, Jie Feng, Yong Li

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09302 2025-10-13 cs.CV cs.AI cs.CL 62%

CapGeo: A Caption-Assisted Approach to Geometric Reasoning

Yuying Li, Siyi Qian, Hao Liang, Leqi Zheng, Ruichuan An, Yongzhen Guo, Wentao Zhang

机构 * THU(清华大学) PKU(北京大学) Ant Group(蚂蚁集团)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08964 2025-10-13 cs.CV cs.CL 57%

Unleashing Perception-Time Scaling to Multimodal Reasoning Models

Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou, Ruipu Luo, Can Zhang, Zhentao He, Yufei Zhan, Wayne Xin Zhao, Minghui Qiu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室) ByteDance(字节跳动) University of California, San Diego(加州大学圣地亚哥分校) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 11 篇

2502.03333 2025-10-13 cs.CV cs.AI 84%

RadVLM: A Multitask Conversational Vision-Language Model for Radiology

Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M. Sutter, Julia E. Vogt, Jonas Kluckert, Thomas Frauenfelder, Christian Blüthgen, Farhad Nooralahzadeh, Michael Krauthammer

机构 * Department of Radiology, Kobe University(金泽大学放射科) Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) Department of Advanced Imaging in Medical Magnetic Resonance, Kyoto University(京都大学医学磁共振高级成像部门) Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系) Diagnostic and Interventional Radiology, University Hospital Zurich(苏黎世大学医院诊断与介入放射科)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05404 2025-10-13 cs.CV cs.AI 81%

AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving

Lianming Huang, Haibo Hu, Yufei Cui, Jiacheng Zuo, Shangyu Wu, Nan Guan, Chun Jason Xue

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments We believe that the contribution of this paper is not enough, so we integrated it into another new paper. The arXiv ID of the new paper is arXiv:2510.01795

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 73%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08978 2025-10-13 cs.CV 70%

HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images

Zichuan Wang, Bo Peng, Songlin Yang, Zhenchen Tang, Jing Dong

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08775 2025-10-13 cs.CV cs.AI 62%

Re-Identifying Kākā with AI-Automated Video Key Frame Extraction

Paula Maddigan, Andrew Lensen, Rachael C. Shaw

机构 * Centre for Data Science and Artificial Intelligence, and School of Engineering and Computer Science(数据科学与人工智能中心,工程与计算机科学学院) Victoria University of Wellington(惠灵顿维多利亚大学) School of Biological Sciences(生物科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09421 2025-10-13 cs.CL cs.AI 57%

On the Representations of Entities in Auto-regressive Large Language Models

Victor Morand, Josiane Mothe, Benjamin Piwowarski

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted at BlackBoxNLP@EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09274 2025-10-13 cs.CV 57%

MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding

Ming Dai, Sen Yang, Boqiang Duan, Wankou Yang, Jingdong Wang

机构 * Southeast University(东南大学) Baidu VIS(百度视觉)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08849 2025-10-13 cs.CV 57%

FOLK: Fast Open-Vocabulary 3D Instance Segmentation via Label-guided Knowledge Distillation

Hongrui Wu, Zhicheng Gao, Jin Cao, Kelu Yao, Wen Shen, Zhihua Wei

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21787 2025-10-13 cs.CV cs.CL 57%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19114 2025-10-13 cs.CL cs.IR cs.LG 57%

Understanding and Improving Information Preservation in Prompt Compression for LLMs

Weronika Łajewska, Momchil Hardalov, Laura Aina, Neha Anna John, Hang Su, Lluís Màrquez

机构 * University of Stavanger(斯塔万格大学) AWS AI Labs(AWS AI实验室) Technical University of Catalonia (UPC)(加泰罗尼亚理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025 (Findings), 22 pages, 6 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10168 2025-10-13 stat.ME math.ST stat.TH 50%

Statistical methods: Basic concepts, interpretations, and cautions

Sander Greenland

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 64 pages. For Pigeot I, Ahrens W, eds., Handbook of Epidemiology, 3rd edn. Springer, 2025, Ch. 54-1

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 2 篇

2508.20582 2025-10-13 cs.IR 71%

SUMMA: A Multimodal Large Language Model for Advertisement Summarization

Weitao Jia, Shuo Yin, Zhoufutu Wen, Han Wang, Zehui Dai, Kun Zhang, Zhenyu Li, Tao Zeng, Xiaohui Lv

专题命中 文档图表理解 :multimodal large language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07545 2025-10-13 cs.CL cs.LG 70%

Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices

Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Ridwan Mahbub, Mizanur Rahman, Amran Bhuiyan, Israt Jahan, Mir Tafseer Nayeem, Shafiq Joty, Enamul Hoque, Jimmy Huang

机构 * York University(约克大学) University of Alberta(阿尔伯塔大学) Salesforce AI Research(Salesforce AI研究)

专题命中 文档图表理解 :vision-language model(abstract);LLaVA(abstract);分类 cs.LG

Comments Accepted to the EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 7 篇

2505.12493 2025-10-13 cs.AI 85%

GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning

Longxi Gao, Li Zhang, Pengzhi Gao, Wei Liu, Jian Luan, Mengwei Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Unaffiliated(无隶属机构)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 82%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 GUI与屏幕智能体 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08783 2025-10-13 cs.HC cs.AI 79%

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

Reuben A. Luera, Ryan Rossi, Franck Dernoncourt, Samyadeep Basu, Sungchul Kim, Subhojyoti Mukherjee, Puneet Mathur, Ruiyi Zhang, Jihyung Kil, Nedim Lipka, Seunghyun Yoon, Jiuxiang Gu, Zichao Wang, Cindy Xiong Bearfield, Branislav Kveton

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 GUI与屏幕智能体 :MLLM(title);multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01795 2025-10-13 cs.RO cs.AI 79%

Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving

Haibo Hu, Lianming Huang, Xinyu Wang, Yufei Cui, Shangyu Wu, Nan Guan, Chun Jason Xue

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Department of Computer Science, McGill University(麦吉尔大学计算机科学系) Department of Computer Science, Mohamed bin Zayed University of Artificial Intelligence(马尔代夫穆罕默德· bin·扎耶德人工智能大学计算机科学系)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09038 2025-10-13 cs.AI cs.CL cs.CV cs.CY cs.LG 67%

Auto-scaling Continuous Memory for GUI Agent

Wenyi Wu, Kun Zhou, Ruoxin Yuan, Vivian Yu, Stephen Wang, Zhiting Hu, Biwei Huang

机构 * University of California, San Diego(加州大学圣地亚哥分校) Fudan University(复旦大学) Abel AI

专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20711 2025-10-13 cs.HC cs.RO 67%

Automating eHMI Action Design with LLMs for Automated Vehicle Communication

Ding Xia, Xinyue Gui, Fan Gao, Dongyuan Li, Mark Colley, Takeo Igarashi

机构 * The University of Tokyo(东京大学) University College London(伦敦大学学院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract)

Comments Accepted as findings for EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19433 2025-10-13 cs.CV cs.AI cs.CL 62%

Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System

Lixuan He, Haoyu Dong, Zhenxing Chen, Yangcheng Yu, Jie Feng, Yong Li

专题命中 GUI与屏幕智能体 :MLLM(abstract);分类 cs.CV、cs.AI

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏