arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-16 至 2025-10-16 共收录 29 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 1 篇

2506.15298 2025-10-16 cs.CV cs.MM 85%

MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, Adrian K. Davision

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(计算与数学系,曼彻斯特 Metropolitan 大学) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院) Department of Psychology, University of the Chinese Academy of Sciences(心理学系,中国科学院大学) National Taiwan University(台湾大学) Zhejiang University(浙江大学) University of Oulu(奥卢大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Micro-Expression Grand Challenge (MEGC) at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 11 篇

2510.12845 2025-10-16 cs.CL cs.AI cs.CV cs.RO 86%

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Jesse Atuhurra, Iqra Ali, Tomoya Iwakura, Hidetaka Kamigaito, Tatsuya Hiraoka

专题命中 视觉推理 :VLM(title,abstract);vision language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05258 2025-10-16 cs.CV cs.LG 79%

Spatio-Temporal LLM: Reasoning about Environments and Actions

Haozhen Zheng, Beitong Tian, Mingyuan Wu, Zhenggang Tang, Klara Nahrstedt, Alex Schwing

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments Code and data are available at https://zoezheng126.github.io/STLLM-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14138 2025-10-16 cs.CV cs.AI 73%

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu, Lei Li, Jiahui Gao, Jiyue Jiang, Lingpeng Kong, Chuan Wu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03173 2025-10-16 cs.CL cs.AI cs.CV 73%

MULTI: Multimodal Understanding Leaderboard with Text and Images

Zichen Zhu, Yang Xu, Lu Chen, Jingkai Yang, Yichuan Ma, Yiming Sun, Hailin Wen, Jiaqi Liu, Jinyu Cai, Yingzi Ma, Situo Zhang, Zihan Zhao, Liangtai Sun, Kai Yu

机构 * X-LANCE Lab, School of Computer Science, Key Laboratory of Artificial Intelligence\ of Education, Shanghai Jiao Tong University, Shanghai 200240 , China Jiangsu Key Lab of Language Computing, Suzhou 215123 , China College of Computing Data Science, Nanyang Technological University, Singapore 639798 , Singapore Suzhou Laboratory, Suzhou 215123 , China

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 24 pages, 19 figures, 10 tables. Details and access are available at: https://OpenDFM.github.io/MULTI-Benchmark/

Journal ref Sci. China Inf. Sci. 68, 200107 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13375 2025-10-16 cs.CV 70%

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Zhuoguang Chen, Tao Jiang, Hang Zhao

机构 * IIIS, Tsinghua University(清华大学智能技术学部)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07882 2025-10-16 cs.RO 67%

Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots

Boyu Li, Siyuan He, Hang Xu, Haoqi Yuan, Xinrun Xu, Yu Zang, Liwei Hu, Junpeng Yue, Zhenxiong Jiang, Pengbo Hu, Börje F. Karlsson, Yehui Tang, Zongqing Lu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) AgiBot School of Computer Science, Peking University(北京大学计算机学院)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13804 2025-10-16 cs.CV cs.AI cs.CL 62%

Generative Universal Verifier as Multimodal Meta-Reasoner

Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao, Renrui Zhang, Ruihang Chu, Ling Yang, Yujiu Yang

机构 * Tsinghua University(清华大学) ByteDance Seed(字节跳动种子) Princeton University(普林斯顿大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12736 2025-10-16 cs.CV cs.AI 62%

Beyond Visual Appearances: Privacy-sensitive Objects Identification via Hybrid Graph Reasoning

Zhuohang Jiang, Bingkui Tong, Xia Du, Ahmed Alhammadi, Jizhe Zhou

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI

Comments I would like to formally request the withdrawal of my manuscript from arXiv. After a further internal review, I realized that the dataset used in this study contains personal or sensitive information that may inadvertently compromise individuals' privacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02803 2025-10-16 cs.CL cs.CV 57%

SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking

Sifan Li, Yujun Cai, Yiwei Wang

机构 * University of California, Merced(加州大学默塞德分校) University of Queensland(昆士兰大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23179 2025-10-16 cs.CV 57%

DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes

Sungjune Park, Hyunjun Kim, Junho Kim, Seongho Kim, Yong Man Ro

机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))

专题命中 视觉推理 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20740 2025-10-16 cs.AI 57%

MSEarth: A Multimodal Scientific Dataset and Benchmark for Phenomena Uncovering in Earth Science

Xiangyu Zhao, Wanghan Xu, Bo Liu, Yuhao Zhou, Fenghua Ling, Ben Fei, Xiaoyu Yue, Lei Bai, Wenlong Zhang, Xiao-Ming Wu

机构 * The Hong Kong Polytechnic University(香港理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 7 篇

2509.26165 2025-10-16 cs.CV 83%

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

Yuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang, Xiaozhong Ji, Qingdong He, Wenbin Wu, Donghao Luo, Zhenye Gan, Junwei Zhu, Yunhang Shen, Chaoyou Fu, Chengjie Wang, Xiaobin Hu, Shuicheng Yan

机构 * Yuan-Hou(元侯)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02477 2025-10-16 cs.RO cs.CV 79%

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng, Lue Fan, Dong An, Changwei Wang, Li Guo, Weiliang Meng, Xiaopeng Zhang, Rongtao Xu, Shibiao Xu

机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts Telecommunications, China [1ex] dKey Laboratory of Computing Power Network Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet Service Computing, Shandong Fundamental Research Center for Computer Science, China

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026

Journal ref Information Fusion, 126 (Part B), February 2026, 103652

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10912 2025-10-16 cs.RO 78%

More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks

Xinyu Shao, Yanzhe Tang, Pengwei Xie, Kaiwen Zhou, Yuzheng Zhuang, Xingyue Quan, Jianye Hao, Long Zeng, Xiu Li

机构 * Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院) Noah’s Ark Lab, Huawei, China(华为诺亚实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments More details and videos can be found at https://robo-map.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21976 2025-10-16 cs.CV cs.AI 62%

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, Xiang Li

机构 * College of Computer Science and Technology of Zhejiang University(浙江大学计算机科学与技术学院) Polytechnic Institute of Zhejiang University(浙江大学Polytechnic学院) Om AI Research(Om AI研究机构) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究机构) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) School of Mathematical Sciences of Zhejiang University(浙江大学数学科学学院) China Academy of Space Technology(中国航天科技研究院) University of Bristol(布里斯托大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19972 2025-10-16 cs.CV cs.AI cs.CL 62%

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

Seongheon Park, Sharon Li

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10300 2025-10-16 cs.CC cs.AI cs.IT cs.SY eess.SY math.IT q-bio.NC 57%

The Algorithmic Regulator

Giulio Ruffini

机构 * Neuroelectrics(神经电医学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 2 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02527 2025-10-16 cs.CL cs.LG 57%

I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2

Oliver McLaughlin, Arjun Khurana, Jack Merullo

机构 * Brown University(布朗大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 3 篇

2510.13054 2025-10-16 cs.RO cs.AI 70%

VLA-0: Building State-of-the-Art VLAs with Zero Modification

Ankit Goyal, Hugo Hadfield, Xuning Yang, Valts Blukis, Fabio Ramos

机构 * NVIDIA

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13778 2025-10-16 cs.RO cs.AI cs.CV 62%

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Xinyi Chen, Yilun Chen, Yanwei Fu, Ning Gao, Jiaya Jia, Weiyang Jin, Hao Li, Yao Mu, Jiangmiao Pang, Yu Qiao, Yang Tian, Bin Wang, Bolun Wang, Fangjing Wang, Hanqing Wang, Tai Wang, Ziqin Wang, Xueyuan Wei, Chao Wu, Shuai Yang, Jinhui Ye, Junqiu Yu, Jia Zeng, Jingjing Zhang, Jinyu Zhang, Shi Zhang, Feng Zheng, Bowen Zhou, Yangkun Zhu

机构 * Intern Robotics Shanghai AI Laboratory(Intern Robotics上海AI实验室)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV、cs.AI

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06207 2025-10-16 cs.RO 50%

EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model

Zefu Lin, Rongxu Cui, Chen Hanning, Xiangyu Wang, Junjia Xu, Xiaojuan Jin, Chen Wenbo, Hui Zhou, Lue Fan, Wenling Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) New Laboratory of Pattern Recognition (NLPR)(模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室) Beihang University(北航) Chinese University of Hong Kong(香港大学)

专题命中 GUI与屏幕智能体 :grounding(abstract)

Comments Demo Page: https://embodiedcoder.github.io/EmbodiedCoder/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 1 篇

2510.13190 2025-10-16 cs.CL 50%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 3 篇

2509.07613 2025-10-16 cs.CV 85%

Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease

Fangqi Cheng, Surajit Ray, Xiaochen Yang

机构 * School of Mathematics and Statistics, University of Glasgow, UK(数学与统计学学院,格拉斯哥大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted at MICAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13359 2025-10-16 cs.IR cs.CV cs.LG 84%

Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models

Yuki Yada, Sho Akiyama, Ryo Watanabe, Yuta Ueno, Yusuke Shido, Andre Rusli

机构 * Mercari, Inc.(Mercari公司)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to ACM RecSys 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 57%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他VLM 3 篇

2510.13276 2025-10-16 cs.CV cs.CL 79%

MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models

Keyan Zhou, Zecheng Tang, Lingfeng Ming, Guanghao Zhou, Qiguang Chen, Dan Qiao, Zheming Yang, Libo Qin, Minghui Qiu, Juntao Li, Min Zhang

机构 * Soochow University(苏州大学) ByteDance(字节跳动) Harbin Institute of Technology(哈尔滨工业大学) Central South University(中南大学)

专题命中 其他VLM :vision-language model(title);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13364 2025-10-16 cs.CV cs.AI 62%

Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity

MingZe Tang, Jubal Chandy Jacob

机构 * Department of Computing Science University of Aberdeen(计算科学系阿伯丁大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13211 2025-10-16 cs.CV cs.AI 62%

MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance

Subin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon, Gary Geunbae Lee

机构 * KT Corporation, Republic of Korea(韩国KT公司) Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(POSTECH人工智能研究生院) Computer Science and Engineering, POSTECH, Republic of Korea(POSTECH计算机科学与工程系)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏