arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-08 至 2025-10-08 共收录 23 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 1 篇

2506.07966 2025-10-08 cs.CV 83%

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence

Ziyang Gong, Wenhao Li, Oliver Ma, Songyuan Li, Zhaokai Wang, Songyuan Li, Jiayi Ji, Xue Yang, Gen Luo, Junchi Yan, Rongrong Ji

机构 * SJTU(上海交通大学) XMU(厦门大学) Shanghai AI Lab(上海人工智能实验室) SYSU(南方科技大学) FDU(福建大学) NUS(新加坡国立大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 5 篇

2509.23250 2025-10-08 cs.AI cs.CV 79%

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

Brandon Ong, Tej Deep Pala, Vernon Toh, William Chandra Tjhi, Soujanya Poria

机构 * AI Singapore(AI新加坡) Nanyang Technological University(南洋理工大学)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06077 2025-10-08 cs.CV cs.AI 76%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视觉推理 :grounding(title);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06131 2025-10-08 cs.CV cs.AI 73%

Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation

Jiawei Mao, Yuhan Wang, Lifeng Chen, Can Zhao, Yucheng Tang, Dong Yang, Liangqiong Qu, Daguang Xu, Yuyin Zhou

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 16 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05593 2025-10-08 cs.CV cs.AI cs.CL 62%

Improving Chain-of-Thought Efficiency for Autoregressive Image Generation

Zeqi Gu, Markos Georgopoulos, Xiaoliang Dai, Marjan Ghazvininejad, Chu Wang, Felix Juefei-Xu, Kunpeng Li, Yujun Shi, Zecheng He, Zijian He, Jiawei Zhou, Abe Davis, Jialiang Wang

机构 * Meta Superintelligence Labs(Meta超智能实验室) Meta FAIR Cornell University(康奈尔大学) Stony Brook University(石溪大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07086 2025-10-08 cs.LG cs.CL 57%

A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility

Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, Matthias Bethge

机构 * Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.LG

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 7 篇

2510.05153 2025-10-08 cs.AI cs.IT math.IT 79%

An Algorithmic Information-Theoretic Perspective on the Symbol Grounding Problem

Zhangchi Liu

机构 * Zhangchi Liu(刘志强)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 7 pages, 1 table (in appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06040 2025-10-08 cs.CV cs.AI 76%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06056 2025-10-08 cs.AI 57%

Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research

Gang Liu, Yihan Zhu, Jie Chen, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学) MIT-IBM Watson AI Lab, IBM Research(麻省理工-IBM Watson AI实验室,IBM研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 25 pages, 17 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22304 2025-10-08 cs.CV 57%

CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction

Jiali Chen, Xusen Hei, HongFei Liu, Yuancheng Wei, Zikun Deng, Jiayuan Xie, Yi Cai, Li Qing

机构 * School of Software Engineering, South China University of Technology(华南理工大学软件学院) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05722 2025-10-08 cs.CV 57%

Data Factory with Minimal Human Effort Using VLMs

Jiaojiao Ye, Jiaxing Zhong, Qian Xie, Yuzhou Zhou, Niki Trigoni, Andrew Markham

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22688 2025-10-08 cs.CV 57%

Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization

Xu Jia

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05619 2025-10-08 eess.AS 50%

Teaching Machines to Speak Using Articulatory Control

Akshay Anand, Chenxu Guo, Cheol Jun Cho, Jiachen Lian, Gopala Anumanchipalli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 2 篇

2510.05586 2025-10-08 cs.CV 57%

CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval

Bin Kang, Bin Chen, Junjie Wang, Yulin Li, Junzhi Zhao, Zhuotao Tian

机构 * Chengdu Institute of Computer Applications, Chinese Academy of Sciences(成都计算机应用研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) International Research Institute for Artificial Intelligence, Harbin Institute of Technology (Shenzhen)(人工智能国际研究院,哈尔滨工业大学(深圳)) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Southwest Jiaotong University(西南交通大学) Tencent(腾讯)

专题命中 幻觉与鲁棒性 :visual language model(abstract);分类 cs.CV

Comments ACMMM2025(oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00055 2025-10-08 eess.IV cs.CV cs.CY 57%

Adapting Large Language Models to Mitigate Skin Tone Biases in Clinical Dermatology Tasks: A Mixed-Methods Study

Kiran Nijjer, Ryan Bui, Derek Jiu, Adnan Ahmed, Peter Wang, Kevin Zhu, Lilly Zhu

机构 * Stanford University(斯坦福大学) Algoverse AI Research(Algoverse人工智能研究)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments Accepted to EADV (European Academy of Dermatology) and SID (Society for Investigative Dermatology)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. VLM训练与架构 5 篇

2509.00192 2025-10-08 cs.CV 85%

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学) Department of Civil Environmental and Construction Engineering, University of Central Florida, USA(土木环境与建设工程系,中央佛罗里达大学) Department of Computer Science, University of Central Florida, USA(计算机科学系,中央佛罗里达大学)

专题命中 VLM训练与架构 :LLaVA(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05283 2025-10-08 cs.AI cs.CL cs.CV 81%

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

Radha Gulhane, Sathish Reddy Indurthi

机构 * Radha Gulhane(独立研究者) Sathish Reddy Indurthi(独立研究者)

专题命中 VLM训练与架构 :MLLM(title);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05836 2025-10-08 cs.CV 70%

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

Ruyang Liu, Shangkun Sun, Haoran Tang, Ge Li, Wei Gao

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, 2 Peng Cheng LaboratoryPeking University(1 电子与计算机工程学院,深圳研究生院,2 深圳鹏城实验室,北京大学)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV' 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21653 2025-10-08 cs.CV 70%

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

Ke Zhang, Cihan Xiao, Jiacong Xu, Yiqun Mei, Vishal M. Patel

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19252 2025-10-08 cs.CV 57%

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS2025. Website: https://sites.google.com/view/t2v-dlbs and Code: https://github.com/shim0114/T2V-Diffusion-Search

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他VLM 3 篇

2510.06064 2025-10-08 cs.CV cs.LG 81%

Medical Vision Language Models as Policies for Robotic Surgery

Akshay Muppidi, Martin Radfar

机构 * Department of Computer Science(计算机科学系) Stony Brook University(史泰尼斯布鲁克大学)

专题命中 其他VLM :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.LG

Comments IEEE CAI 2025

Journal ref 2025 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA, USA, 2025, pp. 513,518

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05538 2025-10-08 cs.CV cs.AI 73%

Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work

Owen Henkel, Bill Roberts, Doug Jaffe, Laurence Holt

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17132 2025-10-08 cs.AI 57%

Applications of Large Models in Medicine

YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏