arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-08 至 2025-10-08 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 5 篇

2509.23250 2025-10-08 cs.AI cs.CV 79%

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

Brandon Ong, Tej Deep Pala, Vernon Toh, William Chandra Tjhi, Soujanya Poria

机构 * AI Singapore(AI新加坡) Nanyang Technological University(南洋理工大学)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06077 2025-10-08 cs.CV cs.AI 76%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视觉推理 :grounding(title);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06131 2025-10-08 cs.CV cs.AI 73%

Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation

Jiawei Mao, Yuhan Wang, Lifeng Chen, Can Zhao, Yucheng Tang, Dong Yang, Liangqiong Qu, Daguang Xu, Yuyin Zhou

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 16 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05593 2025-10-08 cs.CV cs.AI cs.CL 62%

Improving Chain-of-Thought Efficiency for Autoregressive Image Generation

Zeqi Gu, Markos Georgopoulos, Xiaoliang Dai, Marjan Ghazvininejad, Chu Wang, Felix Juefei-Xu, Kunpeng Li, Yujun Shi, Zecheng He, Zijian He, Jiawei Zhou, Abe Davis, Jialiang Wang

机构 * Meta Superintelligence Labs(Meta超智能实验室) Meta FAIR Cornell University(康奈尔大学) Stony Brook University(石溪大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07086 2025-10-08 cs.LG cs.CL 57%

A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility

Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, Matthias Bethge

机构 * Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.LG

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏