arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-19 至 2025-11-19 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 8 篇

2511.14631 2025-11-19 cs.CL cs.AI cs.CV cs.MA 84%

Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities

Kahaan Gandhi, Boris Bolliet, Inigo Zubeldia

机构 * Department of Physics, University of Cambridge, Cambridge, United Kingdom(剑桥大学物理系) Kavli Institute for Cosmology, University of Cambridge, Cambridge, United Kingdom(剑桥大学卡弗利天文研究所) Department of Physics and Astronomy, Haverford College, 370 Lancaster Avenue, Haverford, PA 19041, USA(哈弗福德学院物理与天文学系) Division of Physics, Mathematics and Astronomy, California Institute of Technology, Pasadena, CA 91125, USA(加州理工学院物理、数学与天文学系) Institute of Astronomy, University of Cambridge, Cambridge, United Kingdom(剑桥大学天文研究所)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14120 2025-11-19 cs.CV cs.AI 84%

Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning Framework with Vision-Language Models

Hao Zhen, Yunxiang Yang, Jidong J. Yang

机构 * College of Engineering University of Georgia(工程学院 乔治亚大学)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13782 2025-11-19 cs.AI 83%

Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models

Xiaoxing Lian, Aidong Yang, Jun Zhu, Peng Wang, Yue Zhang

专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract);分类 cs.AI

Comments 10 pages,a detail and effective benchmark for spatial reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14446 2025-11-19 cs.CV cs.AI 73%

Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding

Hong Gao, Yiming Bao, Xuezhen Tu, Yutong Xu, Yue Jin, Yiyang Mu, Bin Zhong, Linan Yue, Min-Ling Zhang

机构 * SouthEast University(东南大学) ZTE Corporation(中兴通讯)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11239 2025-11-19 cs.CV 70%

Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression

Zhongbin Guo, Jiahe Liu, Yushan Li, Wenyu Gao, Zhen Yang, Chenzhi Li, Xinyue Zhang, Ping Jian

机构 * School of Computer Science & Technology(计算机科学与技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14160 2025-11-19 cs.CV cs.AI cs.RO 66%

RynnEC: Bringing MLLMs into Embodied World

Ronghao Dang, Yuqian Yuan, Yunxuan Mao, Kehan Li, Jiangpin Liu, Zhikai Wang, Xin Li, Fan Wang, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎盘实验室) Zhejiang University(浙江大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI;MLLM(comments)

Comments The technical report of RynnEC, an embodied cognition MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13970 2025-11-19 cs.AI cs.CV 62%

Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios

Sanjay Acharjee, Abir Khan Ratul, Diego Patino, Md Nazmus Sakib

机构 * Ph.D. Student, Dept. of Civil Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Computer Sci. \& Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Civil Eng., University of Texas at Arlington. E-mail

专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14153 2025-11-19 cs.CV cs.AI 62%

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng, Rui Hu, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Li Yu, Wenyu Liu, Xinggang Wang

机构 * vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Code is released at https://github.com/hustvl/LENS

详情

展开后加载摘要…

URL PDF HTML 收藏