arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-19 至 2025-11-19 共收录 34 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2511.13892 2025-11-19 cs.AI 83%

Jailbreaking Large Vision Language Models in Intelligent Transportation Systems

Badhan Chandra Das, Md Tasnim Jawad, Md Jueal Mia, M. Hadi Amini, Yanzhao Wu

机构 * KFSCIS, Florida International University(凯斯-西储大学信息科学学院) Knight Foundation School of Computing and Information Sciences(骑士基金会计算与信息科学学院) Learning for InterDependent Networks Laboratory (solid lab)(依赖网络学习实验室)

专题命中 视觉问答 :vision language model(title,abstract);visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13135 2025-11-19 cs.CV 70%

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

Junjie Yang, Yuhao Yan, Gang Wu, Yuxuan Wang, Ruoyu Liang, Xinjie Jiang, Xiang Wan, Fenglei Fan, Yongquan Zhang, Feiwei Qin, Changmiao Wang

机构 * South China University of Technology(华南理工大学) Sun Yat-sen University(中山大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Finance & Economics(浙江财经大学) National University of Singapore(新加坡国立大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) City University of Hong Kong(香港城市大学)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments CVPR 2026 Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13881 2025-11-19 cs.CV 70%

VLMs Guided Interpretable Decision Making for Autonomous Driving

Xin Hu, Taotao Jing, Renran Tian, Zhengming Ding

机构 * Department of Computer Science, Tulane University(路易斯安那大学计算机科学系) Qualcomm(高通公司) Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 8 篇

2511.14631 2025-11-19 cs.CL cs.AI cs.CV cs.MA 84%

Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities

Kahaan Gandhi, Boris Bolliet, Inigo Zubeldia

机构 * Department of Physics, University of Cambridge, Cambridge, United Kingdom(剑桥大学物理系) Kavli Institute for Cosmology, University of Cambridge, Cambridge, United Kingdom(剑桥大学卡弗利天文研究所) Department of Physics and Astronomy, Haverford College, 370 Lancaster Avenue, Haverford, PA 19041, USA(哈弗福德学院物理与天文学系) Division of Physics, Mathematics and Astronomy, California Institute of Technology, Pasadena, CA 91125, USA(加州理工学院物理、数学与天文学系) Institute of Astronomy, University of Cambridge, Cambridge, United Kingdom(剑桥大学天文研究所)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14120 2025-11-19 cs.CV cs.AI 84%

Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning Framework with Vision-Language Models

Hao Zhen, Yunxiang Yang, Jidong J. Yang

机构 * College of Engineering University of Georgia(工程学院 乔治亚大学)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13782 2025-11-19 cs.AI 83%

Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models

Xiaoxing Lian, Aidong Yang, Jun Zhu, Peng Wang, Yue Zhang

专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract);分类 cs.AI

Comments 10 pages,a detail and effective benchmark for spatial reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14446 2025-11-19 cs.CV cs.AI 73%

Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding

Hong Gao, Yiming Bao, Xuezhen Tu, Yutong Xu, Yue Jin, Yiyang Mu, Bin Zhong, Linan Yue, Min-Ling Zhang

机构 * SouthEast University(东南大学) ZTE Corporation(中兴通讯)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11239 2025-11-19 cs.CV 70%

Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression

Zhongbin Guo, Jiahe Liu, Yushan Li, Wenyu Gao, Zhen Yang, Chenzhi Li, Xinyue Zhang, Ping Jian

机构 * School of Computer Science & Technology(计算机科学与技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14160 2025-11-19 cs.CV cs.AI cs.RO 66%

RynnEC: Bringing MLLMs into Embodied World

Ronghao Dang, Yuqian Yuan, Yunxuan Mao, Kehan Li, Jiangpin Liu, Zhikai Wang, Xin Li, Fan Wang, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎盘实验室) Zhejiang University(浙江大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI;MLLM(comments)

Comments The technical report of RynnEC, an embodied cognition MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13970 2025-11-19 cs.AI cs.CV 62%

Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios

Sanjay Acharjee, Abir Khan Ratul, Diego Patino, Md Nazmus Sakib

机构 * Ph.D. Student, Dept. of Civil Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Computer Sci. \& Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Civil Eng., University of Texas at Arlington. E-mail

专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14153 2025-11-19 cs.CV cs.AI 62%

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng, Rui Hu, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Li Yu, Wenyu Liu, Xinggang Wang

机构 * vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Code is released at https://github.com/hustvl/LENS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 10 篇

2412.15925 2025-11-19 cs.CV cs.AI 84%

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

机构 * Department of Electronics, Information, and Bioengineering(电子、信息与生物工程系) Polytechnic University of Milan(米兰理工学院) Department of Industrial, and Information Engineering(工业与信息工程系) University of Pavia(帕维亚大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref Moglia, A., Nastasio, E.C., Mainardi, L. et al. MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Observation and Localization in CT Images. J Healthc Inform Res (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14086 2025-11-19 cs.CV cs.AI cs.CL 81%

Error-Driven Scene Editing for 3D Grounding in Large Language Models

Yue Zhang, Zun Wang, Han Lin, Jialu Li, Jianing Yang, Yonatan Bitton, Idan Szpektor, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) University of Michigan(密歇根大学) Google Research(谷歌研究)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Code: https://github.com/zhangyuejoslin/Deer-3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13924 2025-11-19 cs.CV 79%

Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding

Qingyang Yan, Guangyao Chen, Yixiong Zou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13442 2025-11-19 cs.CV cs.AI 73%

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu

机构 * Zhejiang University(浙江大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05430 2025-11-19 cs.CV cs.AI cs.LG 67%

Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

机构 * University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学) LMU Munich(慕尼黑大学) MCML DFKI(德累斯顿大学) Bielefeld University(比勒菲尔德大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments NeurIPS 2025. Code: https://github.com/hbaniecki/fixlip

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14027 2025-11-19 cs.CL 67%

HiEAG: Evidence-Augmented Generation for Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13759 2025-11-19 cs.LG cs.AI 62%

Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection

Han Wang, Deyi Ji, Junyu Lu, Lanyun Zhu, Hailong Zhang, Haiyang Wu, Liqun Liu, Peng Shu, Roy Ka-Wei Lee

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 4 figures, Fortieth AAAI Conference on Artificial Intelligence (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19082 2025-11-19 cs.CV 57%

Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference

Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11777 2025-11-19 cs.RO cs.CV 57%

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

Vinit Mehta, Charu Sharma, Karthick Thiyagarajan

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) SensR Lab Western Sydney University(Western Sydney University SensR实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 45 pages, 15 figures, MDPI Sensors Journal

Journal ref Sensors 2025, 25(20), 6394

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09030 2025-11-19 cs.LG 57%

Contextual Learning for Anomaly Detection in Tabular Data

Spencer King, Zhilu Zhang, Ruofan Yu, Baris Coskun, Wei Ding, Qian Cui

机构 * Amazon Web Services, Seattle, WA, USA(亚马逊网络服务,西雅图,WA,USA)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to TMLR. 26 pages, 4 figures, 8 tables, 1 algorithm, 8 datasets, contextual anomaly detection framework for tabular data

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 1 篇

2511.14499 2025-11-19 cs.CV cs.RO 83%

Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM

Jack Qin, Zhitao Wang, Yinan Zheng, Keyu Chen, Yang Zhou, Yuanxin Zhong, Siyuan Cheng

机构 * Tsinghua University(清华大学) Laboratories, Huawei Technologies(华为技术有限公司2012实验室)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 3 篇

2509.07463 2025-11-19 cs.RO cs.AI cs.CV 81%

DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving

Sven Kirchner, Nils Purschke, Ross Greer, Alois C. Knoll

机构 * Chair of Robotics, Artificial Intelligence and Real-time Systems, Technical University of Munich(机器人学、人工智能与实时系统教授会,慕尼黑技术大学) Computer Science and Engineering Department, University of California Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 幻觉与鲁棒性 :vision-language model(title);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13378 2025-11-19 cs.CV 57%

Governance-Ready Small Language Models for Medical Imaging: Prompting, Abstention, and PACS Integration

Yiting Wang, Ziwei Wang, Di Zhu, Jiachen Zhong, Weiyi Li

机构 * Department of Data Science, University of Southern California(数据科学系,南加州大学) Department of Electrical and Computer Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学) Department of Computer Science and Engineering, Santa Clara University(计算机科学与工程系,圣克拉拉大学) Department of Applied Mathematics, University of Washington(应用数学系,华盛顿大学) School of Computer Science, Georgia Institute of Technology(计算机科学学院,佐治亚理工学院)

专题命中 幻觉与鲁棒性 :LLaVA(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16597 2025-11-19 cs.CV 57%

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Na Zhao, Zhiyu Tan, Hao Li, Xingjun Ma, Jingjing Chen

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 7 篇

2511.10701 2025-11-19 cs.CV cs.RO 79%

CARScenes: Semantic VLM Dataset for Safe Autonomous Driving

Yuankai He, Weisong Shi

机构 * st Yuankai He(第一作者) nd Weisong Shi(第二作者)

专题命中 VLM训练与架构 :VLM(title);vision-language model(abstract);分类 cs.CV

Comments 8 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01293 2025-11-19 cs.CV cs.AI 79%

GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification

Ngoc Bui Lam Quang, Nam Le Nguyen Binh, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Quan Nguyen, Ulas Bagci

机构 * AI VIETNAM(AI越南) Carnegie Mellon University(卡内基梅隆大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) PTIT Northwestern University(西北大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Acccepted in MICCAI Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14760 2025-11-19 cs.CV 70%

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Apple(苹果公司)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23982 2025-11-19 cs.CV cs.RO 70%

StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving

Ruiyang Hao, Bowen Jing, Haibao Yu, Zaiqing Nie

机构 * AIR, Tsinghua University(空气动力学研究所,清华大学) King’s College London(伦敦国王学院) The University of Manchester(曼彻斯特大学) The University of Hong Kong(香港大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 25 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10683 2025-11-19 cs.RO cs.AI cs.CV 62%

MotIF: Motion Instruction Fine-tuning

Minyoung Hwang, Joey Hejna, Dorsa Sadigh, Yonatan Bisk

机构 * Massachusetts Institute of Technology(麻省理工学院) Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏