arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2602.15650 2026-02-18 cs.CV 77%

Concept-Enhanced Multimodal RAG: Towards Interpretable and Accurate Radiology Report Generation

概念增强的多模态RAG:迈向可解释且准确的放射科报告生成

Marco Salmè, Federico Siciliano, Fabrizio Silvestri, Paolo Soda, Rosa Sicilia, Valerio Guarrasi

机构 * Department of Engineering(工程系) Research Unit of Artificial Intelligence and Computer Systems(人工智能与计算机系统研究单位) Università Campus Bio-Medico of Roma(罗马大学生物医学校园) Department of Computer, Control and Management Engineering(计算机、控制与管理工程系) Sapienza University of Rome(罗马萨皮恩扎大学) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入、辐射物理、生物医学工程系) Umeå University(乌梅拉大学) UniCamillus-Saint Camillus International University of Health Sciences(UniCamillus-圣卡米卢斯国际健康科学大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 概念增强的多模态RAG通过分解视觉表示为可解释的临床概念,提升放射科报告生成的可解释性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12260 2026-01-21 cs.AI 77%

Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding

Docs2Synth: 一种用于扫描视觉丰富文档理解的合成数据训练检索框架

Yihao Ding, Qiang Sun, Puzhen Wu, Sirui Li, Siwen Luo, Wei Liu

机构 * University of Western Australia(西澳大学) The University of Hong Kong(香港大学) Murdoch University(默多克大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

AI总结 Docs2Synth通过合成监督框架实现私有和低资源领域文档理解,利用检索引导推理提升接地能力和领域泛化,无需人工标注。

Comments Accepted at WWW 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02358 2026-01-19 cs.CV 77%

VINO: A Unified Visual Generator with Interleaved OmniModal Context

VINO:一个具有交错多模态上下文的统一视觉生成器

Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 VINO通过统一的视觉生成框架实现图像和视频的生成与编辑,利用交错多模态上下文条件处理,提升多任务视觉创作能力。

Comments Project page: https://sotamak1r.github.io/VINO-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04734 2026-01-09 cs.CV 77%

AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection

AIVD:自适应边缘-云协作用于准确高效的工业视觉检测

Yunqing Hu, Zheming Yang, Chang Zhao, Qi Guo, Meng Gao, Pengcheng Li, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries, Chinese Academy of Sciences(中国科学院工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :visual reasoning(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 AIVD通过边缘-云协作提升工业视觉检测的精度与效率,采用轻量边缘检测与云MLLM协同,结合高效微调策略和动态调度算法,实现高吞吐低延迟的资源优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15971 2025-12-19 cs.CV 77%

From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection

从词语到波长:用于少样本多光谱目标检测的视觉语言模型

Manuel Nkegoum, Minh-Tan Pham, Élisa Fromont, Bruno Avignon, Sébastien Lefèvre

机构 * Univ Bretagne Sud, IRISA, UMR 6074(布列塔尼大学,IRISA,UMR 6074) Univ Rennes, IRISA, UMR 6074(里尔大学,IRISA,UMR 6074) ATERMES UiT The Arctic University of Norway(北欧大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出利用视觉语言模型进行少样本多光谱目标检测,通过整合文本、视觉和热模态,在数据稀缺情况下实现高效检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08294 2025-12-11 cs.CV 77%

OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation

OpenSubject: 利用视频衍生的身份和多样性先验进行主体驱动的图像生成与操纵

Yexin Liu, Manyuan Zhang, Yueze Wang, Hongyu Li, Dian Zheng, Weiming Zhang, Changsheng Lu, Xunliang Cai, Yan Feng, Peng Pei, Harry Yang

机构 * HKUST(香港科技大学) Meituan(美团) HKUST(GZ)(香港科技大学(广州))

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 OpenSubject通过视频衍生的大规模数据集提升主体驱动图像生成与操纵的性能,尤其在复杂场景中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06663 2025-12-09 cs.CV 77%

CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks

CoT4Det:面向感知导向视觉-语言任务的链式思维框架

Yu Qi, Yumeng Zhang, Chenting Gong, Xiao Tan, Weiming Zhang, Wei Zhang, Jingdong Wang

机构 * Baidu Inc.(百度公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV

AI总结 CoT4Det通过将感知任务分解为分类、计数和定位三个步骤,提升视觉-语言模型在目标检测等任务中的性能,使mAP从19%提升至33%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01946 2025-12-09 cs.CV 77%

3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding

3DRS: MLLMs 需要 3D 意识表示监督以实现场景理解

Xiaohu Huang, Jingjing Wu, Qunyi Xie, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室) Department of Computer Vision Technology (VIS), Baidu Inc.(百度公司计算机视觉技术部)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 3DRS 通过引入预训练 3D 基础模型的监督,提升 MLLM 的 3D 表示能力,从而增强场景理解性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16412 2025-12-09 cs.CV cs.CE 77%

A Survey on Industrial Anomalies Synthesis

工业异常合成方法综述

Yanshu Wang, Xichen Xu, Jiaqi Liu, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu

机构 * Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) Department of Intelligent Manufacturing, CATL(CATL智能制造部门)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出工业异常合成方法的统一综述,引入首个分类法并探讨多模态学习的应用,为未来研究提供框架和路线图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19221 2025-12-01 cs.CV cs.RO 77%

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

Percept-WAM:感知增强的环境感知-行动模型用于鲁棒的端到端自动驾驶

Jianhua Han, Meng Tian, Jiangtong Zhu, Fan He, Huixin Zhang, Sitong Guo, Dechang Zhu, Hao Tang, Pei Xu, Yuze Guo, Minzhe Niu, Haojie Zhu, Qichao Dong, Xuechao Yan, Siyuan Dong, Lu Hou, Qingqiu Huang, Xiaosong Jia, Hang Xu

机构 * Yinwang Intelligent Technology Co. Ltd.(亿网通智能科技有限公司) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 Percept-WAM通过整合2D/3D场景理解能力,提升自动驾驶的感知与行动决策,实现端到端鲁棒性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11177 2025-11-26 cs.CV 77%

Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation

Viper-F1:基于交叉模态状态空间调制的高效细粒度多模态理解

Quoc-Huy Trinh

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 Viper-F1通过引入高效液态状态空间动力学和令牌-网格相关模块,实现了高效细粒度多模态理解。

Comments arXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 77%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02866 2025-11-04 cs.CV 77%

OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery

Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki

机构 * Department of Architecture, National University of Singapore(建筑系,新加坡国立大学) Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学) School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学) Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18262 2025-10-22 cs.CV 77%

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

Da Zhang, Chenggang Rong, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI)、中国电信)

专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV

Comments We have released V1, which only reports the test results. Our work is still ongoing, and the next version will be coming soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17332 2025-10-21 cs.CV 77%

iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA

Zhaoran Zhao, Xinli Yue, Jianhui Sun, Yuhao Xie, Tao Shao, Liangchao Yao, Fan Xia, Yuetang Deng

机构 * Tencent, WeChat(腾讯,微信)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22827 2025-10-21 cs.CV 77%

ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents

Yilei Jiang, Yaozhi Zheng, Yuxuan Wan, Jiaming Han, Qunzhong Wang, Michael R. Lyu, Xiangyu Yue

机构 * CUHK(香港中文大学) MMLab(多模态实验室) ARISE Lab(ARISE实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments ScreenCoder-v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14374 2025-10-17 cs.CV 77%

Spatial Preference Rewarding for MLLMs Spatial Understanding

Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory(上海人工智能实验室) Sensetime Research(商汤科技研究院) Zhejiang University of Technology(浙江工业大学) UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01954 2025-10-03 cs.CV 77%

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu, Jingyi Liao, Junyi Pan, Yuan Liu, Xiaofen Xing, Chong Sun, Chen Li, Nancy F. Chen, Shuicheng Yan, Xulei Yang, Xun Xu

机构 * South China University of Technology(华南理工大学) Institute for Infocomm Research (I 2 R), A*STAR(信息通信研究所(I 2 R),A*STAR) WeChat Vision, Tencent Inc.(微信视觉,腾讯公司) Foshan University(佛山大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments 24 pages, 12 figures and 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24791 2025-09-30 cs.CV 77%

Vision Function Layer in Multimodal LLMs

Cheng Shi, Yizhou Yu, Sibei Yang

机构 * Sun Yat-sen University(中山大学) School of Computing and Data Science(计算与数据科学学院) The University of Hong Kong(香港大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025 (preview; camera-ready in preparation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24298 2025-09-30 cs.HC cs.AI cs.CL cs.CY cs.MM 77%

Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports

Changde Du, Yizhuo Lu, Zhongyu Huang, Yi Sun, Zisen Zhou, Shaozheng Qin, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学) State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(认知神经科学与学习国家重点实验室,北京师范大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05976 2025-08-11 cs.CV cs.RO 77%

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation

Zhihao Zhu, Yifan Zheng, Siyu Pan, Yaohui Jin, Yao Mu

机构 * MoE key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能教育部重点实验室、人工智能研究院、上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 8 pages main paper, 8 figures, plus supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04175 2025-08-07 cs.CV 77%

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

Jingyi Liao, Yongyi Su, Rong-Cheng Tu, Zhao Jin, Wenhao Sun, Yiting Li, Dacheng Tao, Xun Xu, Xulei Yang

专题命中 视觉定位与Grounding :vision-language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01921 2025-08-05 cs.CV 77%

InspectVLM: Unified in Theory, Unreliable in Practice

Conor Wallace, Isaac Corley, Jonathan Lwowski

机构 * Zeitview

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to 2025 ICCV VISION Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19262 2025-07-28 cs.CV 77%

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

Monika Wysoczańska, Shyamal Buch, Anurag Arnab, Cordelia Schmid

机构 * Google DeepMind(谷歌DeepMind) Warsaw University of Technology(华沙技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02859 2025-07-04 cs.CV 77%

Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation

Jiaer Xia, Bingkui Tong, Yuhang Zang, Rui Shao, Kaiyang Zhou

机构 * Hong Kong Baptist University(香港 Baptist 大学) Sichuan University(四川大学) Shanghai AI Lab(上海人工智能实验室) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13337 2025-07-03 cs.CV 77%

Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

Qihang Fan, Huaibo Huang, Mingrui Chen, Ran He

机构 * MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(自动化研究所,中国科学院,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京)

专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24102 2025-07-01 cs.CV 77%

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Xiangtai Li, Tao Zhang, Yanwei Li, Haobo Yuan, Shihao Chen, Yikang Zhou, Jiahao Meng, Yueyi Sun, Shilin Xu, Lu Qi, Tianheng Cheng, Yi Lin, Zilong Huang, Wenhao Huang, Jiashi Feng, Guang Shi

机构 * ByteDance Seed(字节跳动种子) Wuhan University(武汉大学) Peking University(北京大学)

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Datasets and Models: https://github.com/lxtGH/DenseWorld-1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07519 2025-04-11 cs.CV 77%

VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong, Zechao Li

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15190 2025-04-08 cs.CV 77%

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D Watson, Levente J Klein, Fahad Shahbaz Khan, Salman Khan

专题命中 视觉定位与Grounding :vision-language model(abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03314 2025-03-28 cs.CV cs.CL cs.DB 77%

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan, Huangji Wang, Zhicai Wang, Shangwen Zhu, Han Zhang, Jie Xiao, Pingyu Wu, Kai Zhu, Jixuan Chen, Chen-Wei Xie, Yue Yang, Hongyang Zhang, Yu Liu, Fan Cheng

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏