arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2507.07939 2025-07-23 cs.CL 85%

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

Guoxin Zang, Xue Li, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song, Lei Fan

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of New South Wales(新南威尔士大学)

专题命中 视觉定位与Grounding :visual language model(title);vision-language model(abstract);VLM(abstract);visual reasoning(abstract)

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14538 2025-04-02 eess.IV cs.AI cs.CV cs.LG 85%

Vision-Language Models for Acute Tuberculosis Diagnosis: A Multimodal Approach Combining Imaging and Clinical Data

Ananya Ganapthy, Praveen Shastry, Naveen Kumarasami, Anandakumar D, Keerthana R, Mounigasri M, Varshinipriya M, Kishore Prasath Venkatesh, Bargava Subramanian, Kalyan Sivasailam

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08722 2025-03-18 cs.CV cs.AI cs.LG 85%

A Recipe for Improving Remote Sensing VLM Zero Shot Generalization

Aviad Barzilai, Yotam Gigi, Amr Helmy, Vered Silverman, Yehonathan Refael, Bolous Jaber, Tomer Shekel, George Leifman, Genady Beryozkin

专题命中 视觉定位与Grounding :VLM(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06912 2025-03-04 cs.CV cs.AI cs.LG 85%

Compositional Entailment Learning for Hyperbolic Vision-Language Models

Avik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, Pascal Mettes

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted as oral paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04559 2024-10-28 cs.CL cs.AI cs.CV cs.LG 85%

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments In proceedings of EMNLP 2024 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19696 2024-05-01 cs.CV cs.AI cs.CL cs.LG 85%

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

Chun Feng, Joy Hsu, Weiyu Liu, Jiajun Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR 2024. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10093 2023-11-08 cs.CV cs.AI cs.CL cs.LG 85%

Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection

Kyle Buettner, Adriana Kovashka

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at Winter Conference on Applications of Computer Vision (WACV), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.06403 2022-07-14 cs.CV cs.AI cs.CL cs.GR cs.LG 85%

3D Concept Grounding on Neural Fields

Yining Hong, Yilun Du, Chunru Lin, Joshua B. Tenenbaum, Chuang Gan

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project page: http://3d-cg.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15925 2025-11-19 cs.CV cs.AI 84%

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

机构 * Department of Electronics, Information, and Bioengineering(电子、信息与生物工程系) Polytechnic University of Milan(米兰理工学院) Department of Industrial, and Information Engineering(工业与信息工程系) University of Pavia(帕维亚大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref Moglia, A., Nastasio, E.C., Mainardi, L. et al. MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Observation and Localization in CT Images. J Healthc Inform Res (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19294 2025-10-01 cs.CV cs.AI cs.CL 84%

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

Ranjan Sapkota, Manoj Karkee

专题命中 视觉定位与Grounding :vision-language model(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI

Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)

Journal ref Information Fusion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09480 2025-04-15 cs.CV cs.AI 84%

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai, Jinqing Zhang, Qiqi Zhan, Ziyue Huang, Hongxi Yan, Qiao Wan, Chenguang Liu, Junzhe Wang, Jiahui Lv, Ziqi Liu, Tengyuan Shi, Qingjie Liu, Yunhong Wang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments A Review and Evaluation about Vision-Language Model for Object Detection and Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01525 2024-07-18 cs.CV cs.AI cs.CL 84%

ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

Chenming Zhu, Tai Wang, Wenwei Zhang, Kai Chen, Xihui Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by ECCV 2024. A comprehensive and hierarchical 3D reasoning grounding benchmark in the era of foundation models. Project page: https://zcmax.github.io/projects/ScanReason

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05704 2024-04-24 cs.CV cs.AI cs.CL 84%

Visual Grounding Methods for VQA are Working for the Wrong Reasons!

Robik Shrestha, Kushal Kafle, Christopher Kanan

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Published in ACL 2020 under the title "A negative case analysis of visual grounding methods for VQA"

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15462 2024-01-09 cs.CV cs.CL cs.LG 84%

Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Ziyan Yang, Kushal Kafle, Franck Dernoncourt, Vicente Ordonez

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

Comments CVPR 2023. Fix ReferIt results. Code: https://github.com/uvavision/AMC-grounding Project Webpage: https://vislang.ai/amc

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20803 2026-08-27 cs.CV 版本更新 84%

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

ARGenSeg:基于自回归图像生成模型的图像分割

Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou

机构 * Ant Group(蚂蚁集团)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 该研究提出ARGenSeg框架,将图像生成融入MLLM,通过并行生成视觉令牌实现高效图像分割,在多数据集上性能优于现有方法且推理速度显著提升。

Comments Accepted to NeurIPS 2025, 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23553 2026-08-26 cs.CV 版本更新 84%

RT-NeuS: Towards Real-Time Neuro-Symbolic Video Understanding via Adaptive Temporal Verification

LE-NeuS: 通过自适应时间验证实现低延迟的神经符号视频理解

Shawn Liang, Sahil Shah, Chengwei Zhou, S P Sharan, Harsh Goel, Arnab Sanyal, Sandeep Chinchali, Gourav Datta

机构 * Case Western Reserve University(凯斯西储大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 LE-NeuS通过自适应时间验证优化,实现低延迟的神经符号视频理解,在保持高准确率的同时大幅减少推理延迟。

Comments Camera-ready version, accepted at NeuS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22885 2026-08-25 cs.CV 新提交 84%

DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

DRAgent:用于指代表达分割的判别推理智能体

Yujie Qi, Luyan Zhang

机构 * School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院) Khoury College of Computer Sciences, Northeastern University(东北大学Khoury计算机科学学院)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文针对指代表达分割中MLLM单次坐标预测导致的定位偏差问题,提出DRAgent判别推理框架,通过两阶段目标选择与LoRA微调提升性能,在多个基准数据集上表现具竞争力。

Comments 5 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25467 2026-08-18 cs.CV 版本更新 84%

GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids

GridVAD: 通过分层帧网格上的空间推理实现开放集视频异常检测

Mohamed Eltahir, Ahmed O. Ibrahim, Obada Siralkhatim, Tabarak Abdallah, Sondos Mohamed

机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) Independent Researcher(独立研究员) National Center for Research (NCR)(国家研究中心)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 GridVAD通过分层帧网格的空间推理生成开放集候选描述,结合自一致性巩固和空间跟踪模块实现像素级异常掩码,其在UCSD Ped2上达到77.59的Pixel-AUROC,优于其他方法。

Comments Accepted at the Large-scale Video Object Segmentation (LVOS) Workshop in conjunction with ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09147 2026-08-11 cs.CV 新提交 84%

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

RefineAny3D:作为语义对齐的深度细化用于单目3D检测

Zhihao Zhang, Gengwei Zhang, Tianlong Chen, Xiaoming Liu

机构 * Michigan State University(密歇根州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12715 2026-08-05 cs.CV 版本更新 84%

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

VLC Fusion:面向鲁棒目标检测的视觉-语言条件传感器融合

Aditya Taparia, Noel Ngu, Mario Leiva, Joshua Shay Kricheli, John Corcoran, Nathaniel D. Bastian, Gerardo Simari, Paulo Shakarian, Ransalu Senanayake

机构 * Arizona State University(亚利桑那州立大学) Department of Computer Science and Engineering, Universidad Nacional del Sur and Institute for Computer Science and Engineering(计算机科学与工程系,国家南方大学和计算机科学与工程研究所) U.S. Department of Defense(美国国防部) United States Military Academy(美国军事学院) Syracuse University(雪城大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出VLC Fusion视觉-语言条件传感器融合框架,利用VLM捕捉环境线索动态调整模态权重,在多传感器融合目标检测任务中,于自动驾驶与军事目标数据集上较传统方法实现更优性能。

Comments 27 pages, 20 figures, Accepted for presentation at ECML PKDD 2026, Shortlisted for the Best Research Track Student Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01113 2026-08-04 cs.CV 新提交 84%

CoT-Edit: Let CoT Guide Instruction Video Editing

CoT-Edit:让思维链(CoT)指导指令视频编辑

Sen Liang, Fengbin Guan, Youliang Zhang, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出CoT-Edit的plan--guide--edit框架,以CoT增强的MLLM为规划器生成空间先验,结合扩散编辑器实现高保真指令视频编辑,性能优于多个基准方法

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00232 2026-08-04 cs.CV 新提交 84%

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

手术增强现实中的实时视觉遮挡检测

Shih-Chin Yang, Yanming Xiu, Hanting Ye, Qi Chen, Elias Rotondo, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 针对手术AR虚拟内容遮挡手术器械的问题,提出结合VLM与分割推理的延迟感知流水线,构建伪AR基准,实现87.43%准确率、479 ms延迟,较云端基线降延迟62.90%。

Comments ISMAR 2026 Mecidal Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14675 2026-07-24 cs.RO cs.AI 版本更新 84%

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

一种用于机器人的智能云边缘多模态交互系统

Zihan Guo, Xiaoqi Li

机构 * Hainan University(海南大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.AI

AI总结 针对复杂环境下资源受限机器人的交互问题,提出云边缘多模态交互框架,集成增强YOLO手势检测器与LLM、VLM智能体,改进手势检测方法,经实验验证该系统在手势检测精度、任务成功率及用户满意度方面表现良好,证明了方法的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25763 2026-06-25 cs.CV 新提交 84%

ShutterMuse: Capture-Time Photography Guidance with MLLMs

ShutterMuse: 基于多模态大语言模型的拍摄时刻摄影指导

Jiayu Li, Yixiao Fang, Tianyu Hu, Wei Cheng, Ping Huang, Zheheng Fan, Gang Yu, Xingjun Ma

机构 * Fudan University(复旦大学) StepFun

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 提出CaptureGuide-Bench基准测试,评估MLLM在摄影师构图与主体姿态推荐方面的能力,并构建ShutterMuse统一模型,通过监督与强化微调实现最佳综合性能。

Comments Project Page:https://lijayutnt.github.io/ShutterMuse

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18609 2026-06-18 cs.CV 新提交 84%

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

基于反事实证据验证的医学视觉语言模型幻觉检测与纠正

Nan Zhou, Ke Zou, Meng Liu, Linchao He, Jiaqi Zhu, Yi Zhang, Hu Chen, Huazhu Fu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨潞龄医学院) Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(四川大学数据保护与智能管理教育部重点实验室) National Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(北京理工大学自主智能无人系统国家重点实验室) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出CoEV框架,通过文本与视觉证据的双向验证检测并纠正医学VLM幻觉,无需重新训练,在四个数据集上显著提升检测和纠正性能。

Comments MICCAI 2026 Accept. Submission Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17433 2026-06-17 cs.CV 新提交 84%

LADBench: A Benchmark for Logical Fault Detection in Images

LADBench: 图像中逻辑故障检测的基准

Sahasra Kondapalli, Lara Radovanovic, Aadi Palnitkar, Mingyang Mao, Xiaomin Lin

机构 * University of South Florida(南佛罗里达大学)

专题命中 视觉定位与Grounding :VLM(summary_cn);vision language model(abstract);visual question answering(abstract);grounding(abstract)

AI总结 提出LAD-Bench基准,包含1000多张合成图像的四域逻辑异常,通过分层提示协议评估模型,揭示现有VLM在隐式逻辑故障检测上的不足。

Comments Accepted to the IEEE International Conference on Development and Learning (ICDL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10528 2026-06-05 cs.CV 84%

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

BareBones: 视觉语言模型中零样本几何理解的基准测试

Aaditya Baranwal, Vishal Yadav, Abhishek Rajora

机构 * University of Central Florida(佛罗里达大学中央分校) University of Calgary(卡尔加里大学)

专题命中 视觉定位与Grounding :LLaVA(abstract,abstract_cn);vision-language model(abstract);VLM(abstract_cn);grounding(abstract)

AI总结 提出BareBones基准,通过去除RGB纹理仅保留轮廓,测试26个视觉语言模型在零样本几何形状理解上的表现,发现模型存在严重的纹理偏置悬崖。

Comments Accepted at CVPR (13th FGVC Workshop) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14799 2026-05-27 cs.CV cs.CR cs.SI 84%

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

视觉Mamba能否提升AI生成图像检测?一项深入研究

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Xianxun Zhu, Abdenour Hadid

机构 * Laboratory of IEMN, CNRS, Centrale Lille, UMR 8520, Univ. Polytechnique Hauts-de-France(伊姆纳实验室,国家科学研究中心,里尔中央理工大学,UMR 8520,法国高等技术大学) Khalifa University(卡利法大学) School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi(索邦人工智能中心,索邦大学阿布扎克分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本研究系统评估了Vision Mamba模型在AI生成图像检测中的性能,与CNN、ViT和VLM检测器进行对比,分析了准确性、效率和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23304 2026-05-25 cs.CV 84%

General Hazard Detection

通用危险检测

Stephanie Ng, CP Lim, SueJen Looi, Hendrik Zurlinden, David Nguyen, Lei Wei, Saeid Nahavandi, Hailing Zhou

机构 * Swinburne University of Technology(斯winburne大学) National Transport Research Organisation(国家交通运输研究组织) Google Cloud(谷歌云) Deakin University(德金大学)

专题命中 视觉定位与Grounding :LLaVA(abstract,abstract_cn);vision-language model(abstract);VLM(abstract_cn);visual reasoning(abstract)

AI总结 针对现有危险检测系统在抽象安全概念上的局限性,提出基于规则合规评估的CompliVision数据集和结合LLaVA视觉推理与人在回路反馈的通用危险检测框架。

Comments 20 pages, 7 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15325 2026-05-18 cs.CV 84%

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

COPRA:基于强化学习的条件参数适应用于视频异常检测

Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan He

机构 * Auburn University(奥本大学) Tencent Hunyuan(腾讯文元)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 COPRA通过生成输入特定的参数更新,动态适应冻结的VLM,提升视频异常检测的适应性和泛化能力,同时拓展到多选视频问答和密集标注等任务。

Comments Manuscript currently under review for publication

详情

展开后加载摘要…

URL PDF HTML 收藏