arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2512.17640 2025-12-22 cs.CV 70%

Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs

通过可微认知引导的多模态大语言模型进行生成式人类-物体交互检测

Zhaolin Cai, Huiyu Duan, Zitong Xu, Fan Li, Zhi Liu, Jing Liu, Wei Shen, Xiongkuo Min, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所) Xi’an Jiao Tong University(西安交通大学) Shandong University(山东大学) Tianjin University(天津大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出GRASP-HO框架,通过可微认知引导的多模态大语言模型实现生成式人类-物体交互检测,解决封闭集分类与开放词汇生成之间的监督不匹配问题,实现判别感知与生成推理的统一。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12492 2025-12-17 cs.CV cs.CL 70%

Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings

面向开放世界设置的零样本息肉检测自适应检测-验证框架

Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu, Honghui Zhang, Junqiao Wang, Runmeng Ding, Guanyu Liu, Tianyu Shi, Zhenyu Yu, Guofeng Pan, Ziqian Bi, Yuqi Ouyang

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Columbia University(哥伦比亚大学) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院) Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院) Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院) Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院) Zhaolong Technology(智龙科技) Purdue University(普渡大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出AdaptiveDetector,通过自适应阈值调整和成本敏感强化学习,实现开放世界中零样本息肉检测,提升召回率并减少假阴性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00858 2025-12-16 cs.RO cs.AI cs.HC 70%

Enhancing Interpretability and Interactivity in Robot Manipulation: A Neurosymbolic Approach

提升机器人操作的可解释性和交互性:一种神经符号方法

Georgios Tziafas, Hamidreza Kasaei

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);分类 cs.AI

AI总结 本文提出一种神经符号方法,通过结合语言引导的视觉推理与机器人操作,提升可解释性和交互性,实现80.2%的成功率。

Comments Published in International Journal of Robotics Research (IJRR) (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11490 2025-12-15 cs.CV cs.IR 70%

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

VLM2GeoVec:迈向通用遥感多模态嵌入

Emanuel Sánchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg

机构 * Linköping University(_linköping大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 VLM2GeoVec通过单编码器对比学习实现遥感多模态嵌入,统一可扩展检索与区域推理,提升遥感场景分析的连贯性。

Comments 21 pages, 7 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10363 2025-12-12 cs.CV 70%

Point to Span: Zero-Shot Moment Retrieval for Navigating Unseen Hour-Long Videos

点到跨度:零样本时刻检索用于导航未见的小时级视频

Mingyu Jeon, Jisoo Yang, Sungjin Han, Jinkwon Hwang, Sunjae Yoon, Jonghee Kim, Junyeoung Kim

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 P2S提出了一种无训练框架,通过适应性跨度生成器和查询分解技术,解决零样本长视频时刻检索中的搜索和细化阶段效率问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07687 2025-12-09 cs.CL cs.CV 70%

HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉

Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das

机构 * Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院) TCG Crest Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位) Indian Statistical Institute(印度统计研究所)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 HalluShift++通过分析MLLM内部表示转移,解决多模态大语言模型中的层级幻觉问题,提升幻觉检测的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06376 2025-12-09 cs.CV 70%

Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework

AI生成的驾驶视频是否准备好用于自动驾驶?一种诊断评估框架

Xinhao Xiang, Abhijeet Rastogi, Jiawei Zhang

机构 * IFM Lab, University of California, Davis(加州大学戴维斯分校IFM实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出了一种诊断评估框架,系统研究AI生成驾驶视频在自动驾驶训练和评估中的可靠性,通过分析失败模式和构建基准测试,验证了过滤AIGVs可提升性能并补充现实数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04504 2025-12-08 cs.CV 70%

AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM

AnyAnomaly: 零样本可定制化视频异常检测与LVLM

Sunghyun Ahn, Youngwan Jo, Kijung Lee, Sein Kwon, Inpyo Hong, Sanghyun Park

机构 * Yonsei University(延世大学)

专题命中 视觉定位与Grounding :vision language model(abstract);visual question answering(abstract);分类 cs.CV

AI总结 AnyAnomaly通过上下文感知的视觉问答模型实现零样本可定制化视频异常检测,无需微调大型视觉语言模型,在多个基准测试中取得最佳性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04599 2025-12-05 cs.CV 70%

Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot

通过视觉-语言分割融合进行恶意图像分析:一次检测、元素识别与定位

Sheng Hang, Chaoxiang He, Hongsheng Hu, Hanqing Hu, Bin Benjamin Zhu, Shi-Feng Sun, Dawu Gu, Shuo Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Microsoft Corporation(微软公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出一种零样本方法,通过视觉-语言分割融合实现恶意图像的检测、元素识别与定位,提升细粒度审核的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22826 2025-12-04 cs.CV 70%

Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs

某些模态比其他模态更平等:在MLLMs中解码和架构多模态整合

Tianle Chen, Chaitanya Chakka, Arjun Reddy Akula, Xavier Thomas, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Google DeepMind(谷歌DeepMind)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文研究了多模态大语言模型对矛盾模态的鲁棒性,提出模态对齐调优策略以提升多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01755 2025-12-02 cs.CV cs.RO 70%

3EED: Ground Everything Everywhere in 3D

3EED: 在三维中万物皆 grounded

Rong Li, Yuhao Dong, Tianshuai Hu, Ao Liang, Youquan Liu, Dongyue Lu, Liang Pan, Lingdong Kong, Junwei Liang, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 3EED提出一个大规模多平台多模态三维 grounding 基准测试,通过提供丰富的户外场景数据和跨平台学习技术,推动语言驱动的三维具身感知研究。

Comments NeurIPS 2025 DB Track; 38 pages, 17 figures, 10 tables; Project Page at https://project-3eed.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23594 2025-12-02 cs.CV 70%

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

PRISM-Bench: 一个包含推理错误检测的基于谜题的视觉任务基准

Yusu Qian, Cheng Wan, Chao Jia, Yinfei Yang, Qingyu Zhao, Zhe Gan

专题命中 视觉定位与Grounding :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 PRISM-Bench通过检测推理错误评估多模态模型的视觉推理能力,揭示了流畅生成与忠实推理之间的差距。

Comments This paper's first error detection task's ground truth data contains hallucination introduced by gpt and needs to be withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11165 2025-12-02 cs.CV 70%

Explainable Deep Convolutional Multi-Type Anomaly Detection

可解释的深度卷积多类型异常检测

Alex George, Lyudmila Mihaylova, Sean Anderson

机构 * School of Electrical and Electronic Engineering(电子工程学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出 MultiTypeFCDD 框架,通过图像级别标签生成多通道热图,实现多类型异常检测,有效解决传统方法的计算和内存限制问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22170 2025-12-01 cs.CV 70%

Partially Shared Concept Bottleneck Models

部分共享概念瓶颈模型

Delong Zhao, Qiang Huang, Di Yan, Yiqun Sun, Jun Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 PS-CBM通过部分共享的概念策略和概念效率准确性度量,提升模型的分类准确性和可解释性。

Comments 14 pages, 7 figures, 11 tables, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21984 2025-12-01 cs.CV 70%

PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation

PPBoost: 逐步提示增强用于文本驱动的医学图像分割

Xuchen Li, Hengrui Gu, Mohan Zhang, Qin Liu, Zhen Tan, Xinyuan Zhu, Huixue Zhou, Tianlong Chen, Kaixiong Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 PPBoost通过逐步增强弱文本提示为强空间指导,提升医学图像分割的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21705 2025-12-01 cs.CL cs.CV 70%

Insight-A: Attribution-aware for Multimodal Misinformation Detection

Insight-A: 多模态虚假信息检测中的归因意识

Junjie Wu, Yumeng Fu, Chen Gong, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 Insight-A通过归因意识和分层推理提升多模态虚假信息检测效果,有效识别伪造来源并增强跨模态一致性检查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19111 2025-11-27 cs.CV 70%

DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection

DiffSeg30k: 一种用于局部AIGC检测的多轮扩散编辑基准

Hai Ci, Ziheng Peng, Pei Yang, Yingxin Xuan, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) South China University of Technology(华南理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 DiffSeg30k通过引入多轮扩散编辑基准,推动AIGC检测从二元分类到语义分割的转变,提升局部编辑定位和模型识别的可靠性。

Comments 16 pages, 10 figures; typos corrected, references added

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19920 2025-11-26 cs.CV 70%

Intelligent Image Search Algorithms Fusing Visual Large Models

融合视觉大模型的智能图像搜索算法

Kehan Wang, Tingqiong Cui, Yang Zhang, Yu Chen, Shifeng Wu, Zhenzhang Li

机构 * Chongqing University(重庆大学) CRRC Chongqing Co., Ltd.(CRRC重庆公司) Guangdong Polytechnic Normal University(广东 polytechnic 正规大学)

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 DetVLM融合视觉大模型与物体检测,实现细粒度图像检索中的状态搜索和零样本搜索,取得高准确率。

Comments 31 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15165 2025-11-25 cs.CR cs.AI 70%

Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments

MLLMs能否检测钓鱼?一个全面的安全基准套件,聚焦于动态威胁和学术环境中的多模态评估

Jingzhuo Zhou

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出AdapT-Bench,一个针对学术环境动态钓鱼攻击的多模态安全基准套件,旨在评估MLLM的防御能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12363 2025-11-18 cs.CV 70%

Explainable AI-Generated Image Detection RewardBench

Michael Yang, Shijian Deng, William T. Doan, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特大学) Stony Brook University(石溪大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09955 2025-11-14 cs.CV 70%

Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching

Uday Bhaskar, Rishabh Bhattacharya, Avinash Patel, Sarthak Khoche, Praveen Anil Kulkarni, Naresh Manwani

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) Bosch Global Software Technologies(博世全球软件技术公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08971 2025-11-13 cs.HC cs.CV cs.MM 70%

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

Sicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun, You He, Jiankang Deng, Hang Zhang, Jifei Song, Zhensong Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 16 pages, 9 figures, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07744 2025-11-12 cs.CV 70%

VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics

Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24378 2025-11-12 cs.LG 70%

AXIS: Explainable Time Series Anomaly Detection with Large Language Models

Tian Lan, Hao Duong Le, Jinbo Li, Wenjun He, Meng Wang, Chenghao Liu, Chen Zhang

机构 * Huawei(华为)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14686 2025-11-12 cs.CV 70%

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06281 2025-11-11 cs.CV 70%

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang, Daizong Liu, Yu Cheng

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19875 2025-11-11 cs.CV 70%

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny

机构 * KAUST(卡塔尔科技大学) Monash University(墨尔本大学) RICE University(里士满大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted for oral presentation at the EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16495 2025-11-06 cs.CV 70%

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

Lingfeng Wang, Hualing Lin, Senda Chen, Tao Wang, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00940 2025-11-04 cs.RO cs.AI 70%

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

Zhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu, Che Xu, Ying Li, Chengkai Hou, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) University of Washington(华盛顿大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00613 2025-11-04 cs.CV 70%

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

Yating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng, Jie Li, Yuxin Li, Zihao Wei, Zhongpei Shen, Jiajun Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏