arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2503.06442 2025-05-21 cs.CV cs.MM 57%

OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection

Yu Liu, Hao Tang, Haiqi Zhang, Jing Qin, Zechao Li

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Centre for Smart Health(智能健康中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to the 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12001 2025-05-20 cs.AI cs.MA 57%

Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework

Ruta Binkyte

机构 * Ruta Binkyte(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11676 2025-05-20 cs.CV 57%

DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation

Ziyu Zhao, Xiaoguang Li, Linjia Shi, Nasrin Imanpour, Song Wang

机构 * University of South Carolina(南卡罗来纳大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21186 2025-05-20 cs.LG 57%

GLIP-OOD: Zero-Shot Graph OOD Detection with Graph Foundation Model

Haoyan Xu, Zhengtao Yao, Xuzhi Zhang, Ziyi Wang, Langzhou He, Yushun Dong, Philip S. Yu, Mengyuan Li, Yue Zhao

机构 * University of Southern California(南加州大学) University of Maryland, College Park(马里兰大学学院公园分校) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Florida State University(佛罗里达州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14162 2025-05-19 cs.AI 57%

EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language Models

Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10836 2025-05-19 cs.CL cs.CV 57%

Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs

Abhishek Dey, Aabha Bothera, Samhita Sarikonda, Rishav Aryan, Sanjay Kumar Podishetty, Akshay Havalgi, Gaurav Singh, Saurabh Srivastava

机构 * George Mason University(乔治·马歇尔大学) Amazon(亚马逊公司)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

Comments Accepted at NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13274 2025-05-19 cs.HC cs.AI 57%

Augmented Object Intelligence with XR-Objects

Mustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du, Andrea Colaço, Johnny Lee, Mar Gonzalez-Franco, David Kim

机构 * Google(谷歌)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

Comments 15 pages, 15 figures, 2024 ACM Symposium on User Interface Software and Technology (UIST)

Journal ref ACM UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14347 2025-05-16 cs.CV 57%

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng, Yuda Xiong, Wenlong Liu, Zhengyu Ma, Junyi Shen, Yuan Gao, Xiaoke Jiang, Xingyu Chen, Zhuheng Song, Yuhong Zhang, Hongjie Huang, Han Gao, Shilong Liu, Hao Zhang, Feng Li, Kent Yu, Lei Zhang

机构 * IDEA Research Team International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09155 2025-05-15 cs.CV 57%

AMSnet 2.0: A Large AMS Database with AI Segmentation for Net Detection

Yichen Shi, Zhuofu Tao, Yuhao Gao, Li Huang, Hongyang Wang, Zhiping Yu, Ting-Jung Lin, Lei He

机构 * University of California, Los Angeles, USA(加州大学洛杉矶分校) Tsinghua University, Beijing, China(清华大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments accepted by LAD25

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16696 2025-05-15 cs.CY cs.AI 57%

Public Constitutional AI

Gilad Abiri

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06663 2025-05-13 cs.CV 57%

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection

Yongqi Wang, Xinxiao Wu, Shuo Yang

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments IJCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08507 2025-05-13 cs.CV 57%

Referring to Any Person

Qing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren, Yuda Xiong, Yihao Chen, Qin Liu, Lei Zhang

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02590 2025-05-13 cs.CV cs.RO 57%

Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling

Xiaowen Qiu, Jincheng Yang, Yian Wang, Zhehuan Chen, Yufei Wang, Tsun-Hsuan Wang, Zhou Xian, Chuang Gan

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11928 2025-05-12 cs.RO cs.AI 57%

"Set It Up!": Functional Object Arrangement with Compositional Generative Models

Yiqing Xu, Jiayuan Mao, Yilun Du, Tomas Lozáno-Pérez, Leslie Pack Kaelbling, David Hsu

机构 * School of Computing, Smart System Institute, National University of Singapore(计算学院、智能系统研究所、新加坡国立大学) CSAIL, Massachusetts Institute of Technology(计算机科学与人工智能实验室、麻省理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 10 pages main paper, 21 pages appendix, RSS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05343 2025-05-09 cs.CV cs.SD eess.AS 57%

Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization

Sooyoung Park, Arda Senocak, Joon Son Chung

机构 * School of Electrical Engineering, KAIST, South Korea(韩国延世大学电气工程学院) ETRI, South Korea(韩国电子技术研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Journal Extension of WACV 2024 paper (arXiv:2311.04066). Code is available at ACL-SSL" target="_blank" rel="noopener">https://github.com/swimmiing/ACL-SSL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05197 2025-05-09 cs.AI cs.CY 57%

Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt

Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham, Sébastien Krier, Manfred Diaz, Simon Osindero

机构 * Google DeepMind(谷歌DeepMind) University of Toronto(多伦多大学) Mila - Québec AI Institute(魁北克AI研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05180 2025-05-09 cs.LG 57%

OpenworldAUC: Towards Unified Evaluation and Optimization for Open-world Prompt Tuning

Cong Hua, Qianqian Xu, Zhiyong Yang, Zitai Wang, Shilong Bao, Qingming Huang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

Comments This paper has been accepted by ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04410 2025-05-08 cs.CV 57%

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

Junjie Wang, Bin Chen, Yulin Li, Bin Kang, Yichi Chen, Zhuotao Tian

机构 * School of Computer Science and Technology, HIT, Shenzhen(哈尔滨工业大学深圳校区计算机科学与技术学院) International Research Institute for Artificial Intelligence, HIT, Shenzhen(哈尔滨工业大学深圳国际人工智能研究院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02829 2025-05-06 cs.AI 57%

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta, Trevor Darrell, David M. Chan

机构 * Department of Electrical Engineering and Computer Sciences, University of California-Berkeley, Berkeley, CA, USA(电气工程与计算机科学系,加州大学伯克利分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

Comments 28 pages, 10 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02448 2025-05-06 cs.CV 57%

Recent Advances in Out-of-Distribution Detection with CLIP-Like Models: A Survey

Chaohua Li, Enhao Zhang, Chuanxing Geng, Songcan Chen

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息产业部模式分析与机器智能重点实验室) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07577 2025-05-06 cs.CV 57%

3D Vision-Language Gaussian Splatting

Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng, Anwesa Choudhuri, Terrence Chen, Chen Chen, Ziyan Wu

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted at ICLR 2025. Main paper + supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14467 2025-05-02 cs.CV 57%

LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation

Jiachen Li, Qing Xie, Renshu Gu, Jinyu Xu, Yongjian Liu, Xiaohan Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21495 2025-05-01 cs.CV cs.MM 57%

Consistency-aware Fake Videos Detection on Short Video Platforms

Junxi Wang, Jize liu, Na Zhang, Yaxiong Wang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments 2025 icic

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03118 2025-05-01 cs.HC cs.CV 57%

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller, Junwei Zheng, Kailun Yang, Anhong Guo, Kathrin Gerling, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19754 2025-04-29 cs.IR cs.AI cs.CL 57%

Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation

Carlo Merola, Jaspinder Singh

机构 * Department of Computer Science and Engineering, University of Bologna(计算机科学与工程系,博洛尼亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 13 pages, 2 figures, Second Workshop on Knowledge-Enhanced Information Retrieval, ECIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19127 2025-04-29 cs.CV cs.MM 57%

DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning

Jialang Lu, Huayu Zhao, Huiyu Zhai, Xingxing Yang, Shini Han

机构 * School of Cyber Science and Technology, Hubei University(湖北大学计算机科学与技术学院) Department of Electrical Automation Design, Beijing Shougang International Engineering Technology(北京首钢国际工程技术部) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学部) School of Computer Science and Technology, Harbin University of Science and Technology(哈尔滨理工大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICMR 2025 Main track. Code is available at https://github.com/Wenyuzhy/DeepSPG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19086 2025-04-29 cs.CV 57%

Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction

Xiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi, Shichu Sun, Jing Xing, Jian Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(先进交叉学科研究院,中国科学院大学) Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16574 2025-04-24 cs.CL cs.AI 57%

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression

Lizhe Chen, Binjia Zhou, Yuyao Ge, Jiayi Chen, Shiguang NI

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Zhejiang University(浙江大学) CAS Key Laboratory of AI Security, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03563 2025-04-23 cs.DB cs.AI 57%

A Conceptual Model for Attributions in Event-Centric Knowledge Graphs

Florian Plötzky, Katarina Britz, Wolf-Tilo Balke

机构 * ifis.cs.tu-bs.de(技术科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by Data & Knowledge Engineering, 22 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21131 2025-04-23 cs.AI cs.CL 57%

Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments

Marharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico, Kadi Tulver, Eduard Barbu, Raul Vicente

机构 * AAAI-25

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This paper extends the AAAI-2025 version by including the Appendix

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 39(15), 16308-16316, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏