arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2509.12145 2025-09-16 cs.CV 74%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06564 2025-08-14 cs.CV 74%

Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

Guanyu Hu, Dimitrios Kollias, Xinyu Yang

机构 * Xi'an Jiaotong University(西安交通大学) Queen Mary University of London(伦敦女王玛丽大学) Center for Multimodal AI(多模态人工智能中心) Digital Environment Research Institute(数字环境研究院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments accepted for publication at ACM Multimedia (ACM MM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08761 2025-08-13 cs.CL cs.AI 74%

DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation

Stavros Doropoulos, Stavros Vologiannidis, Ioannis Magnisalis

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 74%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 74%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视觉定位与Grounding :MLLM(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18276 2025-07-25 cs.RO cs.CV 74%

Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

Xiaojie Zhang, Yuanfei Wang, Ruihai Wu, Kunqi Xu, Yu Li, Liuyu Xiang, Hao Dong, Zhaofeng He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) School of Computer Science, Peking University(北京大学计算机学院) School of EECS, Peking University(北京大学电子工程学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02600 2025-07-22 cs.CV cs.RO eess.IV 74%

Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts

Yizhou Huang, Fan Yang, Guoliang Zhu, Gen Li, Hao Shi, Yukun Zuo, Wenrui Chen, Zhiyong Li, Kailun Yang

机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(人工智能与机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted to IROS 2025. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05336 2025-07-08 cs.CV 74%

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Linköping University(林奈大学) Australian National University(澳大利亚国立大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16769 2025-07-03 cs.CV 74%

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models

Muhammad Atta ur Rahman, Dooseop Choi, Seung-Ik Lee, KyoungWook Min

机构 * Artificial Intelligence Creative Research Lab, ETRI University of Science(人工智能创意研究实验室,ETRI大学) Field Robotics Research Section, ETRI University of Science(机器人领域研究部,ETRI大学) Artificial Intelligence Creative Research Lab, ETRI Daejeon, South Korea(人工智能创意研究实验室,ETRI大田,韩国)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments Accepted at the 17th IEEE International Conference on Advanced Computational Intelligence (ICACI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07007 2025-07-02 cs.CV 74%

Grounding Creativity in Physics: A Brief Survey of Physical Priors in AIGC

Siwei Meng, Yawei Luo, Ping Liu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted by IJCAI 2025 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22442 2025-07-01 cs.LG 74%

Features-based embedding or Feature-grounding

Piotr Makarevich

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05266 2025-06-25 cs.CV q-bio.NC 74%

Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision Transformers

Andrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan, Margaret M. Henderson, Leila Wehbe, Michael J. Tarr

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted at ICLR 2025, code: https://github.com/aluo-x/BrainSAIL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15802 2025-06-24 cs.CV 74%

Visual Prompt Engineering for Vision Language Models in Radiology

Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, Klaus Maier-Hein

机构 * Division of Medical Image Computing, German Cancer Research Center, Heidelberg, Germany(德国癌症研究中心医学图像计算部) Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany(海德堡大学数学与计算机科学学院) Medical Faculty Heidelberg, University of Heidelberg, Heidelberg, Germany(海德堡大学医学学院)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop on Emergent Visual Abilities and Limits of Foundation Models & Medical Imaging with Deep Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17375 2025-06-24 q-bio.NC cs.AI 74%

Challenges in Grounding Language in the Real World

Peter Lindes, Kaoutar Skiker

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16201 2025-06-23 cs.RO cs.CV 74%

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

Sen Wang, Le Wang, Sanping Zhou, Jingyi Tian, Jiayi Li, Haowen Sun, Wei Tang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人类-机器混合增强智能重点实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13282 2025-06-17 cs.CV 74%

Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling

Daichi Tanaka, Takumi Karasawa, Shu Takenouchi, Rei Kawakami

机构 * Institute Science of Tokyo(东京科学研究所)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01953 2025-06-17 cs.LG stat.ME 74%

The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications

Philippe Brouillard, Chandler Squires, Jonas Wahl, Konrad P. Kording, Karen Sachs, Alexandre Drouin, Dhanya Sridhar

机构 * Mila-Québec, Université de Montréal(蒙特利尔大学魁北克分校) Carnegie Mellon University(卡内基梅隆大学) Deutsches Forschungszentrum für künstliche Intelligenz (DFKI)(德国人工智能研究中心) University of Pennsylvania(宾夕法尼亚大学) Next Generation Analytics and Modulo Bio(下一代分析与Modulo Bio) ServiceNow Research(ServiceNow研究) Mila-Québec, Université Laval(魁北克蒙特利尔大学拉瓦尔分校)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments 39 pages, 8 figures; CLeaR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06622 2025-06-13 cs.CE cs.AI 74%

QuantMCP: Grounding Large Language Models in Verifiable Financial Reality

Yifan Zeng

机构 * Sun Yat-sen University, Guangzhou, China(中山大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18463 2025-05-30 cs.CV 74%

A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models

Shiho Noda, Atsuyuki Miyai, Qing Yu, Go Irie, Kiyoharu Aizawa

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments Accepted at ICIP2025 Dataset and Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19761 2025-05-27 cs.AI 74%

Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning

Zican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue, Chunlin Chen, Zhi Wang, Yu Cheng

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted by ICML 2025, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07304 2025-04-11 cs.CL cs.AI 74%

PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games

Santiago Góngora, Luis Chiruzzo, Gonzalo Méndez, Pablo Gervás

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Presented at the 15th International Conference on Computational Creativity (ICCC'24)

Journal ref Proceedings of the Fifteenth International Conference on Computational Creativity (2024) 101-106

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13727 2025-04-02 cs.CL cs.AI 74%

LLM-Human Pipeline for Cultural Context Grounding of Conversations

Rajkumar Pujari, Dan Goldwasser

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Oral at NAACL 2025 Main conference. Albuquerque, USA. Apr 29 - May 4, 2025. 19 pages, 9 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05804 2025-04-01 cs.CV 74%

CASA: Class-Agnostic Shared Attributes in Vision-Language Models for Efficient Incremental Object Detection

Mingyi Guo, Yuyang Liu, Zhiyuan Yan, Zongying Lin, Peixi Peng, Yonghong Tian

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20188 2025-03-27 cs.CV 74%

Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector

Xiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu, Xiaoming Liu

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments 8 figures; 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19240 2025-03-26 cs.CV cs.HC 74%

Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding

Hao Guo, Jianfei Zhu, Wei Fan, Chunzhi Yi, Feng Jiang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19801 2025-03-18 cs.SE cs.AI cs.CL 74%

CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells

Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn Rose

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01601 2025-03-04 cs.CV 74%

Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR

Muhammad Musab Ansari

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02089 2025-02-19 cs.CL cs.AI 74%

RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Jonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella, Quentin Carbonneaux, Taco Cohen, Gabriel Synnaeve

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Add repair model ablation, update related work

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00486 2025-02-04 cs.CV 74%

GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval

Yuxuan Wang, Difei Gao, Licheng Yu, Stan Weixian Lei, Matt Feiszli, Mike Zheng Shou

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Updated in Jan. 2025, In Proceedings of the European Conference on Computer Vision 2022 [ECCV 2022], 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14035 2024-12-12 cs.CL cs.AI 74%

Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models

Sherzod Hakimov, Yerkezhan Abdullayeva, Kushal Koshti, Antonia Schmidt, Yan Weiser, Anne Beyer, David Schlangen

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted at COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏