arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2509.10837 2025-11-13 cs.AI 79%

Exploring the Paradigm Shift from Grounding to Skolemization for Complex Query Answering on Knowledge Graphs

Yuyin Lu, Hegang Chen, Shanrui Xie, Yanghui Rao, Haoran Xie, Fu Lee Wang, Qing Li

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院) School of Science and Technology, Hong Kong Metropolitan University(香港理工大学科技学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06908 2025-11-11 cs.CV cs.MM 79%

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Zhaoyang Li, Yuan Bian, Xueping Wang, Erbo Zhai, Yaonan Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21844 2025-11-11 cs.CV 79%

Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation

Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers

机构 * LIVIA, ÉTS Montréal, Canada International Laboratory on Learning Systems (ILLS)(LIVIA,蒙特利尔ÉTS,加拿大国际学习系统实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02044 2025-11-11 cs.CV 79%

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Chinese Medicine Guangdong Laboratory(广东中医药实验室) Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22672 2025-10-29 cs.CV cs.CL cs.RO 79%

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

Anna Deichler, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15256 2025-10-29 cs.CV 79%

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

Jinsol Song, Jiamu Wang, Anh Tien Nguyen, Keunho Byeon, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * Korea University(韩国大学) The Catholic University of Korea(韩国天主大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 79%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08216 2025-10-28 cs.AI 79%

Grounding Methods for Neural-Symbolic AI

Rodrigo Castellano Ontiveros, Francesco Giannini, Marco Gori, Giuseppe Marra, Michelangelo Diligenti

机构 * University of Siena(锡耶纳大学) Scuola Normale Superiore(正规大学) KU Leuven(卢森堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), pp. 4806-4814, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21076 2025-10-28 cs.CV 79%

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Waseda University(早稻田大学) Wuhan University(武汉大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21083 2025-10-27 cs.CV 79%

Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease

Youssef Megahed, Atallah Madi, Dina El Demellawy, Adrian D. C. Chan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted into the ICAAI 2025 - The 9th International Conference on Advances in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19678 2025-10-24 cs.CL cs.CV 79%

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

Hao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao, Bin Chen, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18582 2025-10-23 cs.CV 79%

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) Adobe(Adobe公司)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12718 2025-10-23 cs.CV cs.MM 79%

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17384 2025-10-21 cs.CV 79%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17023 2025-10-21 cs.CV cs.MM 79%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17007 2025-10-21 cs.CV 79%

An empirical study of the effect of video encoders on Temporal Video Grounding

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Felipe Bravo-Marquez

机构 * Department of Computer Science, University of Chile(计算机科学系,智利大学) CENIA and IMFD(CENIA和IMFD) Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 79%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11375 2025-10-21 cs.CV cs.MM 79%

Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion

Xinghan Wang, Zixi Kang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing (TIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16080 2025-10-21 q-bio.QM cs.AI 79%

TriAgent: Automated Biomarker Discovery with Deep Research Grounding for Triage in Acute Care by LLM-Based Multi-Agent Collaboration

Kerem Delikoyun, Qianyu Chen, Win Sen Kuan, John Tshon Yit Soong, Matthew Edward Cove, Oliver Hayden

机构 * Technical University of Munich(慕尼黑技术大学) National University of Singapore(新加坡国立大学) National University Hospital(新加坡国立大学医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14965 2025-10-17 cs.CV 79%

ChangingGrounding: 3D Visual Grounding in Changing Scenes

Miao Hu, Zhiwei Huang, Tai Wang, Jiangmiao Pang, Dahua Lin, Nanning Zheng, Runsen Xu

机构 * Xi’an Jiaotong University(西安交通大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13800 2025-10-17 cs.CV 79%

Reasoning in Space via Grounding in the World

Yiming Chen, Zekun Qi, Wenyao Zhang, Xin Jin, Li Zhang, Peidong Liu

机构 * Westlake University(西湖大学) Shanghai Innovation Institute(上海创新研究院) Zhejiang University(浙江大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(技术研究院) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02477 2025-10-16 cs.RO cs.CV 79%

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng, Lue Fan, Dong An, Changwei Wang, Li Guo, Weiliang Meng, Xiaopeng Zhang, Rongtao Xu, Shibiao Xu

机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts Telecommunications, China [1ex] dKey Laboratory of Computing Power Network Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet Service Computing, Shandong Fundamental Research Center for Computer Science, China

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026

Journal ref Information Fusion, 126 (Part B), February 2026, 103652

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10587 2025-10-14 cs.CV 79%

A Simple and Better Baseline for Visual Grounding

Jingchao Wang, Wenlong Zhang, Dingjiang Huang, Hong Wang, Yefeng Zheng

机构 * School of Data Science(数据科学学院) Engineering East China Normal University Shanghai, China(工程 东华师范大学 上海中国) OpenScience Lab Shanghai AI Laboratory Shanghai, China(OpenScience Lab 上海AI实验室 上海中国) School of Life Science(生命科学学院) Technology Xi'an Jiaotong University Xi'an, China(技术 西安交通大学 西安中国) Medical Artificial Intelligence Laboratory Westlake University Hangzhou, China(医学人工智能实验室 西湖大学 杭州中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05153 2025-10-08 cs.AI cs.IT math.IT 79%

An Algorithmic Information-Theoretic Perspective on the Symbol Grounding Problem

Zhangchi Liu

机构 * Zhangchi Liu(刘志强)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 7 pages, 1 table (in appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03840 2025-10-07 cs.CV 79%

Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models

Pranav Sharma, Shivank Garg, Durga Toshniwal

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥里分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments ACM MM'25, MALLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03376 2025-10-07 cs.CV eess.IV 79%

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

Sanjukta Ghosh

专题命中 视觉定位与Grounding :visual language model(title,abstract);分类 cs.CV

Comments Pre-review version submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00907 2025-10-06 cs.AI 79%

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

Ram Ramrakhya, Matthew Chang, Xavier Puig, Ruta Desai, Zsolt Kira, Roozbeh Mottaghi

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta FAIR

专题命中 视觉定位与Grounding :grounding(title);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11955 2025-09-30 cs.CV 79%

Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

Seunghun Lee, Jiwan Seo, Jeonghoon Kim, Sungho Moon, Siwon Kim, Haeun Yun, Hyogyeong Jeon, Wonhyeok Choi, Jaehoon Jeong, Zane Durante, Sang Hyun Park, Sunghoon Im

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project page: https://seung-hun-lee.github.io/projects/TGL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04511 2025-09-30 cs.CV 79%

FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection

Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Sun Yat-sen University(中山大学) Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Key Laboratory of Machine Intelligence and Advanced Computing(人工智能与先进计算重点实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 12 pages, 4 figures, Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23768 2025-09-29 cs.CL cs.CV 79%

Texture or Semantics? Vision-Language Models Get Lost in Font Recognition

Zhecheng Li, Guoxian Song, Yujun Cai, Zhen Xiong, Junsong Yuan, Yiwei Wang

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) The University of Queensland(昆士兰大学) University of Southern California(南加州大学) University at Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏