arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-05 至 2025-08-05 共收录 16 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 16 篇

2508.01008 2025-08-05 cs.CV 85%

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation

Cihang Peng, Qiming Hou, Zhong Ren, Kun Zhou

机构 * State Key Lab of CAD&CG(计算机辅助设计与图形学国家重点实验室)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15542 2025-08-05 cs.CV 83%

HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation

Qinqian Lei, Bo Wang, Robby T. Tan

机构 * National University of Singapore(国立新加坡大学) University of Mississippi(密苏里大学) ASUS Intelligent Cloud Services (AICS)(ASUS智能云服务(AICS))

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01723 2025-08-05 cs.RO 82%

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Danyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang, Guangyong Shang, Qiang Ma, Zheng Yang

机构 * School of Software, Tsinghua University(清华大学软件学院) School of computer science and engineering, Central South University(中南大学计算机科学与工程学院) Inspur Yunzhou Industrial Internet Co., Ltd(Inspur Yunzhou工业互联网有限公司) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

Comments ACM MM '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02479 2025-08-05 cs.CV 79%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01699 2025-08-05 cs.CV 79%

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01408 2025-08-05 cs.CY 78%

Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?

Tarian Fu, Javier Conde, Gonzalo Martínez, Pedro Reviriego, Elena Merino-Gómez, Fernando Moral

专题命中 视觉定位与Grounding :vision language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06493 2025-08-05 stat.AP 78%

Near-real-time ship grounding damage assessment using Bayesian networks

Dimitris G. Georgiadis, Manolis S. Samuelides, Daniel Straub

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Preprint submitted to Elsevier journal

Journal ref Ocean Engineering 339 (2025) 122132

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01921 2025-08-05 cs.CV 77%

InspectVLM: Unified in Theory, Unreliable in Practice

Conor Wallace, Isaac Corley, Jonathan Lwowski

机构 * Zeitview

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to 2025 ICCV VISION Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18937 2025-08-05 cs.CV cs.CL 70%

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

Mahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡斯特大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01932 2025-08-05 cs.CV cs.AI 62%

Proactive Disentangled Modeling of Trigger-Object Pairings for Backdoor Defense

Kyle Stein, Andrew A. Mahyari, Guillermo Francia, Eman El-Sheikh

机构 * Department of Intelligent Systems and Robotics, University of West Florida(智能系统与机器人系,西佛罗里达大学) Florida Institute For Human and Machine Cognition (IHMC)(佛罗里达人类与机器认知研究所) Center for Cybersecurity, University of West Florida(网络安全中心,西佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Journal ref Computers, Materials & Continua, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01338 2025-08-05 cs.CV cs.AI 62%

Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework

Ziqi Sheng, Junyan Wu, Wei Lu, Jiantao Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04154 2025-08-05 cs.CV cs.AI 62%

CA-W3D: Leveraging Context-Aware Knowledge for Weakly Supervised Monocular 3D Detection

Chupeng Liu, Runkai Zhao, Weidong Cai

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney(计算机科学学院,工程学院,悉尼大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02405 2025-08-05 cs.RO cs.CV 57%

Improving Generalization of Language-Conditioned Robot Manipulation

Chenglin Cui, Chaoran Zhu, Changjae Oh, Andrea Cavallaro

机构 * Centre for Intelligent Sensing, Queen Mary University of London(智能传感中心,伦敦女王玛丽大学) Idiap Research Institute and École Polytechnique Fédérale de Lausanne(Idiap研究机构和日内瓦联邦理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 7 pages,18 figures,2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01150 2025-08-05 cs.CV 57%

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

Dianyi Yang, Xihan Wang, Yu Gao, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang

机构 * School of Automation, Beijing Institute of Technology, Beijing, China(自动化学院,北京理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07234 2025-08-05 math.AG 50%

Spectral Fingerprints of Algebraic Cycles: A Hodge-Theoretic Approach to the Hodge Conjecture and Special L-Values

Bita Hajebi, Pooya Hajebi

专题命中 视觉定位与Grounding :grounding(abstract)

Comments arXiv admin note: This submission has been withdrawn due to violation of arXiv policies for acceptable submissions

详情

展开后加载摘要…

URL PDF HTML 收藏