arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-26 至 2025-08-26 共收录 18 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 18 篇

2505.15123 2025-08-26 cs.CV cs.AI 84%

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

Ta Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi, Qi Chen, Phi Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan W. Verjans, Vu Minh Hieu Phan

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Macquarie University(麦考瑞大学) Hanoi University of Science and Technology(河内科学技术大学) University of Wollongong(沃林根大学) Flinders University(弗林德斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17976 2025-08-26 cs.CV eess.IV 83%

Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization

Keyang Zhang, Chenqi Kong, Hui Liu, Bo Ding, Xinghao Jiang, Haoliang Li

机构 * Department of Electrical Engineering, City University of Hong Kong(香港城市大学电子工程系) Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technology University(南洋理工大学电子与电气工程学院快速丰富对象搜索(ROSE)实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :MLLM(title);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16974 2025-08-26 cs.CV 83%

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren, Zheyu Song

机构 * Zhongkai University of Agriculture and Engineering(仲恺农业工程大学) EDP University of Puerto Rico: San Sebastian(波多黎各圣塞巴斯蒂安EDP大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02943 2025-08-26 cs.CR cs.AI cs.CL cs.LG 81%

PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding

Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, Xuebing Zhou

机构 * Huawei Munich Research Center(华为慕尼黑研究中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments Accepted at PrivateNLP Workshop at ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 79%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 73%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13357 2025-08-26 cs.CL 71%

Adaptive Linguistic Prompting (ALP) Enhances Phishing Webpage Detection in Multimodal Large Language Models

Atharva Bhargude, Ishan Gonehal, Dave Yoon, Kaustubh Vinnakota, Chandler Haney, Aaron Sandoval, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title)

Comments Published at ACL 2025 SRW, 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02259 2025-08-26 cs.CV 70%

T*: Re-thinking Temporal Search for Long-Form Video Understanding

Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristobal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, Manling Li

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted by CVPR 2025; A real-world long video needle-in-haystack benchmark; long-video QA with human ref frames

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18742 2025-08-26 cs.CV cs.RO 70%

3D Feature Distillation with Object-Centric Priors

Georgios Tziafas, Yucheng Xu, Zhibin Li, Hamidreza Kasaei

机构 * Department of Artificial Intelligence University of Groningen, the Neteherlands(格罗宁根大学人工智能系) School of Informatics University of Edinburgh, United Kingdom(爱丁堡大学信息学院) Department of Computer Science University College London, United Kingdom(伦敦大学学院计算机科学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17667 2025-08-26 cs.CV cs.AI 62%

Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection

Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Nottingham Malaysia(诺丁汉大学(马来西亚)) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(机器智能与高级计算重点实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 2 figures, Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16643 2025-08-26 cs.LG cs.AI 62%

From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective

Tianhua Chen

机构 * School of Computing and Engineering University of Huddersfield(计算与工程学院赫德斯菲尔德大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments This is a substantially improved and expanded version of an earlier manuscript hosted on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5244929

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18067 2025-08-26 cs.CV 57%

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

Kaiyu Li, Xiangyong Cao, Ruixun Liu, Shihong Wang, Zixuan Jiang, Zhi Wang, Deyu Meng

机构 * School of Software Engineering, Xi’an Jiaotong University(软件工程学院,西安交通大学) School of Computer Science and Technology and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学) School of Automation, Xi’an Jiaotong University(自动化学院,西安交通大学) College of Artificial Intelligence, Xi’an Jiaotong University(人工智能学院,西安交通大学) School of Mathematics and Statistics and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学院和教育部智能网络与网络安全重点实验室,西安交通大学)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments All codes and models will be released at https://github.com/earth-insights/SegEarth-OV-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18050 2025-08-26 cs.CV 57%

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou, Hongfei Wang, Ye Li

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17255 2025-08-26 cs.CV cs.RO 57%

SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality

Yuzhi Lai, Shenghai Yuan, Peizheng Li, Jun Lou, Andreas Zell

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17044 2025-08-26 cs.CV cs.RO 57%

M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments

Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 29 pages, 3 figures, 13 tables. Preprint of the accepted article in Optical Memory and Neural Network Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08525 2025-08-26 cs.AI cs.CL 57%

Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks

Ye Ye

机构 * Ye Ye(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 14 pages, 5 figures. Preprint prepared for future submission. Includes implementation and token-efficiency analysis. Code at https://github.com/biubiutomato/TME-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17724 2025-08-26 cond-mat.soft 50%

Adhesion Control through Electric Field-Induced Water Adsorption at Oxidized Silicon Interfaces

Tunç Çiftçi, Jonathon Cottom, Rachid Hahury, Emilia Olsson, Bart Weber

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05718 2025-08-26 cs.CL 50%

A Factuality and Diversity Reconciled Decoding Method for Knowledge-Grounded Dialogue Generation

Chenxu Yang, Zheng Lin, Chong Tian, Liang Pang, Lanrui Wang, Zhengyang Tong, Qirong Ho, Yanan Cao, Weiping Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏