arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-08 至 2025-08-08 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2508.05021 2025-08-08 cs.RO 82%

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

Weifan Zhang, Tingguang Li, Yuzhen Liu

机构 * Tencent Robotics X(腾讯机器人X) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05323 2025-08-08 cs.CV 70%

Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

Frank Ruis, Gertjan Burghouts, Hugo Kuijf

机构 * TNO(荷兰技术院) Intelligent Imaging(智能成像)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05409 2025-08-08 cs.CV cs.SD eess.AS 57%

From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization

Farah Wahida, M. A. P. Chamikara, Yashothara Shanmugarasa, Mohan Baruwal Chhetri, Thilina Ranbaduge, Ibrahim Khalil

机构 * RMIT University, Australia(皇家墨尔本理工大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments 19 Pages, 24 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04868 2025-08-08 cs.CV 57%

Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications

Noreen Anwar, Guillaume-Alexandre Bilodeau, Wassim Bouachir

机构 * LITIV, Polytechnique Montréal(Polytechnique Montréal 的 LITIV) Data Science Laboratory, Université du Québec (TELUQ)(Université du Québec (TELUQ) 的 Data Science Laboratory)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05534 2025-08-08 cs.CL 50%

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim, Pranav Shetty, Dongsheng Wang, Zhiqiang Ma, Armineh Nourbakhsh, Xiaomo Liu

机构 * School of Computation, Information, and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学) Graduate Institute of International and Development Studies, Geneva(国际与发展研究研究生院,日内瓦) JPMorgan AI Research(摩根大通人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to ACL 2025-Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05061 2025-08-08 cs.DB cs.IR 50%

Data-Aware Socratic Query Refinement in Database Systems

Ruiyuan Zhang, Chrysanthi Kosyfaki, Xiaofang Zhou

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02068 2025-08-08 cs.RO 50%

"Set It Up": Functional Object Arrangement with Compositional Generative Models (Journal Version)

Yiqing Xu, Jiayuan Mao, Linfeng Li, Yilun Du, Tomas Lozáno-Pérez, Leslie Pack Kaelbling, David Hsu

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments This is the journal version accepted to the International Journal of Robotics Research (IJRR). It extends our prior work presented at Robotics: Science and Systems (RSS) 2024, with a new compositional program induction pipeline from natural language, and expanded evaluations on personalized bookshelf and bedroom furniture layout tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05236 2025-08-08 cs.MA 50%

Towards Language-Augmented Multi-Agent Deep Reinforcement Learning

Maxime Toquebiau, Jae-Yun Jun, Faïz Benamar, Nicolas Bredeche

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accespted at the European Conference on Artificial Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏