arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-21 至 2025-10-21 共收录 27 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 27 篇

2510.16017 2025-10-21 cs.CV cs.AI cs.CL cs.RO 86%

InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects

Ibrahim Sheikh Mohamed, Abdullah Yahya Abdullah Omaisan

机构 * Independent Researchers(独立研究者)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16538 2025-10-21 cs.CV cs.RO 85%

Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

Bastian Pätzold, Jan Nogga, Sven Behnke

机构 * Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统中心) Lamarr Institute for Machine Learning and AI(拉马尔人工智能与机器学习研究所) Center for Robotics, University of Bonn(博恩大学机器人中心)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments IEEE Robotics and Automation Letters (RA-L), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16290 2025-10-21 cs.CV cs.CL 85%

Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models

Yue Zheng, Xiufang Shi, Jiming Chen, Yuanchao Shu

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 83%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 83%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16924 2025-10-21 cs.CL 82%

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

机构 * Beijing Normal University(北京师范大学) Tencent AI Lab(腾讯AI实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11302 2025-10-21 cs.CV cs.AI cs.LG 82%

When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models

Samer Al-Hamadani

机构 * Automated Manufacturing Department(自动化制造部门) Al-Khwarizmi College of Engineering(阿尔·卡瓦尔米工程学院) University of Baghdad(巴格达大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments 30 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16455 2025-10-21 cs.CL 82%

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯公司) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)

Comments ACL 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 80%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17384 2025-10-21 cs.CV 79%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17023 2025-10-21 cs.CV cs.MM 79%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17007 2025-10-21 cs.CV 79%

An empirical study of the effect of video encoders on Temporal Video Grounding

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Felipe Bravo-Marquez

机构 * Department of Computer Science, University of Chile(计算机科学系,智利大学) CENIA and IMFD(CENIA和IMFD) Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 79%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11375 2025-10-21 cs.CV cs.MM 79%

Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion

Xinghan Wang, Zixi Kang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing (TIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16080 2025-10-21 q-bio.QM cs.AI 79%

TriAgent: Automated Biomarker Discovery with Deep Research Grounding for Triage in Acute Care by LLM-Based Multi-Agent Collaboration

Kerem Delikoyun, Qianyu Chen, Win Sen Kuan, John Tshon Yit Soong, Matthew Edward Cove, Oliver Hayden

机构 * Technical University of Munich(慕尼黑技术大学) National University of Singapore(新加坡国立大学) National University Hospital(新加坡国立大学医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17332 2025-10-21 cs.CV 77%

iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA

Zhaoran Zhao, Xinli Yue, Jianhui Sun, Yuhao Xie, Tao Shao, Liangchao Yao, Fan Xia, Yuetang Deng

机构 * Tencent, WeChat(腾讯,微信)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22827 2025-10-21 cs.CV 77%

ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents

Yilei Jiang, Yaozhi Zheng, Yuxuan Wan, Jiaming Han, Qunzhong Wang, Michael R. Lyu, Xiangyu Yue

机构 * CUHK(香港中文大学) MMLab(多模态实验室) ARISE Lab(ARISE实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments ScreenCoder-v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16785 2025-10-21 cs.CV 70%

Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs

Jiazhen Liu, Long Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17218 2025-10-21 cs.CV cs.AI 62%

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

Zhuo Cao, Heming Du, Bingqing Zhang, Xin Yu, Xue Li, Sen Wang

机构 * The University of Queensland, Australia(昆士兰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16972 2025-10-21 cs.CV cs.AI 62%

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji

机构 * Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

Comments The 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15998 2025-10-21 cs.LG cs.AI 62%

AMStraMGRAM: Adaptive Multi-cutoff Strategy Modification for ANaGRAM

Nilo Schwencke, Cyriaque Rousselot, Alena Shilova, Cyril Furtlehner

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03621 2025-10-21 cs.CV 57%

DynVFX: Augmenting Real Videos with Dynamic Content

Danah Yatim, Rafail Fridman, Omer Bar-Tal, Tali Dekel

机构 * Weizmann Institute of Science(魏兹曼科学研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Project page: https://dynvfx.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16530 2025-10-21 cs.LG stat.ML 57%

Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks

Ashutosh Srivastava, Lokesh Nagalapatti, Gautam Jajoo, Aniket Vashishtha, Parameswari Krishnamurthy, Amit Sharma

机构 * Dept. of Computer Science IIIT Hyderabad(IIIT Hyderabad 计算机科学系) Dept. of Computer Science IIT Bombay(IIT Bombay 计算机科学系) Dept. of Computer Science BITS Pilani(BITS Pilani 计算机科学系) Dept. of Computer Science UIUC(UIUC 计算机科学系) Microsoft Research Bengaluru, India(微软研究院(印度班加罗尔))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16374 2025-10-21 cs.AI 57%

Before you <think>, monitor: Implementing Flavell's metacognitive framework in LLMs

Nick Oh

机构 * socius labs(socius实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Presented at the Workshop on the Application of LLM Explainability to Reasoning and Planning at COLM 2025 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15895 2025-10-21 cs.HC cs.AI cs.SD 57%

BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation

Yunzhe Wang, Xinyu Tang, Zhixun Huang, Xiaolong Yue, Yuxin Zeng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by LLM4Music @ ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16500 2025-10-21 cs.RO 50%

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

Chen Min, Jilin Mei, Heng Zhai, Shuai Wang, Tong Sun, Fanjie Kong, Haoyang Li, Fangyuan Mao, Fuyang Liu, Shuo Wang, Yiming Nie, Qi Zhu, Liang Xiao, Dawei Zhao, Yu Hu

机构 * Research Center for Intelligent Computing Systems, SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China, 100190(中国科学院计算技术研究所,智能计算系统研究中心,SKLP,北京,中国,100190) Tongji University, Shanghai, China, 200092(同济大学,上海,中国,200092) Xi’an Jiaotong University, Shaanxi, China, 710049(西安交通大学,陕西,中国,710049) Nanchang University, Jiangxi, China, 330047(南昌大学,江西,中国,330047) Defense Innovation Institute, Beijing, China, 100073(国防科技创新院,北京,中国,100073)

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Off-road robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09542 2025-10-21 cs.CL 50%

KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs

Dingjun Wu, Yukun Yan, Zhenghao Liu, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏