arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-21 至 2025-10-21 共收录 66 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2510.16292 2025-10-21 cs.LG 85%

QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models

Yutong Wang, Haiyu Wang, Sai Qian Zhang

机构 * Tandon School of Engineering, New York University(纽约大学工程学院) Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)

专题命中 视觉问答 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.LG

Comments Accepted as Spotlight paper by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17771 2025-10-21 cs.AI cs.CV 85%

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Benoit Dumoulin, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Penn State University(宾夕法尼亚州立大学)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);LLaVA(abstract);InternVL(abstract)

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14605 2025-10-21 cs.CV cs.AI 73%

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 视觉问答 :visual language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13946 2025-10-21 cs.AI 70%

Visual Instruction Bottleneck Tuning

Changdae Oh, Jiatong Li, Shawn Im, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin–Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 10 篇

2510.16907 2025-10-21 cs.AI cs.CL 83%

VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

Kangrui Wang, Pingyue Zhang, Zihan Wang, Yaning Gao, Linjie Li, Qineng Wang, Hanyang Chen, Chi Wan, Yiping Lu, Zhengyuan Yang, Lijuan Wang, Ranjay Krishna, Jiajun Wu, Li Fei-Fei, Yejin Choi, Manling Li

机构 * Northwestern University(西北大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学) Microsoft(微软) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17274 2025-10-21 cs.CV 79%

Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models

Katie Luo, Jingwei Ji, Tong He, Runsheng Xu, Yichen Xie, Dragomir Anguelov, Mingxing Tan

机构 * Computer and Information Sciences Department, Cornell University(康奈尔大学计算机与信息科学系) Waymo LLC(Waymo公司) UC Berkeley(伯克利大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments In proceedings of IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14671 2025-10-21 cs.CV 70%

UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens

Ruichuan An, Sihan Yang, Renrui Zhang, Zijun Shen, Ming Lu, Gaole Dai, Hao Liang, Ziyu Guo, Shilin Yan, Yulin Luo, Bocheng Zou, Chaoqun Yang, Wentao Zhang

机构 * Peking University(北京大学) Xi’an JiaoTong University(西安交通大学) CUHK(香港中文大学) Intel Labs, China(中国英特尔实验室) Nanjing University(南京大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Tsinghua University(清华大学)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17354 2025-10-21 cs.CL cs.AI cs.IR cs.LG 62%

Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation

Chenghao Zhang, Guanting Dong, Xinyu Yang, Zhicheng Dou

机构 * Renmin University of China(中国人民大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI、cs.LG

Comments This work is in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17157 2025-10-21 cs.CV cs.AI 62%

GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image

Yinghui Wang, Xinyu Zhang, Peng Du

专题命中 视觉推理 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16772 2025-10-21 cs.CV cs.AI 62%

Region in Context: Text-condition Image editing with Human-like semantic reasoning

Thuy Phuong Vu, Dinh-Cuong Hoang, Minhhuy Le, Phan Xuan Tan

机构 * Greenwich Vietnam FPT University(越南格林威治FPT大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16643 2025-10-21 cs.CV cs.AI cs.RO 62%

Structured Interfaces for Automated Reasoning with 3D Scene Graphs

Aaron Ray, Jacob Arkin, Harel Biggie, Chuchu Fan, Luca Carlone, Nicholas Roy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16540 2025-10-21 cs.CV cs.AI 62%

Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions

Jihoon Kwon, Kyle Min, Jy-yong Sohn

机构 * Seoul National University(首尔国立大学) Oracle(Oracle公司) Yonsei University(延世大学) Intel Labs(英特尔实验室)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025 (poster). This is the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00568 2025-10-21 cs.CV 57%

CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning

Ke Niu, Zhuofan Chen, Haiyang Yu, Yuwen Chen, Teng Fu, Mengyang Zhao, Bin Li, Xiangyang Xue

机构 * College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16641 2025-10-21 cs.CV 57%

MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models

Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung-Ho Han, Bowon Ko, Ho-Jin Choi

机构 * KAIST(韩国科学技术院) WHU(华中科技大学) NAVER CMU(卡内基梅隆大学)

专题命中 视觉推理 :VLM(abstract);分类 cs.CV

Comments Project website: https://passing2961.github.io/multiverse-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 27 篇

2510.16017 2025-10-21 cs.CV cs.AI cs.CL cs.RO 86%

InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects

Ibrahim Sheikh Mohamed, Abdullah Yahya Abdullah Omaisan

机构 * Independent Researchers(独立研究者)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16538 2025-10-21 cs.CV cs.RO 85%

Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

Bastian Pätzold, Jan Nogga, Sven Behnke

机构 * Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统中心) Lamarr Institute for Machine Learning and AI(拉马尔人工智能与机器学习研究所) Center for Robotics, University of Bonn(博恩大学机器人中心)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments IEEE Robotics and Automation Letters (RA-L), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16290 2025-10-21 cs.CV cs.CL 85%

Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models

Yue Zheng, Xiufang Shi, Jiming Chen, Yuanchao Shu

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 83%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 83%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16924 2025-10-21 cs.CL 82%

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

机构 * Beijing Normal University(北京师范大学) Tencent AI Lab(腾讯AI实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11302 2025-10-21 cs.CV cs.AI cs.LG 82%

When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models

Samer Al-Hamadani

机构 * Automated Manufacturing Department(自动化制造部门) Al-Khwarizmi College of Engineering(阿尔·卡瓦尔米工程学院) University of Baghdad(巴格达大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments 30 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16455 2025-10-21 cs.CL 82%

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯公司) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)

Comments ACL 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 80%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17384 2025-10-21 cs.CV 79%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17023 2025-10-21 cs.CV cs.MM 79%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17007 2025-10-21 cs.CV 79%

An empirical study of the effect of video encoders on Temporal Video Grounding

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Felipe Bravo-Marquez

机构 * Department of Computer Science, University of Chile(计算机科学系,智利大学) CENIA and IMFD(CENIA和IMFD) Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 79%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11375 2025-10-21 cs.CV cs.MM 79%

Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion

Xinghan Wang, Zixi Kang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing (TIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16080 2025-10-21 q-bio.QM cs.AI 79%

TriAgent: Automated Biomarker Discovery with Deep Research Grounding for Triage in Acute Care by LLM-Based Multi-Agent Collaboration

Kerem Delikoyun, Qianyu Chen, Win Sen Kuan, John Tshon Yit Soong, Matthew Edward Cove, Oliver Hayden

机构 * Technical University of Munich(慕尼黑技术大学) National University of Singapore(新加坡国立大学) National University Hospital(新加坡国立大学医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17332 2025-10-21 cs.CV 77%

iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA

Zhaoran Zhao, Xinli Yue, Jianhui Sun, Yuhao Xie, Tao Shao, Liangchao Yao, Fan Xia, Yuetang Deng

机构 * Tencent, WeChat(腾讯,微信)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏