arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-21 至 2025-10-21 共收录 66 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 27 篇

2507.22827 2025-10-21 cs.CV 77%

ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents

Yilei Jiang, Yaozhi Zheng, Yuxuan Wan, Jiaming Han, Qunzhong Wang, Michael R. Lyu, Xiangyu Yue

机构 * CUHK(香港中文大学) MMLab(多模态实验室) ARISE Lab(ARISE实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments ScreenCoder-v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16785 2025-10-21 cs.CV 70%

Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs

Jiazhen Liu, Long Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17218 2025-10-21 cs.CV cs.AI 62%

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

Zhuo Cao, Heming Du, Bingqing Zhang, Xin Yu, Xue Li, Sen Wang

机构 * The University of Queensland, Australia(昆士兰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16972 2025-10-21 cs.CV cs.AI 62%

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji

机构 * Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

Comments The 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15998 2025-10-21 cs.LG cs.AI 62%

AMStraMGRAM: Adaptive Multi-cutoff Strategy Modification for ANaGRAM

Nilo Schwencke, Cyriaque Rousselot, Alena Shilova, Cyril Furtlehner

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03621 2025-10-21 cs.CV 57%

DynVFX: Augmenting Real Videos with Dynamic Content

Danah Yatim, Rafail Fridman, Omer Bar-Tal, Tali Dekel

机构 * Weizmann Institute of Science(魏兹曼科学研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Project page: https://dynvfx.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16530 2025-10-21 cs.LG stat.ML 57%

Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks

Ashutosh Srivastava, Lokesh Nagalapatti, Gautam Jajoo, Aniket Vashishtha, Parameswari Krishnamurthy, Amit Sharma

机构 * Dept. of Computer Science IIIT Hyderabad(IIIT Hyderabad 计算机科学系) Dept. of Computer Science IIT Bombay(IIT Bombay 计算机科学系) Dept. of Computer Science BITS Pilani(BITS Pilani 计算机科学系) Dept. of Computer Science UIUC(UIUC 计算机科学系) Microsoft Research Bengaluru, India(微软研究院(印度班加罗尔))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16374 2025-10-21 cs.AI 57%

Before you <think>, monitor: Implementing Flavell's metacognitive framework in LLMs

Nick Oh

机构 * socius labs(socius实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Presented at the Workshop on the Application of LLM Explainability to Reasoning and Planning at COLM 2025 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15895 2025-10-21 cs.HC cs.AI cs.SD 57%

BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation

Yunzhe Wang, Xinyu Tang, Zhixun Huang, Xiaolong Yue, Yuxin Zeng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by LLM4Music @ ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16500 2025-10-21 cs.RO 50%

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

Chen Min, Jilin Mei, Heng Zhai, Shuai Wang, Tong Sun, Fanjie Kong, Haoyang Li, Fangyuan Mao, Fuyang Liu, Shuo Wang, Yiming Nie, Qi Zhu, Liang Xiao, Dawei Zhao, Yu Hu

机构 * Research Center for Intelligent Computing Systems, SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China, 100190(中国科学院计算技术研究所,智能计算系统研究中心,SKLP,北京,中国,100190) Tongji University, Shanghai, China, 200092(同济大学,上海,中国,200092) Xi’an Jiaotong University, Shaanxi, China, 710049(西安交通大学,陕西,中国,710049) Nanchang University, Jiangxi, China, 330047(南昌大学,江西,中国,330047) Defense Innovation Institute, Beijing, China, 100073(国防科技创新院,北京,中国,100073)

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Off-road robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09542 2025-10-21 cs.CL 50%

KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs

Dingjun Wu, Yukun Yan, Zhenghao Liu, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 文档图表理解 2 篇

2510.15872 2025-10-21 cs.AR cs.AI cs.LG 73%

Multimodal Chip Physical Design Engineer Assistant

Yun-Da Tsai, Chang-Yu Chao, Liang-Yeh Shen, Tsung-Han Lin, Haoyu Yang, Mark Ho, Yi-Chen Lu, Wen-Hao Liu, Shou-De Lin, Haoxing Ren

机构 * National Taiwan University(国立台湾大学) NVIDIA Research(NVIDIA研究)

专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15349 2025-10-21 cs.CL 50%

Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing

Baode Wang, Biao Wu, Weizhen Li, Meng Fang, Zuming Huang, Jun Huang, Haozhe Wang, Yanjie Liang, Ling Chen, Wei Chu, Yuan Qi

专题命中 文档图表理解 :vision-language model(abstract)

Comments This submission (arXiv:2510.15349) was mistakenly uploaded as a new article. It was intended to replace our previous work arXiv:2506.03197. All subsequent updates will be made to arXiv:2506.03197

详情

展开后加载摘要…

URL PDF HTML 收藏

3. GUI与屏幕智能体 2 篇

2510.17038 2025-10-21 cs.RO cs.AI cs.CV 62%

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi

机构 * Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学) The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院)) Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19131 2025-10-21 cs.RO cs.AI cs.CV 62%

ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments

Shreya Gummadi, Mateus V. Gasparino, Gianluca Capezzuto, Marcelo Becker, Girish Chowdhary

机构 * Field Robotics Engineering and Science Hub (FRESH), Illinois Autonomous Farm, University of Illinois at Urbana-Champaign (UIUC), IL(伊利诺伊大学厄巴纳-香槟分校) Mobile Robotics Group, São Carlos School of Engineering, University of São Paulo (EESC-USP), São Carlos, SP, Brazil(圣保罗大学)

专题命中 GUI与屏幕智能体 :visual reasoning(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 5 篇

2504.13169 2025-10-21 cs.CV 84%

Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling

Tsung-Han Wu, Heekyung Lee, Jiaxin Ge, Joseph E. Gonzalez, Trevor Darrell, David M. Chan

机构 * UC Berkeley(加州大学伯克利分校) POSTECH

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV

Comments Accepted to NeurIPS 2025; Project Page: https://reverse-vlm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14697 2025-10-21 cs.CR cs.RO 67%

AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions

Zonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang, Yuqing Ma, Jinyang Guo, Zhenfei Yin, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * SKLCCSE, Beihang University(北京航空航天大学智能科学与技术研究中心) Zhongguancun Laboratory(中关村实验室) The University of Sydney(悉尼大学) Henan University of Science(河南科技大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02456 2025-10-21 cs.LG cs.AI cs.NA math.NA 62%

Market-Driven Subset Selection for Budgeted Training

Ashish Jha, Valentin Leplat, AH Phan

机构 * Skolkovo Institute of Science and Technology(斯克洛夫诺科学与技术研究所) Innopolis University(因诺波利斯大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI、cs.LG

Comments Retitled major revision of the same work (formerly "Market-Based Data Subset Selection -- Principled Aggregation of Multi-Criteria Example Utility"). Abstract and exposition revised; ablations added; theory clarified. Core results unchanged. Supersedes v1; please process as a replacement

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15948 2025-10-21 cs.AI cs.CR 57%

VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search

MingSheng Li, Guangze Zhao, Sichen Liu

机构 * Independent Researcher(独立研究者) Harbin Institute of Technology(哈尔滨工业大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 50%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. VLM训练与架构 13 篇

2503.06073 2025-10-21 cs.CL cs.AI cs.CV 86%

GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images

Xiang Lan, Feng Wu, Kai He, Qinghao Zhao, Shenda Hong, Mengling Feng

机构 * National University of Singapore(新加坡国立大学) Peking University People’s Hospital(北京大学人民医院) Peking University(北京大学)

专题命中 VLM训练与架构 :MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Camera-Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17777 2025-10-21 cs.CV 83%

SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference

Samir Khaki, Junxian Guo, Jiaming Tang, Shang Yang, Yukang Chen, Konstantinos N. Plataniotis, Yao Lu, Song Han, Zhijian Liu

机构 * NVIDIA MIT(麻省理工学院) UC San Diego(南加州大学圣地亚哥分校) University of Toronto(多伦多大学)

专题命中 VLM训练与架构 :VLM(title,abstract);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16870 2025-10-21 cs.CV 83%

Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding

Yudan Ren, Xinlong Wang, Kexin Wang, Tian Xia, Zihan Ma, Zhaowei Li, Xiangrong Bi, Xiao Li, Xiaowei He

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10090 2025-10-21 cs.RO cs.AI 83%

Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models

Chenrui Tie, Shengxiang Sun, Jinxuan Zhu, Yiwei Liu, Jingxiang Guo, Yue Hu, Haonan Chen, Junting Chen, Ruihai Wu, Lin Shao

机构 * National University of Singapore(新加坡国立大学) University of Toronto(多伦多大学) Peking University(北京大学) Sichuan University(四川大学) Zhejiang University(浙江大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.AI

Journal ref Robotics: Science and Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17197 2025-10-21 cs.CV cs.AI 81%

ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models

Pu Zhang, Yuwei Li, Xingyuan Xian, Guoming Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15430 2025-10-21 cs.CV cs.AI 81%

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17205 2025-10-21 cs.CV cs.CL 70%

$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs

Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT, Ningbo(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT,宁波) Shanghai Jiao Tong University(上海交通大学) Hong Kong Polytechnic University(香港理工大学) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13061 2025-10-21 cs.CV 70%

Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection

Jingyao Wang, Yiming Chen, Lingyu Si, Changwen Zheng

机构 * Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of the Chinese Academy of Sciences(中国科学院大学) Beijing University of Technology(北京理工大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17041 2025-10-21 cs.CV cs.AI cs.LG 67%

Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM

Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye

机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22904 2025-10-21 cs.HC cs.AI 57%

SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches

Ehsan Latif, Zirak Khan, Xiaoming Zhai

机构 * AI4STEM Education Center University of Georgia(AI4STEM教育中心乔治亚大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.AI

Comments Submitted to NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏