arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-22 至 2025-08-22 共收录 27 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2508.15189 2025-08-22 cs.AI cs.CV eess.IV 73%

SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis

Jiahao Xu, Changchang Yin, Odysseas Chatzipanagiotou, Diamantis Tsilimigras, Kevin Clear, Bingsheng Yao, Dakuo Wang, Timothy Pawlik, Ping Zhang

机构 * The Ohio State University(俄亥俄州立大学) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心) Northeastern University(东北大学)

专题命中 视觉问答 :visual question answering(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15717 2025-08-22 cs.CV cs.AI 62%

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla, Aashu Singh, Shlok Kumar Mishra, Lizhu Zhang, Mengye Ren

机构 * Meta AI New York University(纽约大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00329 2025-08-22 cs.CV cs.LG 62%

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.LG

Comments TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05182 2025-08-22 cs.IR 50%

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza, Yifan Nie, Jimmy Lin

专题命中 视觉问答 :MLLM(abstract)

Comments 15 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 4 篇

2508.07470 2025-08-22 cs.CV 74%

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

Siminfar Samakoush Galougah, Rishie Raj, Sanjoy Chowdhury, Sayan Nag, Ramani Duraiswami

专题命中 视觉推理 :visual reasoning(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14923 2025-08-22 cs.AI 57%

A Fully Spectral Neuro-Symbolic Reasoning Architecture with Graph Signal Processing as the Computational Backbone

Andrew Kiruluta

机构 * Andrew Kiruluta(独立研究者)

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15164 2025-08-22 cs.CL 50%

ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following

Seungmin Han, Haeun Kwon, Ji-jun Park, Taeyang Yoon

机构 * Dongguk University(东国大学)

专题命中 视觉推理 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14941 2025-08-22 cs.MM cs.CL 50%

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视觉推理 :grounding(abstract)

Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 10 篇

2506.04562 2025-08-22 cs.GR cs.CV 83%

Handle-based Mesh Deformation Guided By Vision Language Model

Xingpeng Sun, Shiyang Jia, Zherong Pan, Kui Wu, Aniket Bera

机构 * Purdue University(普渡大学) LightSpeed Studios University of California San Diego(加州大学圣地亚哥分校)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03290 2025-08-22 cs.CV cs.AI 81%

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao, Yufan Zhou, Yixin Cao, Qifan Wang, Weifeng Ge, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) Virginia Tech(弗吉尼亚理工大学) The Chinese University of Hong Kong(香港中文大学) NVIDIA Adobe Research(Adobe研究) Fudan University(复旦大学) Meta AI

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10225 2025-08-22 cs.CV 70%

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

Jinglun Li, Kaixun Jiang, Zhaoyu Chen, Bo Lin, Yao Tang, Weifeng Ge, Wenqiang Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai(智能机器人与先进制造学院,复旦大学,上海) Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai(上海智能信息处理重点实验室,计算机科学与人工智能学院,复旦大学,上海) JIIOV Technology, Beijing(JIIOV技术,北京)

专题命中 视觉定位与Grounding :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18246 2025-08-22 cs.CV 70%

Referring Expression Instance Retrieval and A Strong End-to-End Baseline

Xiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo, Ning Jiang, Quan Lu, Ming Tang, Jinqiao Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Mashang Consumer Finance Co, Ltd(马商消费金融有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03847 2025-08-22 cs.LG cs.AI 62%

KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis

Reilly Haskins, Benjamin Adams

机构 * Department of Computer Science and Software Engineering, University of Canterbury(计算机科学与软件工程系,坎特伯雷大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15767 2025-08-22 cs.CV 57%

ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling

Jinhyung Park, Javier Romero, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Yichen Xu, Federica Bogo, Shoou-I Yu, Kris Kitani, Rawal Khirodkar

机构 * Meta Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ICCV 2025; Website: https://jindapark.github.io/projects/atlas/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15641 2025-08-22 cs.CV 57%

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Pengcheng Fang, Yuxia Chen, Rui Guo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15236 2025-08-22 eess.IV cs.CV 57%

Pathology-Informed Latent Diffusion Model for Anomaly Detection in Lymph Node Metastasis

Jiamu Wang, Keunho Byeon, Jinsol Song, Anh Nguyen, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * School of Electrical Engineering, Korea University, Seoul 02841, Korea(韩国大学电子工程学院) Department of Pathology, Korea University Anam Hospital and Department of Biomedical Informatics, Korea University College of Medicine, Seoul 02841, Korea(韩国大学医学院病理学系) Department of Hospital Pathology, Seoul St. Mary’s Hospital, College of Medicine, The Catholic University of Korea, Seoul 06591, Korea(韩国天主大学医学院圣玛丽医院医院病理学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13560 2025-08-22 cs.CV 57%

DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup

Zhen Qu, Xian Tao, Xinyi Gong, ShiChen Qu, Xiaopei Zhang, Xingang Wang, Fei Shen, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Casivision Longmen Laboratory(龙门实验室) HDU UTS UCLA(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025, Project: https://github.com/xiaozhen228/DictAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22668 2025-08-22 cs.CV 57%

Understanding Co-speech Gestures in-the-wild

Sindhu B Hegde, K R Prajwal, Taein Kwon, Andrew Zisserman

机构 * Visual Geometry Group, Dept. of Engineering Science, University of Oxford(牛津大学视觉几何组)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Main paper - 11 pages, 4 figures, Supplementary - 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 2 篇

2508.15232 2025-08-22 cs.CV 70%

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, Si Liu

机构 * Beihang University(北京航空航天大学) Sangfor Technologies Inc.(深信服科技有限公司) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 GUI与屏幕智能体 :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15663 2025-08-22 cs.RO cs.AI 57%

Mind and Motion Aligned: A Joint Evaluation IsaacSim Benchmark for Task Planning and Low-Level Policies in Mobile Manipulation

Nikita Kachaev, Andrei Spiridonov, Andrey Gorodetsky, Kirill Muravyev, Nikita Oskolkov, Aditya Narendra, Vlad Shakhuro, Dmitry Makarov, Aleksandr I. Panov, Polina Fedotova, Alexey K. Kovalev

机构 * AIRI Moscow Institute of Physics and Technology (MIPT)(莫斯科物理技术学院) Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学) Sberbank, Robotics Center(Sberbank机器人中心) Skoltech

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 2 篇

2508.15370 2025-08-22 cs.CL cs.AI 79%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20077 2025-08-22 cs.CV cs.CL 57%

The Devil is in the EOS: Sequence Training for Detailed Image Captioning

Abdelrahman Mohamed, Yova Kementchedjhieva

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 3 篇

2410.06154 2025-08-22 cs.CV 85%

GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao, Sivan Doveh, Wei Lin, Paul Gavrikov, Michael Dorkenwald, Shiqi Yang, Saurav Jha, Hiromi Wakaki, Yuki Mitsufuji, Horst Possegger, Rogerio Feris, Leonid Karlinsky, James Glass

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Sony(索尼) Weizmann Institute of Science(魏茨曼科学研究院) JKU(约翰纳斯堡大学) Tübingen AI Center(图宾根人工智能中心) UVA(乌得勒支大学) UNSW(新南威尔士大学) TU Graz(格拉茨技术大学) MIT-IBM(麻省理工学院-IBM)

专题命中 VLM训练与架构 :vision language model(title);vision-language model(abstract);VLM(abstract);LLaVA(abstract)

Comments Code: https://github.com/jmiemirza/GLOV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15688 2025-08-22 cs.CV 83%

LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions

Yongju Jia, Jiarui Ma, Xiangxian Li, Baiqiao Zhang, Xianhui Cao, Juan Liu, Yulong Bian

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Comments accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15036 2025-08-22 cs.CR cs.AI 57%

MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs

Ruyi Ding, Tianhong Xu, Xinyi Shen, Aidong Adam Ding, Yunsi Fei

机构 * Louisiana State University(路易斯安那州立大学) Northeastern University(东北大学) Yale University(耶鲁大学)

专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.AI

Comments This paper will appear in CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他VLM 2 篇

2503.11060 2025-08-22 cs.CV 70%

BannerAgency: Advertising Banner Design with Multimodal LLM Agents

Heng Wang, Yotaro Shimose, Shingo Takamatsu

机构 * Sony Group Corporation(索尼集团)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15297 2025-08-22 cs.CV cs.AI 62%

DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding

Zhu Wang, Homaira Huda Shomee, Sathya N. Ravi, Sourav Medya

机构 * Department of Computer Science, University of Illinois Chicago(计算机科学系,伊利诺伊大学芝加哥分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by EMNLP 2025. 22 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏