arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-06 至 2025-10-06 共收录 33 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 6 篇

2510.02328 2025-10-06 cs.CL cs.AI cs.MA 83%

AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering

Ziqing Wang, Chengsheng Mao, Xiaole Wen, Yuan Luo, Kaize Ding

机构 * Northwestern University(西北大学) Microsoft(微软公司)

专题命中 视觉问答 :visual question answering(title,abstract);multimodal large language model(abstract);分类 cs.AI

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03232 2025-10-06 cs.CV 79%

LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models

Ci-Siang Lin, Min-Hung Chen, Yu-Yang Sheng, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾国立台湾大学通信工程研究所) NVIDIA

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23899 2025-10-06 cs.CV 79%

Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq, Rakesh Chandra Joshi

机构 * Amity University(阿米蒂大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 12 pages (9 main + 2 references/appendix), 2 figures, conference paper submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06461 2025-10-06 cs.CV 70%

Ranked from Within: Ranking Large Multimodal Models Without Labels

Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

机构 * Australian National University Sydney AI Centre, The University of Sydney Curtin University University of \'OBuda

专题命中 视觉问答 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV

Comments ICML 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02543 2025-10-06 cs.CV 57%

Exploring OCR-augmented Generation for Bilingual VQA

JoonHo Lee, Sunho Park

机构 * KL-Net, South Korea(韩国KL-Net)

专题命中 视觉问答 :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00579 2025-10-06 cs.MM cs.IR 50%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 视觉问答 :vision-language model(abstract)

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 1 篇

2510.02780 2025-10-06 cs.CV 83%

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

Prahitha Movva

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Journal ref COLM 2025: First Workshop on the Application of LLM Explainability to Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 14 篇

2510.02750 2025-10-06 cs.CV 85%

Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models

Lihua Zhou, Mao Ye, Shuaifeng Li, Nianxin Li, Jinlin Wu, Xiatian Zhu, Lei Deng, Hongbin Liu, Jiebo Luo, Zhen Lei

机构 * Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究院,中国科学院,香港,中国) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科学与技术大学计算机科学与工程学院) Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究院,CVSSP, Surrey大学) School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学) University of Rochester(罗切斯特大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03230 2025-10-06 cs.CV cs.AI 81%

Improving GUI Grounding with Explicit Position-to-Coordinate Mapping

Suyuchen Wang, Tianyu Zhang, Ahmed Masry, Christopher Pal, Spandana Gella, Bang Liu, Perouz Taslakian

机构 * ServiceNow Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) York University(约克大学) Polytechnique Montréal(蒙特利尔理工学院) McGill University(麦吉尔大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 81%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00907 2025-10-06 cs.AI 79%

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

Ram Ramrakhya, Matthew Chang, Xavier Puig, Ruta Desai, Zsolt Kira, Roozbeh Mottaghi

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta FAIR

专题命中 视觉定位与Grounding :grounding(title);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02592 2025-10-06 cs.AI 74%

Multimodal Large Language Model Framework for Safe and Interpretable Grid-Integrated EVs

Jean Douglas Carvalho, Hugo Kenji, Ahmad Mohammad Saber, Glaucia Melo, Max Mauro Dias Santos, Deepa Kundur

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

Comments This paper has been presented at the 2025 IEEE PES Conference on Innovative Smart Grid Technologies (ISGT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 62%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03104 2025-10-06 cs.CV cs.RO 57%

Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields

Zhiting Mei, Ola Shorinwa, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02722 2025-10-06 cs.CV 57%

MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context

Junyu Shi, Yong Sun, Zhiyuan Zhang, Lijiang Liu, Zhengjie Zhang, Yuxin He, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02566 2025-10-06 cs.CV 57%

PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction

Qiao Feng, Yiming Huang, Yufu Wang, Jiatao Gu, Lingjie Liu

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02338 2025-10-06 cs.CL cs.AI 57%

Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards

Samyak Jhaveri, Praphul Singh, Jangwon Kim, Tara Taghavi, Krishnaram Kenthapadi

机构 * Oracle Health AI(Oracle健康AI)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02790 2025-10-06 cs.CV cs.CL 57%

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15655 2025-10-06 cs.SE cs.AI cs.CL cs.IR 57%

cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree

Yilin Zhang, Xinran Zhao, Zora Zhiruo Wang, Chenyang Yang, Jiayi Wei, Tongshuang Wu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18525 2025-10-06 cs.SE cs.LG 57%

Programming with Pixels: Can Computer-Use Agents do Software Engineering?

Pranjal Aggarwal, Sean Welleck

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10289 2025-10-06 cs.CV 57%

Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection

Jiawen Zhu, Yew-Soon Ong, Chunhua Shen, Guansong Pang

机构 * Singapore Management University(新加坡管理学院) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025; 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 1 篇

2509.23263 2025-10-06 cs.AI 57%

GUI-PRA: Process Reward Agent for GUI Tasks

Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang

机构 * Zhejiang University(浙江大学) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 3 篇

2510.02677 2025-10-06 cs.AI cs.LG 73%

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

Zhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang, Minzhou Pan, Shuang Yang, Bo Li

机构 * University of Chicago(芝加哥大学) University of Illinois(伊利诺伊大学) Virtue AI Meta

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

Comments 60 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01534 2025-10-06 cs.CV 70%

Toward a Holistic Evaluation of Robustness in CLIP Models

Weijie Tu, Weijian Deng, Tom Gedeon

专题命中 幻觉与鲁棒性 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01494 2025-10-06 cs.LG cs.AI 62%

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan, Ken Ziyu Liu, Sanmi Koyejo

机构 * ETH Zürich(苏黎世联邦理工学院) Stanford CS(斯坦福大学计算机科学系)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 4 篇

2510.02922 2025-10-06 cs.CV cs.AI 84%

Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights

Daphne Tsolissou, Theofanis Ganitidis, Konstantinos Mitsis, Stergios CHristodoulidis, Maria Vakalopoulou, Konstantina Nikita

机构 * National Technical University of Athens(希腊国家技术大学) CentraleSupélec, Université Paris-Saclay(巴黎-萨克雷大学中央理工学院)

专题命中 VLM训练与架构 :vision-language model(title,abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02803 2025-10-06 cs.RO cs.AI cs.CV 84%

Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving

Yifan Liao, Zhen Sun, Xiaoyun Qiu, Zixiao Zhao, Wenbing Tang, Xinlei He, Xinhu Zheng, Tianwei Zhang, Xinyi Huang, Xingshuo Han

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Northwest A&F University(西北农林科技大学) Nanyang Technological University(南洋理工大学) Jinan University(济南大学)

专题命中 VLM训练与架构 :VLM(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI

Comments 13 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02790 2025-10-06 cs.CV cs.AI cs.CL cs.MM 73%

MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding

Jingyuan Deng, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

Comments accepted to emnlp2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02778 2025-10-06 cs.CV 57%

AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding

Xian Zhang, Zexi Wu, Zinuo Li, Hongming Xu, Luqi Gong, Farid Boussaid, Naoufel Werghi, Mohammed Bennamoun

机构 * The University of Western Australia(西澳大学) Dalian University of Technology(大连理工大学) Khalifa University(卡利夫大学) Zhejiang Lab(浙江实验室)

专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他VLM 4 篇

2510.02665 2025-10-06 cs.CL 71%

Self-Improvement in Multimodal Large Language Models: A Survey

Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特丹大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :multimodal large language model(title)

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏