arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-09 至 2025-09-09 共收录 37 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 5 篇

2509.06759 2025-09-09 cs.LG cs.AI 84%

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization

Thanh Thi Nguyen, Campbell Wilson, Janis Dalins

机构 * AiLECS Lab, Monash University Melbourne, Australia(墨尔本大学AiLECS实验室,澳大利亚) AiLECS Lab, Monash University, Australia ICMEC Australia, Sydney, Australia(墨尔本大学AiLECS实验室,澳大利亚 ICMEC澳大利亚,悉尼,澳大利亚)

专题命中 VLM训练与架构 :vision-language model(title,abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05703 2025-09-09 cs.CV cs.AI cs.IR 84%

Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis

Ragib Amin Nihal, Benjamin Yen, Takeshi Ashizawa, Kazuhiro Nakadai

机构 * Systems and Control Engineering, Institute of Science Tokyo(科学理工学院东京)

专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03494 2025-09-09 cs.CV 70%

Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

Yahya Benmahane, Mohammed El Hassouni

机构 * Computer Science Department Faculty of Sciences, Rabat(科学学院计算机科学系,拉巴特) Computer Science Department FLSH(计算机科学系FLSH)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10032 2025-09-09 cs.CV 70%

Osprey: Pixel Understanding with Visual Instruction Tuning

Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, Jianke Zhu

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Microsoft(微软) The HongKong Polytechnical University(香港理工大学)

专题命中 VLM训练与架构 :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

Comments CVPR2024, Code and Demo link:https://github.com/CircleRadon/Osprey

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.03221 2025-09-09 cs.CV eess.IV 57%

ADIR: Adaptive Diffusion for Image Reconstruction

Shady Abu-Hussein, Tom Tirer, Raja Giryes

机构 * Tel Aviv University(特拉维夫大学) Bar Ilan University(巴伊兰大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV

Comments Project page https://shadyabh.github.io/ADIR/

Journal ref BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他VLM 2 篇

2508.17675 2025-09-09 cs.LG 79%

Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models

Victoria Yan, Honor Chotkowski, Fengran Wang, Xinhui Li, Carl Yang, Jiaying Lu, Runze Yan, Xiao Hu, Alex Fedorov

机构 * The Westminster Schools(韦斯敏斯特学校) Center for Data Science, Nell Hodgson Woodruff School of Nursing, Emory University(数据科学中心、恩莫森护理学院、埃默里大学) Department of Computer Science, Emory University(计算机科学系、埃默里大学) School of Electrical and Computer Engineering, Georgia Institute of Technology(电气与计算机工程学院、佐治亚理工学院)

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05714 2025-09-09 cs.AI cs.CV 62%

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs

Zhaoyu Fan, Kaihang Pan, Mingze Zhou, Bosheng Qin, Juncheng Li, Shengyu Zhang, Wenqiao Zhang, Siliang Tang, Fei Wu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏