arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-01 至 2025-10-01 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 6 篇

2509.11662 2025-10-01 cs.CV cs.AI cs.CL eess.IV 84%

MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs

Feilong Chen, Yijiang Liu, Yi Huang, Hao Wang, Miren Tian, Ya-Qi Yu, Minghui Liao, Jihao Wu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26161 2025-10-01 cs.AI cs.SE 79%

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 其他VLM :MLLM(title);multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25818 2025-10-01 cs.CV cs.AI cs.CL 62%

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura

机构 * Keio University(庆应大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26555 2025-10-01 cs.CV 57%

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani

机构 * Stability AI Arizona State University(亚利桑那州立大学) Google DeepMind(谷歌DeepMind)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2025. Project Page : https://stable-cinemetrics.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25863 2025-10-01 cs.CV 57%

MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification

Junjie Zhou, Wei Shao, Yagao Yue, Wei Mu, Peng Wan, Qi Zhu, Daoqiang Zhang

机构 * The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 57%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏