arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-17 至 2025-10-17 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 5 篇

2510.14624 2025-10-17 cs.CV 79%

Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference

Natan Bagrov, Eugene Khvedchenia, Borys Tymchenko, Shay Aharon, Lior Kadoch, Tomer Keren, Ofri Masad, Yonatan Geifman, Ran Zilberstein, Tuomas Rintamaki, Matthieu Le, Andrew Tao

机构 * NVIDIA

专题命中 VLM训练与架构 :VLM(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14968 2025-10-17 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 75%

RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks

Mingxuan Yan, Yuping Wang, Zechun Liu, Jiachen Li

机构 * University of California, Riverside(加州大学河滨分校) University of Michigan(密歇根大学) Meta AI

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025); Project Website: rdd-neurips.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14526 2025-10-17 cs.CV cs.LG 73%

Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models

Yunze Tong, Didi Zhu, Zijing Hu, Jinluan Yang, Ziyu Zhao

机构 * Zhejiang University(浙江大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG

Comments Appendix will be appended soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17692 2025-10-17 cs.CL cs.AI cs.LG 62%

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang

机构 * Beihang University(北航) M-A-P The Hong Kong Polytechnic University(香港理工大学) AIWaves University of Alberta(阿尔伯塔大学) University of Waterloo(滑铁卢大学) University of Manchester(曼彻斯特大学) Chinese Academy of Sciences(中国科学院) Peking University(北京大学) Shanghai AI Lab(上海AI实验室) Nanjing University(南京大学) Kuaishou Technology(快手科技)

专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.AI、cs.LG

Comments EMNLP 2025 (Oral). Codes and models are available in https://github.com/MIO-Team/MIO

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15693 2025-10-17 cs.CV cs.MM 57%

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

Cristian Sbrolli, Matteo Matteucci

机构 * Department of Electronics, Information and Bioengineering(电子、信息与生物工程系)

专题命中 VLM训练与架构 :visual question answering(abstract);分类 cs.CV

Comments to appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏