arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-06 至 2025-10-06 共收录 33 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 4 篇

2510.02815 2025-10-06 cs.CV 57%

Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis

Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) Chinese Academy of Sciences(中国科学院) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR2026 under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02787 2025-10-06 cs.CV 57%

OTR: Synthesizing Overlay Text Dataset for Text Removal

Jan Zdenek, Wataru Shimoda, Kota Yamaguchi

机构 * CyberAgent Tokyo Japan(CyberAgent东京日本)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09047 2025-10-06 cs.CL 50%

Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

Yaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) UC Berkeley(加州大学伯克利分校)

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏