arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-12 至 2025-09-12 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 6 篇

2506.19662 2025-09-12 physics.ed-ph 78%

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Giulia Polverini, Bor Gregorcic

专题命中 其他VLM :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09307 2025-09-12 cs.CV cs.AI cs.CL cs.MM 62%

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

Zhengzhao Lai, Youbin Zheng, Zhenyang Cai, Haonan Lyu, Jinpu Yang, Hongqing Liang, Yan Hu, Benyou Wang

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21831 2025-09-12 cs.CV cs.AI 62%

Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization

Anas Anwarul Haq Khan, Utkarsh Verma, Ganesh Ramakrishnan

机构 * Department of Computer Science and Engineering, IIT Bombay(印度理工学院班加罗尔计算机科学与工程系) Center of Machine Intelligence and Data Science (C-MInDS), IIT Bombay(印度理工学院班加罗尔人工智能与数据科学中心)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09541 2025-09-12 cs.AI 57%

Compositional Concept Generalization with Variational Quantum Circuits

Hala Hawashin, Mina Abbaszadeh, Nicholas Joseph, Beth Pearson, Martha Lewis, Mehrnoosh sadrzadeh

机构 * School of Computer Science Engineering University of New South Wales Sydney, Australia Stanford University California, USA Computer Science University College London London, UK School of Eng. Maths. \& Tech University of Bristol Bristol, UK Inst. Logic Language \& Computation University of Amsterdam Amsterdam, NL

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments Accepted to: 2025 IEEE International Conference on Quantum Artificial Intelligence (QAI), Naples, Italy, Nov 2-5, 2025. This is the authors' accepted manuscript (AAM). An IEEE copyright notice appears on page 1. The final published version will appear in IEEE Xplore; DOI to be added when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11538 2025-09-12 cs.CL cs.AI eess.AS 57%

MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Jinyang Wu, Nancy F. Chen, Ai Ti Aw

机构 * MERaLiON Team(MERaLiON团队) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 50%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 其他VLM :multimodal large language model(abstract)

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏