arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-29 至 2025-08-29 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 2 篇

2505.10583 2025-08-29 cs.CV cs.CL 79%

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti, Jan Arne Telle, José Hernández-Orallo

机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20660 2025-08-29 eess.AS cs.SD 50%

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏