arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-16 至 2025-10-16 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 10 篇

2507.10013 2025-10-16 cs.CV cs.CL 84%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13497 2025-10-16 cs.LG cs.AI 83%

DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation

Zexin Wang, Lin Shi, Haoyu Wu, Junru Luo, Xiangzeng Kong, Jun Qi

机构 * Aliyun School of Big Data, Changzhou University(阿里云大数据学院,长洲大学) Department of Computing, Xi’an JiaoTong-Liverpool University(计算系,西安交通大学-利物浦大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Center for Artificial Intelligence in Agriculture, Fujian Agriculture and Forestry University(农业人工智能中心,福建农林大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14138 2025-10-16 cs.CV cs.AI 81%

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu, Lei Li, Jiahui Gao, Jiyue Jiang, Lingpeng Kong, Chuan Wu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13364 2025-10-16 cs.CV cs.AI 76%

Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity

MingZe Tang, Jubal Chandy Jacob

机构 * Department of Computing Science University of Aberdeen(计算科学系阿伯丁大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12845 2025-10-16 cs.CL cs.AI cs.CV cs.RO 75%

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Jesse Atuhurra, Iqra Ali, Tomoya Iwakura, Hidetaka Kamigaito, Tatsuya Hiraoka

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13235 2025-10-16 cs.CV 70%

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang, Shengsheng Wang

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学教育部长春符号计算与知识工程重点实验室) Yangtze University(扬子大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 62%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15298 2025-10-16 cs.CV cs.MM 62%

MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, Adrian K. Davision

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(计算与数学系,曼彻斯特 Metropolitan 大学) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院) Department of Psychology, University of the Chinese Academy of Sciences(心理学系,中国科学院大学) National Taiwan University(台湾大学) Zhejiang University(浙江大学) University of Oulu(奥卢大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Micro-Expression Grand Challenge (MEGC) at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13359 2025-10-16 cs.IR cs.CV cs.LG 57%

Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models

Yuki Yada, Sho Akiyama, Ryo Watanabe, Yuta Ueno, Yusuke Shido, Andre Rusli

机构 * Mercari, Inc.(Mercari公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM RecSys 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13190 2025-10-16 cs.CL 57%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏