arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-15 至 2025-10-15 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2510.11852 2025-10-15 cs.LG 80%

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) Lancaster University(兰卡斯特大学) University of Surrey(萨里大学)

专题命中 图文多模态 :multimodal(title,abstract)

Comments Accepted to ICDMW 2025 Workshop on Multimodal AI (MMAI). Full workshop info: https://icdmw25mmai.github.io/

Journal ref Proc. IEEE International Conference on Data Mining Workshops (ICDMW 2025), Workshop on Multimodal AI (MMAI 2025), Los Angeles, USA, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07214 2025-10-15 cs.CV cs.AI cs.CL 67%

Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions

Akash Ghosh, Arkadeep Acharya, Sriparna Saha, Vinija Jain, Aman Chadha

机构 * Department of Computer Science and Engineering, IIT Patna, India(印度帕纳大学计算机科学与工程系) Stanford University(斯坦福大学) Amazon AI(亚马逊人工智能) Indian Institute of Technology Patna, India(印度帕纳大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments One of the first survey on Visual Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00378 2025-10-15 cs.AI cs.CV 62%

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

Shixin Yi, Lin Shang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper is not yet mature and needs further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏