arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-26 至 2025-08-26 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2505.16258 2025-08-26 cs.CL cs.AI cs.CV 85%

IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection

Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee

机构 * College of Information Sciences and Technology(信息科学与技术学院) The Pennsylvania State University(宾夕法尼亚州立大学) Department of Computer Science and Engineering(计算机科学与工程系) National Institute of Technology, Tiruchirappalli(特里奇里帕利理工学院)

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted in the COLM First Workshop on Pragmatic Reasoning in Language Models (PragLM), Montreal, Canada, October 2025, https://sites.google.com/berkeley.edu/praglm

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17037 2025-08-26 cs.CV 83%

F4-ITS: Fine-grained Feature Fusion for Food Image-Text Search

Raghul Asokan

机构 * HyperVerge

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16762 2025-08-26 cs.CL cs.CY 83%

Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

Arka Mukherjee, Shreya Ghosh

机构 * KIIT Deemed University(KIIT大学) Indian Institute of Technology (IIT) Bhubaneswar(印度理工学院(Bhubaneswar分校))

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at ASI @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15969 2025-08-26 cs.CV cs.AI cs.CL 82%

Forgotten Polygons: Multimodal Large Language Models are Shape-Blind

William Rudman, Michal Golovanevsky, Amir Bar, Vedant Palit, Yann LeCun, Carsten Eickhoff, Ritambhara Singh

机构 * Brown University(布朗大学) Tel Aviv University(特拉维夫大学) IIT Kharagpur(印度理工学院卡里帕尔分校) New York University(纽约大学) University of Tübingen(图宾根大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10075 2025-08-26 cs.CV cs.AI cs.LG cs.MM 75%

Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework

Ali Can Karaca, M. Enes Ozelbas, Saadettin Berber, Orkhan Karimli, Turabi Yildirim, M. Fatih Amasyali

机构 * Department of Computer Engineering, Yildiz Technical University(计算机工程系,伊兹密尔技术大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This work has been submitted to the IEEE Transactions on Geoscience and Remote Sensing journal for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18179 2025-08-26 cs.AI cs.CV 73%

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Zhenwei Tang, Difan Jiao, Blair Yang, Ashton Anderson

机构 * Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Coolwei AI Lab(Coolwei人工智能实验室)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17595 2025-08-26 cs.CV 57%

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints

Vinh-Thuan Ly, Hoang M. Truong, Xuan-Huong Nguyen

机构 * University of Science, VNU-HCM(越南胡志明市国家大学) Vietnam National University(越南国家大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted for presentation at the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2025

Journal ref IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Hawaii, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13045 2025-08-26 cs.LG cs.CV 57%

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

Haiyang Guo, Fanhu Zeng, Fei Zhu, Jiayi Wang, Xukai Wang, Jingang Zhou, Hongbo Zhao, Wenzhuo Liu, Shijie Ma, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港科学与创新研究院人工智能与机器人中心) School of Computer and Information Engineering, Xiamen University of Technology(厦门理工大学计算机与信息工程学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15798 2025-08-26 cs.CV 57%

MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining

Ruiqi Wu, Na Su, Chenran Zhang, Tengfei Ma, Tao Zhou, Zhiting Cui, Nianfeng Tang, Tianyu Mao, Yi Zhou, Wen Fan, Tianxing Wu, Shenqi Jing, Huazhu Fu

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科部) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Institute of High-Performance Computing, Agency for Science, Technology and Research(科技研究局高性能计算研究所)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏