arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2512.03542 2025-12-04 cs.CV cs.AI 84%

V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention

V-ITI: 通过视觉推理时间干预缓解多模态大语言模型中的幻觉

Nan Sun, Zhenyu Zhang, Xixun Lin, Kun Wang, Yanmin Shang, Naibin Gu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Baidu Inc.(百度公司)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 V-ITI通过视觉推理时间干预框架,有效缓解多模态大语言模型中的视觉相关幻觉问题,同时保持任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21002 2025-11-27 cs.CV cs.AI 84%

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

知识完善视觉:一种多模态实体感知检索增强生成框架用于新闻图像描述

Xiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang, Xiaopeng Liu, Min Zhang, Jun Yu

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MERGE提出了一种多模态实体感知检索增强生成框架,通过构建实体中心知识库和改进跨模态对齐,提升新闻图像描述质量和命名实体识别性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12972 2025-11-24 cs.CV cs.AI 84%

Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning

对齐视觉与语言:无需注释的多模态知识图谱构建以增强大语言模型推理

Junming Liu, Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang, Botian Shi

机构 * Tongji University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(华东师范大学) Stanford University(斯坦福大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出 VaLiK 方法,通过跨模态信息补充构建无需注释的多模态知识图谱,提升大语言模型推理能力。

Comments 14 pages, 7 figures, 6 tables; Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13005 2025-11-18 cs.CV cs.AI 84%

SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

Wenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu, Aidong Zhang

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08971 2025-11-13 cs.HC cs.CV cs.MM 84%

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

Sicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun, You He, Jiankang Deng, Hang Zhang, Jifei Song, Zhensong Zhang

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 16 pages, 9 figures, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07080 2025-11-11 cs.CL cs.AI 84%

Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora

Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00427 2025-11-04 cs.CV cs.AI 84%

Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection

Daichi Zhang, Tong Zhang, Jianmin Bao, Shiming Ge, Sabine Süsstrunk

机构 * School of Computer and Communication Sciences, EPFL(瑞士联邦理工学院计算机与通信科学学院) Microsoft Research Asia(微软亚洲研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04448 2025-10-31 cs.CV cs.MM 84%

TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection

Zehong Yan, Peng Qi, Wynne Hsu, Mong Li Lee

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments EMNLP 2025 Oral; Project Homepage: https://yanzehong.github.io/trust-vl/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19311 2025-10-30 cs.CV cs.AI 84%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15510 2025-10-28 cs.CV cs.CL 84%

Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought

Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, Libo Qin

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心) Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院计算与智能研究所) Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University(贵州大学文字计算与认知智能教育部工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) Peking University(北京大学) ByteDance Seed (China)(字节跳动种子(中国))

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025;

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21346 2025-10-27 cs.CV cs.AI 84%

CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments

Lemin Liu, Fangchao Hu, Honghua Jiang, Yaru Chen, Limin Liu, Yongliang Qiao

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09966 2025-10-21 eess.IV cs.AI cs.CV cs.LG 84%

Multimodal Fusion at Three Tiers: Physics-Driven Data Generation and Vision-Language Guidance for Brain Tumor Segmentation

Mingda Zhang

机构 * Software School, Yunnan University, Kunming 650504, Yunnan, China(云南大学软件学院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 31 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10013 2025-10-16 cs.CV cs.CL 84%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09815 2025-10-14 cs.CV cs.AI 84%

Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning

Yufei Wang, Adriana Kovashka, Loretta Fernández, Marc N. Coutanche, Seth Wiener

机构 * University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to International Conference on Development and Learning (ICDL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14715 2025-10-07 cs.CV cs.AI 84%

Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models

Young Kyun Jang, Ser-nam Lim

机构 * Google DeepMind(谷歌DeepMind) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20961 2025-09-26 cs.CV cs.AI 84%

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya

机构 * Indian Institute of Technology Patna(印度帕纳杰大学) CRISIL LTD(CRISIL公司)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05244 2025-09-22 cs.CV cs.AI 84%

RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding

Tianchen Fang, Guiru Liu

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Upon further review, we identified that our dataset requires optimization to ensure research reliability and accuracy. Additionally, considering the target journal's latest submission policies, we believe comprehensive manuscript revisions are necessary

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15688 2025-09-16 cs.CL cs.AI cs.LG 84%

Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts

Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang

机构 * Department of Computer and Information Science and Engineering, University of Florida(计算机与信息科学与工程系,佛罗里达大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06569 2025-08-21 cs.CL cs.AI 84%

Social Debiasing for Fair Multi-modal LLMs

Harry Cheng, Yangyang Guo, Qingpei Guo, Ming Yang, Tian Gan, Weili Guan, Liqiang Nie

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Project page: https://github.com/xaCheng1996/Social_Debiasing_For_Fair_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19925 2025-08-14 cs.CL cs.CV cs.LG 84%

Improving Multimodal Large Language Models Using Continual Learning

Shikhar Srivastava, Md Yousuf Harun, Robik Shrestha, Christopher Kanan

机构 * University of Rochester(罗切斯特大学) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments CoLLAs 2025 and Scalable Continual Learning for Lifelong Foundation Models, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19015 2025-08-11 cs.CV cs.MM 84%

Can Multimodal Large Language Models Understand Spatial Relations?

Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou, Yinan Zou, Weiyan Zhang, Haiyun Jiang, Tong Ruan

机构 * School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China(信息科学与工程学院,东华大学,上海,中国) School of Computer Science, Fudan University, Shanghai, China(计算机科学学院,复旦大学,上海,中国)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures, published to ACL 2025

Journal ref In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 620-632, 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12591 2025-08-06 cs.CV cs.CL 84%

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao, Thong Nguyen, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) National University of Singapore, Singapore(国立新加坡大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18659 2025-08-01 cs.CV cs.AI 84%

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学、腾讯) Tencent(腾讯) Tencent, University of Macau(腾讯、澳门大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20913 2025-07-29 cs.CV cs.AI 84%

HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

Jialei Cui, Jianwei Du, Yanzhe Li, Lei Gao, Hui Jiang, Chenfu Bao

机构 * Baidu Inc.(百度公司) Southeast University(东南大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20156 2025-07-29 cs.CV cs.AI 84%

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

Daulet Toibazar, Kesen Wang, Sherif Mohamed, Abdulaziz Al-Badawi, Abdulrahman Alfulayt, Pedro J. Moreno

机构 * Humain Riyadh, KSA(利雅得人类,沙特阿拉伯)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17080 2025-07-24 cs.IR cs.AI cs.CV 84%

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

Ramin Giahi, Kehui Yao, Sriram Kollipara, Kai Zhao, Vahid Mirjalili, Jianpeng Xu, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at RecSys 2025; DOI:https://doi.org/10.1145/3705328.3748064

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10202 2025-07-15 cs.CV cs.AI 84%

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Jaeseong Lee, Yeeun Choi, Heechan Choi, Hanjung Kim, Seonjoo Kim

机构 * Yonsei University(延世大学)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CVPR 2025 Workshop on Emergent Visual Abilities and Limits of Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15251 2025-07-08 cs.CL cs.AI 84%

AgentPS: Agentic Process Supervision for Content Moderation with Multimodal LLMs

Mingchao Liu, Yu Sun, Ruixiao Sun, Xin Dong, Xiang Shen, Hongyu Xiong

机构 * TikTok, Inc.(字节跳动公司)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16760 2025-06-23 cs.CL cs.CV 84%

Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models

Lei Jiang, Zixun Zhang, Zizhou Wang, Xiaobing Sun, Zhen Li, Liangli Zhen, Xiaohua Xu

机构 * University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Institute of High Performance Computing, A*STAR, Singapore(新加坡A*STAR高性能计算研究所)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15645 2025-06-19 cs.CV cs.AI 84%

Demystifying the Visual Quality Paradox in Multimodal Large Language Models

Shuo Xing, Lanqing Guo, Hongyuan Hua, Seoyoung Lee, Peiran Li, Yufei Wang, Zhangyang Wang, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏