arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-18 至 2025-08-18 共收录 42 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2508.11616 2025-08-18 cs.CV cs.AI cs.CL cs.LG 85%

Controlling Multimodal LLMs via Reward-guided Decoding

Oscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano, Michal Drozdzal, Aishwarya Agrawal

机构 * Mila - Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) Meta FAIR Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00456 2025-08-18 eess.SP 82%

When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework

Ji Wang, Bin Tang, Jian Xiao, Qimei Cui, Xingwang Li, Tony Q. S. Quek

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05422 2025-08-18 cs.CV cs.AI cs.CL 82%

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Haokun Lin, Teng Wang, Yixiao Ge, Yuying Ge, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG广告实验室) City University of Hong Kong(香港城市大学) Zhejiang University(浙江大学) NLPR & MAIS, Institute of Automation, CAS, Beijing(自动化研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11218 2025-08-18 cs.CV cs.LG 70%

A CLIP-based Uncertainty Modal Modeling (UMM) Framework for Pedestrian Re-Identification in Autonomous Driving

Jialin Li, Shuqi Wu, Ning Wang

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11317 2025-08-18 cs.CV cs.MM 62%

Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Yuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao, Yuqin Dai, Wenhao Yang, Chao Gou, Xiaobo Xia, Tat-Seng Chua

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09346 2025-08-18 cs.CV cs.AI 62%

B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, Kaipeng Zhang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Osaka University(大阪大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11341 2025-08-18 cs.CV cs.CR cs.LG 57%

Semantically Guided Adversarial Testing of Vision Models Using Language Models

Katarzyna Filus, Jorge M. Cruz-Duarte

机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(波兰科学院理论与应用信息学研究所) University of Lille, CNRS, Inria, Centrale Lille, UMR 9189 CRIStAL(里尔大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 12 pages, 4 figures, 3 tables. Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2508.11362 2025-08-18 cs.SD eess.AS 79%

Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge

Honghong Wang, Yankai Wang, Dejun Zhang, Jing Deng, Rong Zheng

机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 2 pages. pubilshed by ICASSP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11187 2025-08-18 eess.AS cs.CL cs.SD 62%

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

Wonjune Kang, Deb Roy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11255 2025-08-18 cs.CV 57%

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

MengChao Wang, Qiang Wang, Fan Jiang, Mu Xu

机构 * Project Leader(项目负责人)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments https://fantasy-amap.github.io/fantasy-talking2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11074 2025-08-18 cs.SD cs.AI cs.CV eess.AS 56%

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

Haomin Zhang, Kristin Qi, Shuxin Yang, Zihao Chen, Chaofan Ding, Xinhan Di

机构 * Giant Network, China(中国巨网) Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系)

专题命中 音频语音多模态 :分类 cs.CV、cs.AI、eess.AS;audio-visual(comments)

Comments Gen4AVC@ICCV: 1st Workshop on Generative AI for Audio-Visual Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2508.11197 2025-08-18 cs.CL cs.AI cs.LG cs.SI 84%

E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection

Ahmad Mousavi, Yeganeh Abdollahinejad, Roberto Corizzo, Nathalie Japkowicz, Zois Boukouvalas

机构 * Department of Mathematics and Statistics, American University, Washington, DC, USA(数学与统计学系,美国大学,华盛顿特区,美国) Department of Computer Science and Mathematics, Pennsylvania State University, Harrisburg, PA, USA(计算机科学与数学系,宾夕法尼亚州立大学,哈里斯堡,宾夕法尼亚州,美国) Department of Computer Science, American University, Washington, DC, USA(计算机科学系,美国大学,华盛顿特区,美国)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11092 2025-08-18 cs.LG 82%

Predictive Multimodal Modeling of Diagnoses and Treatments in EHR

Cindy Shih-Ting Huang, Clarence Boon Liang Ng, Marek Rei

机构 * Imperial College London(帝国理工学院伦敦分校)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 79%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11537 2025-08-18 cs.RO 78%

MultiPark: Multimodal Parking Transformer with Next-Segment Prediction

Han Zheng, Zikang Zhou, Guli Zhang, Zhepei Wang, Kaixuan Wang, Peiliang Li, Shaojie Shen, Ming Yang, Tong Qin

机构 * Shanghai Jiao Tong University(上海交通大学) Zhuoyu Technology, Co., Ltd.(珠海宇科技有限公司) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港理工大学电子与计算机工程系)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01712 2025-08-18 cs.CV cs.AI 62%

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

Han Wang, Zhuoran Wang, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07992 2025-08-18 cs.MM 57%

Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos

Haisong Gong, Bolan Su, Xinrong Zhang, Jing Li, Qiang Liu, Shu Wu, Liang Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments in submission

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2508.11141 2025-08-18 cs.CV cs.AI cs.CL 82%

A Cross-Modal Rumor Detection Scheme via Contrastive Learning by Exploring Text and Image internal Correlations

Bin Ma, Yifei Zhang, Yongjin Xian, Qi Li, Linna Zhou, Gongxun Miao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 62%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20844 2025-08-18 cs.IR cs.CL 57%

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers

Xingyu Deng, Xi Wang, Mark Stevenson

机构 * University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted for ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11246 2025-08-18 cs.ET cs.IR 50%

RAG for Geoscience: What We Expect, Gaps and Opportunities

Runlong Yu, Shiyuan Luo, Rahul Ghosh, Lingyao Li, Yiqun Xie, Xiaowei Jia

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11210 2025-08-18 cs.LG stat.ML 50%

Borrowing From the Future: Enhancing Early Risk Assessment through Contrastive Learning

Minghui Sun, Matthew M. Engelhard, Benjamin A. Goldstein

机构 * Department of Biostatistics & Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学)

专题命中 跨模态检索 :multi-modal(abstract)

Comments accepted by Machine Learning for Healthcare 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2508.11159 2025-08-18 cs.LG 78%

Mitigating Modality Quantity and Quality Imbalance in Multimodal Online Federated Learning

Heqiang Wang, Weihong Yang, Xiaoxiong Zhong, Jia Zhou, Fangming Liu, Weizhe Zhang

机构 * Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(title,abstract)

Comments arXiv admin note: text overlap with arXiv:2505.16138

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 57%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05908 2025-08-18 cs.CV 57%

GBR: Generative Bundle Refinement for High-fidelity Gaussian Splatting with Enhanced Mesh Reconstruction

Jianing Zhang, Yuchao Zheng, Ziwei Li, Qionghai Dai, Xiaoyun Yuan

机构 * College of future information technology, Fudan University(未来信息科技学院,复旦大学) School of Biomedical Engineering, Tsinghua University(生物医学工程学院,清华大学) Key Laboratory for Information Science of Electromagnetic Waves (MoE), Fudan University(电磁波信息科学重点实验室(MoE),复旦大学) Department of Automation, Tsinghua University(自动化系,清华大学) MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能重点实验室,人工智能研究院,上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2508.10955 2025-08-18 cs.CV cs.CL cs.MM 85%

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

Wenbin An, Jiahao Nie, Yaqiang Wu, Feng Tian, Shijian Lu, Qinghua Zheng

机构 * School of Automation Science and Engineering, Xi’an Jiaotong University(自动化科学与工程学院,西安交通大学) Interdisciplinary Graduate Programme, Nanyang Technological University(跨学科研究生项目,南洋理工大学) Lenovo Research, Lenovo(联想研究院,联想) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.MM

Comments 21 pages, 361 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11021 2025-08-18 cs.CV cs.CL 81%

Can Multi-modal (reasoning) LLMs detect document manipulation?

Zisheng Liang, Kidus Zewde, Rudra Pratap Singh, Disha Patil, Zexi Chen, Jiayu Xue, Yao Yao, Yifei Chen, Qinzhe Liu, Simiao Ren

机构 * Duke University(杜克大学) Indian Institute of Technology, Roorkee(印度理工学院,罗尔基) New York University(纽约大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Wisconsin Madison(威斯康星大学麦迪逊分校) Columnbia University(哥伦比亚大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments arXiv admin note: text overlap with arXiv:2503.20084

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18932 2025-08-18 cs.MM cs.CL 81%

MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks

Lei Zhang, Xin Zhou, Chaoyue He, Di Wang, Yi Wu, Hong Xu, Wei Liu, Chunyan Miao

机构 * Nanyang Technological University(南洋理工大学) University College London(伦敦大学学院) Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10916 2025-08-18 cs.HC cs.AI cs.CY cs.MA 74%

Multimodal Quantitative Measures for Multiparty Behaviour Evaluation

Ojas Shirekar, Wim Pouw, Chenxu Hao, Vrushank Phadnis, Thabo Beeler, Chirag Raman

机构 * TU Delft(代尔夫特理工大学) Tilburg University(蒂尔堡大学) Google(谷歌公司)

专题命中 多模态评测 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18057 2025-08-18 cs.CV cs.CL 73%

CLEAR: Character Unlearning in Textual and Visual Modalities

Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin, Boris Mikheev, Denis Bobkov, Aibek Alanov, Oleg Y. Rogov, Ivan Oseledets, Elena Tutubalina

机构 * AIRI HSE University(俄罗斯高等经济大学) Skoltech(斯克里普契克技术大学) Sber AI(俄罗斯储蓄银行人工智能实验室) MTUCI(莫斯科国立大学) MIPT(米哈伊尔戈尔斯基理工学院) ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Journal ref https://aclanthology.org/2025.findings-acl.1058/

详情

展开后加载摘要…

URL PDF HTML 收藏