arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-17 至 2025-11-17 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 12 篇

2511.10671 2025-11-17 cs.CL cs.CV 84%

Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency

Filippo Morbiato, Luca Romano, Alessandro Persona

机构 * University of Padua(帕多瓦大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10568 2025-11-17 cs.LG cs.AI cs.CL cs.CV 82%

MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Ruixiang Jiang, Lingbo Liu, Changwen Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Research Institute of Multiple Agents and Embodied Intelligence, Pengcheng Laboratory(多智能体与具身智能研究院,鹏城实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to IEEE TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11452 2025-11-17 q-bio.QM cs.CV cs.LG eess.IV 79%

Synergy vs. Noise: Performance-Guided Multimodal Fusion For Biochemical Recurrence-Free Survival in Prostate Cancer

Seth Alain Chang, Muhammad Mueez Amjad, Noorul Wahab, Ethar Alzaid, Nasir Rajpoot, Adam Shephard

机构 * Tissue Image Analytics Centre, Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系组织图像分析中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26289 2025-11-17 cs.MM 79%

Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise

Zijing Xu, Yunfeng Kou, Kunming Wu, Hong Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06452 2025-11-17 cs.LG 78%

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains

Leyan Xue, Changqing Zhang, Kecheng Xue, Xiaohong Liu, Guangyu Wang, Zongbo Han

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04870 2025-11-17 cs.LG 78%

NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification

Jun Hu, Yufei He, Yuan Li, Bryan Hooi, Bingsheng He

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11126 2025-11-17 cs.CL cs.CV 73%

Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion

Yi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei, Wenzhi Chen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10046 2025-11-17 cs.CV 70%

FreDFT: Frequency Domain Fusion Transformer for Visible-Infrared Object Detection

Wencong Wu, Xiuwei Zhang, Hanlin Yin, Shun Dai, Hongxi Zhang, Yanning Zhang

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11422 2025-11-17 cs.CV 57%

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

Lukun Wu, Jie Li, Ziqi Ren, Kaifan Zhang, Xinbo Gao

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 21pages,12 figures,published to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10560 2025-11-17 cs.CV 57%

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

Haosong Peng, Hao Li, Yalun Dai, Yushi Lan, Yihang Luo, Tianyu Qi, Zhengshen Zhang, Yufeng Zhan, Junfei Zhang, Wenchao Xu, Ziwei Liu

机构 * HKUST(香港科技大学) NTU(国立台湾大学) SYSU(南方科技大学) NUS(国立新加坡大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://livioni.github.io/OmniVGGT-official/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15888 2025-11-17 cs.CV 57%

MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction

Zhiqiang Wei, Lianqing Zheng, Jianan Liu, Tao Huang, Qing-Long Han, Wenwen Zhang, Fengdeng Zhang

机构 * School of Optical-Electrical and Computer Engineering, University of Shanghai for Science and Technology(光学电子与计算机工程学院,上海科学技术大学) School of Automotive Studies, Tongji University(汽车学院,同济大学) Momoni AI College of Science and Engineering, James Cook University(科学与工程学院,詹姆斯库克大学) School of Engineering, Swinburne University of Technology(工程学院,斯威本技术大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电气与电子工程学院,南洋理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10935 2025-11-17 cs.SD cs.LG q-bio.NC 50%

CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding

Yifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou, Jiawei Ju

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments This is the extended version with technical appendices. The version of record appears in AAAI-26. Please cite the AAAI version

详情

展开后加载摘要…

URL PDF HTML 收藏