arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 25 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 25 篇

2511.12982 2025-11-18 cs.CR cs.CV 83%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17686 2025-11-18 cs.CV 83%

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Yuhang Han, Xuyang Liu, Zihan Zhang, Pengxiang Ding, Junjie Chen, Donglin Wang, Honggang Chen, Qingsen Yan, Siteng Huang

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00374 2025-11-18 cs.CV cs.AI cs.MM 82%

MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention

Tianyi Wang, Jianan Fan, Dingxin Zhang, Dongnan Liu, Yong Xia, Heng Huang, Weidong Cai

机构 * The University of Sydney(悉尼大学) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) School of Computer Science and Engineering, Northwestern Polytechnical University(西北工业大学计算机科学与工程学院) University of Maryland(马里兰大学) Ningbo Institute of Northwestern Polytechnical University(西北工业大学宁波学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by IEEE Transactions on Medical Imaging (TMI). Code available at https://github.com/TianyiFranklinWang/MIRROR. Project page: https://tianyifranklinwang.github.io/MIRROR

Journal ref IEEE Trans. Med. Imaging (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10218 2025-11-18 cs.AI 79%

MTP: Exploring Multimodal Urban Traffic Profiling with Modality Augmentation and Spectrum Fusion

Haolong Xiang, Peisi Wang, Xiaolong Xu, Kun Yi, Xuyun Zhang, Quanzheng Sheng, Amin Beheshti, Wei Fan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13637 2025-11-18 cs.LG 79%

Towards Multimodal Representation Learning in Paediatric Kidney Disease

Ana Durica, John Booth, Ivana Drobnjak

机构 * Institute of Health Informatics(健康信息学研究所) University College London(伦敦大学学院) Data Research, Innovation and Virtual Environments Unit(数据研究、创新与虚拟环境单位) Great Ormond Street Hospital(格雷特奥蒙德医院) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 4 pages, 3 figures. EurIPS 2025 Multimodal Representation Learning for Healthcare (MMRL4H) workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10962 2025-11-18 cs.IR 78%

LEMUR: Large scale End-to-end MUltimodal Recommendation

Xintian Han, Honggang Chen, Quan Lin, Jingyue Gao, Xiangyuan Ren, Lifei Zhu, Zhisheng Ye, Shikang Wu, XiongHang Xie, Xiaochu Gan, Bingzheng Wei, Peng Xu, Zhe Wang, Yuchao Zheng, Jingjian Lin, Di Wu, Junfeng Ge

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10644 2025-11-18 cs.LG 78%

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

Yihua Wang, Qi Jia, Cong Xu, Feiyu Chen, Yuhan Liu, Haotian Zhang, Liang Jin, Lu Liu, Zhichun Wang

机构 * IEIT SYSTEMS Co., Ltd.(IEIT SYSTEMS公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at AAAI 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12917 2025-11-18 cs.CV 77%

Explore How to Inject Beneficial Noise in MLLMs

Ruishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13283 2025-11-18 cs.CV 70%

TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing

Jongha Kim, Minseong Bae, Sanghyeok Lee, Jinsung Yoon, Hyunwoo J. Kim

机构 * Korea University(韩国大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments AAAI 2026 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13575 2025-11-18 cs.CV cs.AI 62%

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

Linhan Zhou, Shuang Li, Neng Dong, Yonghang Tai, Yafei Zhang, Huafeng Li

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11730 2025-11-18 cs.CV cs.AI 62%

GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion

Yongjun Xiao, Dian Meng, Xinlei Huang, Yanran Liu, Shiwei Ruan, Ziyue Qiao, Xubin Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11622 2025-11-18 cs.LG cs.AI cs.CL 62%

Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

Alexis Roger, Gwen Legate, Kashif Rasul, Yuriy Nevmyvaka, Irina Rish

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13150 2025-11-18 cs.CV 57%

Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification

Rifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13047 2025-11-18 cs.CV cs.RO 57%

DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation

Yan Gong, Jianli Lu, Yongsheng Gao, Jie Zhao, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学) Institute for Infocomm Research, A*STAR(信息通信研究机构,A*STAR) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12525 2025-11-18 cs.CV 57%

MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics

Jing Li, Yifan Wang, Jiafeng Yan, Renlong Zhang, Bin Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 7 figures. Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12368 2025-11-18 cs.CV 57%

Fast Reasoning Segmentation for Images and Videos

Yiqing Shen, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰·霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12365 2025-11-18 cs.CV 57%

Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning

Yiqing Shen, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12220 2025-11-18 cs.CV cs.LG 57%

Suppressing VLM Hallucinations with Spectral Representation Filtering

Ameen Ali, Tamim Zoabi, Lior Wolf

机构 * Tel Aviv University(特拉维夫大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06298 2025-11-18 cs.CV 57%

SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection

Xin Zuo, Chenyu Qu, Haibo Zhan, Jifeng Shen, Wankou Yang

机构 * School of Computer, Jiangsu University of Science and Technology(江苏科技大学计算机学院) School of Electrical and Information Engineering, Jiangsu University(江苏大学电气与信息工程学院) School of Automation, Southeast University(东南大学自动化学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 11 pages,8 figures, accepted by IEEE TGRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11750 2025-11-18 cs.LG cs.AI 57%

IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation

Hanting Yan, Pan Mu, Shiqi Zhang, Yuchao Zhu, Jinglin Zhang, Cong Bai

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05933 2025-11-18 eess.IV cs.CV 57%

Beyond H&E: Unlocking Pathological Insights with Polarization Imaging

Yao Du, Jiaxin Zhuang, Xiaoyu Zheng, Jing Cong, Limei Guo, Chao He, Lin Luo, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科技大学) Beijing Institute of Collaborative Innovation(北京协同创新研究院) Peking University Health Science Center, Peking University Third Hospital(北京大学人民医院) University of Oxford(牛津大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a regular paper at IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13327 2025-11-18 cs.RO 50%

ZeroDexGrasp: Zero-Shot Task-Oriented Dexterous Grasp Synthesis with Prompt-Based Multi-Stage Semantic Reasoning

Juntao Jian, Yi-Lin Wei, Chengjie Mou, Yuhao Lin, Xing Zhu, Yujun Shen, Wei-Shi Zheng, Ruizhen Hu

机构 * Shenzhen University(深圳大学) Sun Yat-sen University(中山大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12778 2025-11-18 cs.RO 50%

DR. Nav: Semantic-Geometric Representations for Proactive Dead-End Recovery and Navigation

Vignesh Rajagopal, Kasun Weerakoon Kulathun Mudiyanselage, Gershom Devake Seneviratne, Pon Aswin Sankaralingam, Mohamed Elnoor, Jing Liang, Rohan Chandra, Dinesh Manocha

机构 * Dept. of Computer Science at the University of Virginia(弗吉尼亚大学计算机科学系) Dept. of Computer Science at the University of Maryland College Park(马里兰大学学院市计算机科学系)

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12729 2025-11-18 eess.SP 50%

Bridging the Modality Gap: Enhancing Channel Prediction with Semantically Aligned LLMs and Knowledge Distillation

Zhaoyang Li, Qianqian Yang, Zehui Xiong, Zhiguo Shi, Tony Q. S. Quek

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11673 2025-11-18 cs.LG 50%

Synergistic Feature Fusion for Latent Lyrical Classification: A Gated Deep Learning Architecture

M. A. Gameiro

机构 * Independent Researcher(独立研究者)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏