arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-02 至 2026-02-02 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 14 篇

2510.25801 2026-02-02 cs.LG cs.AI cs.CL cs.CV 85%

Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start

Metis-SPECS: 通过基于偏好自我蒸馏的冷启动解耦多模态学习

Kun Chen, Peng Shi, Haibo Qiu, Zhixiong Zeng, Siqi Yang, Wenji Mao, Lin Ma

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Meituan(美团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Metis-SPECS通过基于偏好的自我蒸馏冷启动框架解耦多模态学习,提升泛化能力和下游RL表现。

Comments Published as a conference paper at ICLR 2026!

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21280 2026-02-02 cs.CV 83%

Token Entropy Regularization for Multi-modal Antenna Affiliation Identification

基于令牌熵正则化的多模态天线归属识别

Dong Chen, Ruoyu Li, Xinyan Zhang, Jialei Xu, Ruosen Zhao, Zhikang Zhang, Lingyun Li, Zizhuang Wei

机构 * Huawei(华为) The University of Hong Kong(香港大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于令牌熵正则化的多模态天线归属识别方法,通过融合视频、几何特征和PCI信号,提升通信网络优化效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22498 2026-02-02 cs.IR 82%

FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning

FITMM: 一种基于信息论的多模态推荐方法

Wei Yang, Rui Zhong, Yiqun Chen, Shixuan Li, Heng Ping, Chi Lu, Peng Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 FITMM通过频域信息论框架提升多模态推荐效果,采用频谱分解与信息瓶颈目标实现频带分离与融合,有效减少冗余并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06777 2026-02-02 cs.CV cs.AI 81%

MolX: Enhancing Large Language Models for Molecular Understanding With A Multi-Modal Extension

MolX: 通过多模态扩展增强大型语言模型的分子理解能力

Khiem Le, Zhichun Guo, Kaiwen Dong, Xiaobao Huang, Bozhao Nan, Roshni Iyer, Xiangliang Zhang, Olaf Wiest, Wei Wang, Ting Hua, Nitesh V. Chawla

机构 * University of Notre Dame, IN, USA(诺丁汉大学) University of California, Los Angeles, CA, USA(加州大学洛杉矶分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 MolX通过多模态扩展提升LLM对分子的理解能力,利用SMILES和分子图提取细粒度特征,并结合分子指纹提升性能,有效提升分子相关任务表现。

Comments MLoG-GenAI@KDD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22729 2026-02-02 cs.CV 79%

GaussianOcc3D: A Gaussian-Based Adaptive Multi-modal 3D Occupancy Prediction

GaussianOcc3D: 一种基于高斯的自适应多模态3D占用预测

A. Enes Doruk, Hasan F. Ates

机构 * Department of Artificial Intelligence and Data Engineering, Ozyegin University(人工智能与数据工程系,奥祖根大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 GaussianOcc3D通过高斯表示实现多模态3D占用预测,提升自动驾驶环境感知的鲁棒性和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17564 2026-02-02 eess.IV cs.CV cs.LG 79%

ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology

ModalTune: 通过多模态信息细调滑片级基础模型以实现数字病理学中的多任务学习

Vishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed, Maged Goubran, Anne L. Martel

机构 * Sunnybrook Research Institute(辛普森布鲁斯研究所在) University of Toronto(多伦多大学) Google Research(谷歌研究)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 ModalTune通过引入多模态信息和大型语言模型,实现数字病理学中多任务学习的统一细调框架,提升癌症生存和亚型预测性能。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22856 2026-02-02 cs.LG 75%

OptiMAG: Structure-Semantic Alignment via Unbalanced Optimal Transport

OptiMAG: 通过不平衡最优传输实现结构-语义对齐

Yilong Zuo, Xunkai Li, Zhihan Zhang, Qiangqiang Dai, Ronghua Li, Guoren Wang

机构 * Beijing Institute of Technology, Beijing, China(北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract)

AI总结 OptiMAG通过不平衡最优传输解决多模态图中结构与语义不一致问题,提升节点表示学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22406 2026-02-02 cs.RO 71%

Accurate Pedestrian Tracking in Urban Canyons: A Multi-Modal Fusion Approach

城市峡谷中精确的人行道跟踪:一种多模态融合方法

Shahar Dubiner, Peng Ren, Roberto Manduchi

专题命中 多模态训练与对齐 :multi-modal(title)

AI总结 本文提出一种多模态融合方法,通过融合GNSS和惯性数据提升城市峡谷中行人定位精度,优于单独使用GNSS或惯性导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23004 2026-02-02 eess.AS 57%

Layer-Aware Early Fusion of Acoustic and Linguistic Embeddings for Cognitive Status Classification

面向层的语音与语言嵌入早期融合用于认知状态分类

Krystof Novotny, Laureano Moro-Velázquez, Jiri Mekyska

专题命中 多模态训练与对齐 :multimodal(abstract);分类 eess.AS

AI总结 本文提出通过语音与语言嵌入的早期融合,结合编码器层深度优化,提升认知状态分类性能,发现中等编码层效果最佳。

Comments 5 pages, 3 figures, paper accepted for ICASSP 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27647 2026-02-02 cs.CV 57%

NegoCollab: A Common Representation Negotiation Approach for Heterogeneous Collaborative Perception

NegoCollab: 一种用于异构协作感知的共同表示协商方法

Congzhang Shao, Quan Yuan, Guiyang Luo, Yue Hu, Danni Wang, Yilin Liu, Rui Pan, Bo Chen, Jinglin Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 NegoCollab通过协商共同表示方法解决异构协作感知中的领域差距问题,提升多智能体协作性能。

Comments 23 pages, Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22696 2026-02-02 cs.CV cs.LG 57%

Bi-MCQ: Reformulating Vision-Language Alignment for Negation Understanding

Bi-MCQ:重新表述视觉-语言对齐以理解否定

Tae Hun Kim, Hyun Gyu Lee

机构 * Department of Electrical and Computer Engineering, Inha University, Republic of Korea(电气与计算机工程系,印哈大学,大韩民国) College of Medicine, Inha University, Republic of Korea(医学学院,印哈大学,大韩民国)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

AI总结 Bi-MCQ通过重新表述视觉-语言对齐为条件语义比较,提升医学VLM对否定理解的性能。

Comments 15 pages, 4 figures, Submitted to ICPR 2026 (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03370 2026-02-02 q-bio.QM cs.AI cs.CE 57%

InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions

InstructPLM-mu:在1小时内微调ESM2在蛋白质突变预测中优于ESM3

Junde Xu, Yapin Shi, Lijun Lang, Taoyong Cui, Zhiming Zhang, Guangyong Chen, Jiezhong Qiu, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Hangzhou Institute of Medicine, CAS(杭州医学研究所,中国科学院) Hangzhou Institute for Advanced Study, UCAS(杭州先进研究 institute,中国科学院大学) University of Chinese Academy of Sciences(中国科学院大学) Tianjin University(天津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 InstructPLM-mu通过1小时微调ESM2在蛋白质突变预测中超越ESM3,揭示了结构输入对模型性能的关键影响。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14957 2026-02-02 cs.CV 57%

DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection

DF-LLaVA: 通过知识注入和冲突驱动的自我反思解锁MLLMs用于合成图像检测

Zhuokang Shen, Kaisen Zhang, Bohan Jia, Heming Jia, Yuan Fang, Zhou Yu, Shaohui Lin

机构 * East China Normal University(东华师范大学) Sanming University(三明大学) The 27th Research Institute of CETC(中国电子科技集团第27研究所)

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CV

AI总结 DF-LLaVA通过知识注入和冲突驱动的自我反思,提升MLLMs在合成图像检测中的准确性和可解释性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20321 2026-02-02 cs.RO 50%

TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation

TaF-VLA:面向力感知操控的视觉-语言-动作模型中的触觉-力对齐

Yuzhe Huang, Pei Lin, Wanlin Li, Daohan Li, Jiajun Li, Jiaming Jiang, Chenxi Xiao, Ziyuan Jiao

机构 * Beihang University(北航大学) ShanghaiTech University(上海科技大学) Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 TaF-VLA通过触觉-力对齐提升视觉-语言-动作模型在力感知操控中的性能。

Comments 17pages,9fig

详情

展开后加载摘要…

URL PDF HTML 收藏