arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-24 至 2025-10-24 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2510.14349 2025-10-24 cs.CV 83%

Vision-Centric Activation and Coordination for Multimodal Large Language Models

Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang, Ziqiang Li, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(人工智能MoE实验室,上海交通大学) Ant Group(蚂蚁集团) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究所,东部技术研究所,宁波)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20736 2025-10-24 cs.LG 82%

Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process

Tsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin, Lequan Yu

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by NeruIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20540 2025-10-24 cs.LG 82%

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis

机构 * Center for Wireless Communications, University of Oulu(无线通信中心,奥卢大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 5 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20256 2025-10-24 cs.CV cs.CL cs.LG cs.MM 82%

Calibrating Multimodal Consensus for Emotion Recognition

Guowei Zhong, Junjie Li, Huaiyu Zhu, Ruohong Huan, Yun Pan

机构 * College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学) Zhejiang University Jinhua Research Institute(浙江大学金华研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19201 2025-10-24 cs.CL 79%

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院) Tandon School of Engineering, New York University(纽约大学工程学院) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05993 2025-10-24 cs.IR 79%

Efficient Multimodal Streaming Recommendation via Expandable Side Mixture-of-Experts

Yunke Qu, Liang Qu, Tong Chen, Quoc Viet Hung Nguyen, Hongzhi Yin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted to CIKM 2025. Code is available at https://github.com/qykcq/Efficient-Multimodal-Streaming-Recommendation-via-Expandable-Side-Mixture-of-Experts

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20347 2025-10-24 cs.RO cs.MA 71%

Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots

Ashutosh Mishra, Shreya Santra, Elian Neppel, Edoardo M. Rossi Lombardi, Shamistan Karimov, Kentaro Uno, Kazuya Yoshida

机构 * Space Robotics Lab. (SRL) in Department of Aerospace Engineering, Graduate School of Engineering, Tohoku University(太空机器人实验室(SRL)位于东东北大学航空航天工程系研究生院) Politecnico di Milano(米兰理工大学)

专题命中 多模态训练与对齐 :multi-modal(title)

Comments Accepted in IEEE iSpaRo 2025. Awaiting Publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 70%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22038 2025-10-24 cs.CV cs.AI 62%

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

Kaiyuan Li, Xiaoyue Chen, Chen Gao, Yong Li, Xinlei Chen

机构 * Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) BNRist, Tsinghua University(清华大学北京研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20345 2025-10-24 cs.AI 57%

LLM-empowered knowledge graph construction: A survey

Haonan Bian

机构 * Xidian University(西安电子科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20615 2025-10-24 cs.LG 50%

MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation

Yang Han, Pengyu Wang, Kai Yu, Xin Chen, Lu Chen

机构 * X-LANCE Lab, School of Computer Science MoE Key Lab of Artificial Intelligence, SJTU AI Institute Shanghai Jiao Tong University, Shanghai, China(X-LANCE实验室,计算机科学学院,人工智能教育部重点实验室,上海交通大学人工智能研究院,上海,中国) Suzhou Laboratory, Suzhou, China(苏州实验室,苏州,中国) Shanghai Innovation Institute, Shanghai, China(上海创新研究院,上海,中国) Jiangsu Key Lab of Language Computing, Suzhou, China(江苏省语言计算重点实验室,苏州,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract)

Comments NeurIPS 2025, We provide the data and code at https://github.com/OpenDFM/MS-BART

详情

展开后加载摘要…

URL PDF HTML 收藏