arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-13 至 2025-10-13 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 10 篇

2510.08606 2025-10-13 cs.CL cs.AI 88%

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun, Yuxuan Ding, Linlin Gong, Leyuan Qu, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究 institute,中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Under review for ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 84%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08589 2025-10-13 cs.CV cs.AI 84%

Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes

Nirmal Elamon, Rouzbeh Davoudi

机构 * Artificial Creative intelligence (ACI)(人工创意智能(ACI)) Expedia Group(Expedia集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.06145 2025-10-13 cs.CV cs.AI cs.HC cs.LG stat.ML 81%

Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1165-1174

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06498 2025-10-13 cs.LG cs.AI cs.CV stat.ML 81%

Deep Multimodal Subspace Clustering Networks

Mahdi Abavisani, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 12, no. 6, pp. 1601-1614, Dec. 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07979 2025-10-13 cs.CV 79%

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) New York University(纽约大学) Chung-Ang University(Chung-Ang 大学) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://cvlab-kaist.github.io/VIRAL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02912 2025-10-13 cs.CV 70%

Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention

Xin Zou, Di Lu, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu, Xu Zheng, Linfeng Zhang, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08802 2025-10-13 cs.LG 67%

Edu-EmotionNet: Cross-Modality Attention Alignment with Temporal Feedback Loops

S M Rafiuddin

机构 * Department of Computer Science Oklahoma State University Stillwater, Oklahoma, USA(计算机科学系 奥克拉荷马州立大学 斯蒂尔沃特 奥克拉荷马州 美国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

Comments 6 Pages, 6 Figures, 3 Tables, Accepted as a Regular Research paper at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09435 2025-10-13 cs.LG cs.IR 50%

Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models

Hyunin Lee, Yong Zhang, Hoang Vu Nguyen, Xiaoyi Liu, Namyong Park, Christopher Jung, Rong Jin, Yang Wang, Zhigang Wang, Somayeh Sojoudi, Xue Feng

机构 * Meta

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25641 2025-10-13 physics.med-ph 50%

Semi-Supervised Radiomics for Glioblastoma IDH Mutation: Limited Labels, Data Sensitivity, and SHAP Interpretation

Amir Hossein Pouria, Shahram Taeb, Somayeh Sadat Mehrnia, Sajad Jabarzadeh Ghandilu, Mehrdad Oveisi, Arman Rahmim, Mohammad R. Salmanpour

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 13 Pages and 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏