arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-05 至 2025-09-05 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6 篇

2506.09556 2025-09-05 cs.CL 83%

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos

机构 * National Technical University of Athens(希腊国家技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 81%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03999 2025-09-05 cs.CV 79%

SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

Han Huang, Han Sun, Ningzhong Liu, Huiyu Zhou, Jiaquan Shen

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Leicester(莱斯特大学) Luoyang Normal University(洛阳师范学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, accepted by PRCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03837 2025-09-05 cs.LG cs.IT math.IT 78%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03872 2025-09-05 cs.CV 57%

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection

Nan Yang, Yang Wang, Zhanwen Liu, Yuchao Dai, Yang Liu, Xiangmo Zhao

机构 * School of Information Engineering, Chang’an University(信息工程学院,长安大学) School of Electronics and Information, Northwestern Polytechnical University(电子信息学院,西北工业大学) School of Vehicle and Mobility, Tsinghua University(车辆与移动学院,清华大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03635 2025-09-05 cs.CV 57%

Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding

Hongpei Zheng, Lintao Xiang, Qijun Yang, Qian Lin, Hujun Yin

机构 * Department of Electrical and Electronic Engineering, The University of Manchester(曼彻斯特大学电子与电气工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏