arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 20 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 20 篇

2510.11175 2025-10-14 cs.CV 83%

Reliable Cross-modal Alignment via Prototype Iterative Construction

Xiang Ma, Litian Xu, Lexin Fang, Caiming Zhang, Lizhen Cui

机构 * Shandong University(山东大学) The University of Exeter(埃克塞特大学) The Joint SDU-NTU Centre for Artificial Intelligence Research(SDU-NTU联合人工智能研究中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17040 2025-10-14 cs.CV 83%

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

机构 * Peking University(北京大学) Northeastern University(东北大学) Sydney Smart Technology College(悉尼智能技术学院) School of Computer Science, Peking University(北京大学计算机学院) The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IJCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15595 2025-10-14 cs.RO 82%

Grasping Deformable Objects via Reinforcement Learning with Cross-Modal Attention to Visuo-Tactile Inputs

Yonghyun Lee, Sungeun Hong, Min-gu Kim, Gyeonghwan Kim, Changjoo Nam

机构 * Dept. of Electronic Engineering at Sogang University(ソガン大学电子工程系) Dept. of Immersive Media and Engineering at Sungkyunkwan University(顺天大学沉浸媒体与工程系) College of Medicine, Yonsei University(延世大学医学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10406 2025-10-14 cs.CV cs.AI cs.LG 81%

Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes

Zhao-Yang Wang, Jieneng Chen, Jiang Liu, Yuxiang Guo, Rama Chellappa

机构 * Johns Hopkins University(约翰霍普金斯大学) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11693 2025-10-14 cs.CL cs.AI cs.CV 80%

Scaling Language-Centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11112 2025-10-14 cs.CV 79%

Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment

Chen Liu, Wenfang Yao, Kejing Yin, William K. Cheung, Jing Qin

机构 * School of Nursing, The Hong Kong Polytechnic University(香港理工大学护理学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10524 2025-10-14 cs.CV 79%

Unified Open-World Segmentation with Multi-Modal Prompts

Yang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen, Yuling Xi, Bo Feng, Hao Wang, Shiyu Li, Chunhua Shen

机构 * Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Technology(浙江工业大学) Apple(苹果公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06538 2025-10-14 cs.CL 79%

Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model

Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通联合实验室(北京交通大学)) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11666 2025-10-14 eess.SP cs.LG 78%

Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion

Timothy Oladunni, Ehimen Aneni

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19147 2025-10-14 cs.CL cs.AI cs.CV 67%

Shifting AI Efficiency From Model-Centric to Data-Centric Compression

Xuyang Liu, Zichen Wen, Shaobo Wang, Junjie Chen, Zhishan Tao, Yubo Wang, Tailai Chen, Xiangqi Jin, Chang Zou, Yiyu Wang, Chenfei Liao, Xu Zheng, Honggang Chen, Weijia Li, Xuming Hu, Conghui He, Linfeng Zhang

机构 * EPIC Lab, Shanghai Jiao Tong University(上海交通大学EPIC实验室) Sichuan University(四川大学) University of Electronic Science & Technology of China(电子科技大学) Shanghai AI Laboratory(上海人工智能实验室) Sun Yat-sen University(中山大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project: \url{https://github.com/xuyang-liu16/Awesome-Token-level-Model-Compression}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23590 2025-10-14 cs.CV cs.AI cs.CL 67%

Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles

Zifu Wang, Junyi Zhu, Bo Tang, Zhiyu Li, Feiyu Xiong, Jiaqian Yu, Matthew B. Blaschko

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11632 2025-10-14 cs.CV cs.AI cs.LG 62%

NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection

Krittin Chaowakarn, Paramin Sangwongngam, Nang Htet Htet Aung, Chalie Charoenlarpnopparut

机构 * The School of Information, Computer, and Communication Technology, Sirindhorn International Institute of Technology, Thammasat University(信息、计算机与通信技术学院,Sirindhorn国际技术学院,泰国朱拉隆梭大学) National Electronics and Computer Technology Center, National Science and Technology Development Agency(国家电子与计算机技术中心,国家科学技术发展局) Department of Electrical Engineering, Faculty of Engineering, Chulalongkorn University(电气工程系,工程学院,朱拉隆梭大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10264 2025-10-14 cs.CV cs.AI 62%

MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs

Haonan Ge, Yiwei Wang, Ming-Hsuan Yang, Yujun Cai

机构 * Department of Computer Science and Engineering, University of California at Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12661 2025-10-14 cs.LG cs.CL cs.CV 62%

VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization

Menglan Chen, Xianghe Pang, Jingjing Dong, WenHao Wang, Yaxin Du, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11456 2025-10-14 cs.CV 57%

Coupled Degradation Modeling and Fusion: A VLM-Guided Degradation-Coupled Network for Degradation-Aware Infrared and Visible Image Fusion

Tianpei Zhang, Jufeng Zhao, Yiming Zhu, Guangmang Cui

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11449 2025-10-14 cs.CV 57%

Enhancing Maritime Domain Awareness on Inland Waterways: A YOLO-Based Fusion of Satellite and AIS for Vessel Characterization

Geoffery Agorku, Sarah Hernandez, Hayley Hames, Cade Wagner

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09826 2025-10-14 cs.CV 57%

Isolated Channel Vision Transformers: From Single-Channel Pretraining to Multi-Channel Finetuning

Wenyi Lian, Patrick Micke, Joakim Lindblad, Nataša Sladoje

机构 * Department of Information Technology Uppsala University(信息科技系乌普萨拉大学) Department of Immunology, Genetics and Pathology Uppsala University(免疫学、遗传学和病理学系乌普萨拉大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Paper has been accepted by BMVC as an Oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17360 2025-10-14 cs.CV 57%

UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning

Maoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang, Shan Fu, Xue Yang, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) CTTL-Terminal, China Academy of Information and Communications Technology(信息通信技术中国科学院CTTL终端) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10731 2025-10-14 cs.RO cs.LG 50%

Controllable Generative Trajectory Prediction via Weak Preference Alignment

Yongxi Cao, Julian F. Schumann, Jens Kober, Joni Pajarinen, Arkady Zgonnikov

机构 * Aalto University(阿alto大学) TU Delft(代尔夫特理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18837 2025-10-14 cs.SI 50%

Sentiment and Social Signals in the Climate Crisis: A Survey on Analyzing Social Media Responses to Extreme Weather Events

Pouya Shaeri, Yasaman Mohammadpour, Alimohammad Beigi, Ariane Middel

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments Accepted and Published in SBP-BRiMS 2025. 18th International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation

详情

展开后加载摘要…

URL PDF HTML 收藏