arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2412.18962 2024-12-30 cs.IR cs.MM 79%

Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network

Zheyu Chen, Jinfeng Xu, Haibo Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02394 2024-12-30 eess.IV cs.CV 79%

Cohort-Individual Cooperative Learning for Multimodal Cancer Survival Analysis

Huajun Zhou, Fengtao Zhou, Hao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17297 2024-12-24 cs.CV 79%

Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural Perspective

Kaifang Long, Guoyang Xie, Lianbo Ma, Jiaqi Liu, Zhichao Lu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15652 2024-12-23 cs.CL 79%

Error-driven Data-efficient Large Multimodal Model Tuning

Barry Menglong Yao, Qifan Wang, Lifu Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14978 2024-12-20 cs.IR cs.MM 79%

Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation

Rongqing Kenneth Ong, Andy W. H. Khong

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.MM

Comments Accepted to ACM Web Search and Data Mining (WSDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01179 2024-12-20 cs.CV 79%

Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information

Yi Chen, Jian Xu, Xu-Yao Zhang, Wen-Zhuo Liu, Yang-Yang Liu, Cheng-Lin Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments AAAI2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12545 2024-12-19 cs.CL 79%

Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment

Tianyu Peng, Jiajun Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by COLING 2025, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13008 2024-12-18 cs.CL 79%

RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection

Tongguan Wang, Junkai Li, Guixin Su, Yongcheng Zhang, Dongyu Su, Yuxue Hu, Ying Sha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02703 2024-12-18 cs.CV 79%

Multi-modal Sensor Fusion for Auto Driving Perception: A Survey

Keli Huang, Botian Shi, Xiang Li, Xin Li, Siyuan Huang, Yikang Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11618 2024-12-17 cs.LG cs.AI 79%

EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

Nuowei Liu, Changzhi Sun, Tao Ji, Junfeng Tian, Jianxin Tang, Yuanbin Wu, Man Lan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10650 2024-12-17 cs.CV 79%

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

Yuhao Wang, Yang Liu, Aihua Zheng, Pingping Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by AAAI2025. More motifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10033 2024-12-16 cs.CV 79%

Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving

Zhihang Song, Lihui Peng, Jianming Hu, Danya Yao, Yi Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05025 2024-12-16 cs.AI 79%

Debiased Multimodal Understanding for Human Language Sequences

Zhi Xu, Dingkang Yang, Mingcheng Li, Yuzheng Wang, Zhaoyu Chen, Jiawei Chen, Jinjie Wei, Lihua Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09511 2024-12-13 cs.CV 79%

GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Dongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee Lee

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 22 pages, 8 figures, 12 tables; Project Page at https://dylanorange.github.io/projects/geal

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08979 2024-12-13 cs.LG cs.CV 79%

A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter

Zirun Guo, Xize Cheng, Yangyang Wu, Tao Jin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03681 2024-12-13 cs.CL 79%

Acquired TASTE: Multimodal Stance Detection with Textual and Structural Embeddings

Guy Barel, Oren Tsur, Dan Vilenchik

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17418 2024-12-12 cs.CV 79%

Multimodal Outer Arithmetic Block Dual Fusion of Whole Slide Images and Omics Data for Precision Oncology

Omnia Alwazzan, Amaya Gallagher-Syed, Thomas O. Millner, Sebastian Brandner, Ioannis Patras, Silvia Marino, Gregory Slabaugh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Revised to 10 pages, with corrected typos, updated references (some added, others removed), improved figure quality, modified text for better method validation, added one more co-author, and identified the IEEE member

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07475 2024-12-11 cs.CV 79%

Progressive Multi-Modal Fusion for Robust 3D Object Detection

Rohit Mohan, Daniele Cattaneo, Florian Drews, Abhinav Valada

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref 8th Annual Conference on Robot Learning, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06212 2024-12-10 cs.LG cs.AI 79%

A Self-guided Multimodal Approach to Enhancing Graph Representation Learning for Alzheimer's Diseases

Zhepeng Wang, Runxue Bao, Yawen Wu, Guodong Liu, Lei Yang, Liang Zhan, Feng Zheng, Weiwen Jiang, Yanfu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15896 2024-12-09 cs.CV 79%

Multimodal Instruction Tuning with Conditional Mixture of LoRA

Ying Shen, Zhiyang Xu, Qifan Wang, Yu Cheng, Wenpeng Yin, Lifu Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 7 figures, ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04337 2024-12-06 cs.CV 79%

Reflective Teacher: Semi-Supervised Multimodal 3D Object Detection in Bird's-Eye-View via Uncertainty Measure

Saheli Hazra, Sudip Das, Rohit Choudhary, Arindam Das, Ganesh Sistu, Ciaran Eising, Ujjwal Bhattacharya

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02531 2024-12-04 cs.CV 79%

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

Jinjin Cai, Kexin Meng, Baijian Yang, Gang Shao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14809 2024-12-04 cs.CL 79%

GSIFN: A Graph-Structured and Interlaced-Masked Multimodal Transformer-based Fusion Network for Multimodal Sentiment Analysis

Yijie Jin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Withdraw for the error in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01248 2024-12-03 cs.CV 79%

Multimodal Fusion Learning with Dual Attention for Medical Imaging

Joy Dhar, Nayyar Zaidi, Maryam Haghighat, Puneet Goyal, Sudipta Roy, Azadeh Alavi, Vikas Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17773 2024-12-03 cs.CV 79%

Efficient Multi-modal Large Language Models via Visual Token Grouping

Minbin Huang, Runhui Huang, Han Shi, Yimeng Chen, Chuanyang Zheng, Xiangguo Sun, Xin Jiang, Zhenguo Li, Hong Cheng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18669 2024-12-02 cs.CV 79%

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality

Chenyang Lei, Liyi Chen, Jun Cen, Xiao Chen, Zhen Lei, Felix Heide, Qifeng Chen, Zhaoxiang Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments project page: https://mt-cly.github.io/SimCMF.github.io/. arXiv admin note: substantial text overlap with arXiv:2409.08083

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17995 2024-11-28 cs.CV 79%

Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion

Taeheon Kim, Sangyun Chung, Youngjoon Yu, Yong Man Ro

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17837 2024-11-28 cs.CV 79%

OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion

Hanqi Jiang, Yi Pan, Junhao Chen, Zhengliang Liu, Yifan Zhou, Peng Shu, Yiwei Li, Huaqin Zhao, Stephen Mihm, Lewis C Howe, Tianming Liu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16577 2024-11-28 cs.LG cs.AI 79%

Seeking the Sufficiency and Necessity Causal Features in Multimodal Representation Learning

Boyu Chen, Junjie Liu, Zhu Li, Mengyue Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11909 2024-11-26 cs.CV 79%

SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization

Hongrui Jia, Chaoya Jiang, Haiyang Xu, Wei Ye, Mengfan Dong, Ming Yan, Ji Zhang, Fei Huang, Shikun Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏