arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2504.09448 2025-04-23 cs.CV cs.LG 83%

Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization

Lin Zhu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13645 2025-04-21 cs.CV cs.LG 83%

Efficient Parameter Adaptation for Multi-Modal Medical Image Segmentation and Prognosis

Numan Saeed, Shahad Hardan, Muhammad Ridzuan, Nada Saadi, Karthik Nandakumar, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence(机器学习系,Mohamed bin Zayed人工智能大学) Michigan State University(密歇根州立大学) M42 Health(M42健康)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10538 2025-04-16 cs.IR cs.AI 83%

Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation

Jiajie Su, Qiyong Zhong, Yunshan Ma, Weiming Liu, Chaochao Chen, Xiaolin Zheng, Jianwei Yin, Tat-Seng Chua

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02430 2025-04-11 cs.CV 83%

FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance

Haicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju, Victor Quétu, Shuai Xiao, Enzo Tartaglione

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15011 2025-04-08 cs.CV 83%

CrossOver: 3D Scene Cross-Modal Alignment

Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Project Page: https://sayands.github.io/crossover/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23721 2025-04-01 cs.LG cs.AI 83%

Unimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion

Jiagen Li, Rui Yu, Huihao Huang, Huaicheng Yan

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23456 2025-04-01 cs.CV 83%

CADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation

Maofu Liu, Xin Jiang, Xiaokang Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14455 2025-03-31 cs.CV 83%

MM-GTUNets: Unified Multi-Modal Graph Deep Learning for Brain Disorders Prediction

Luhui Cai, Weiming Zeng, Hongyu Chen, Hua Zhang, Yueyang Li, Yu Feng, Hongjie Yan, Lingbin Bian, Wai Ting Siok, Nizhuan Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18100 2025-03-25 cs.CV 83%

M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving

Xuesong Chen, Shaoshuai Shi, Tao Ma, Jingqiu Zhou, Simon See, Ka Chun Cheung, Hongsheng Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15284 2025-03-20 cs.CV 83%

EdgeRegNet: Edge Feature-based Multimodal Registration Network between Images and LiDAR Point Clouds

Yuanchao Yue, Hui Yuan, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08992 2025-03-18 cs.CV 83%

Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection

Xuzhong Hu, Zaipeng Duan, Pei An, Jun zhang, Jie Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12560 2025-03-18 cs.CL 83%

Multi-Granular Multimodal Clue Fusion for Meme Understanding

Li Zheng, Hao Fei, Ting Dai, Zuquan Peng, Fei Li, Huisheng Ma, Chong Teng, Donghong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01217 2025-03-18 cs.CV 83%

CromSS: Cross-modal pre-training with noisy labels for remote sensing image segmentation

Chenying Liu, Conrad Albrecht, Yi Wang, Xiao Xiang Zhu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The 1st short version was accepted as an oral presentation by ICLR 2024 ML4RS workshop. The 2nd extended version was accepted by IEEE TGRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11183 2025-03-17 cs.CV 83%

Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation

Leideng Shi, Juan Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 5 figures, accepted in IEEE Geoscience and Remote Sensing Letters (GRSL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10501 2025-03-14 cs.CV 83%

TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Xudong Tan, Peng Ye, Chongjun Tu, Jianjian Cao, Yaoxin Yang, Lin Zhang, Dongzhan Zhou, Tao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09873 2025-03-14 cs.CV 83%

FDCT: Frequency-Aware Decomposition and Cross-Modal Token-Alignment for Multi-Sensor Target Classification

Shoaib Meraj Sami, Md Mahedi Hasan, Nasser M. Nasrabadi, Raghuveer Rao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 12 pages Accepted in the IEEE Transactions on Aerospace and Electronic Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07593 2025-03-11 cs.CV 83%

Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Youjun Zhao, Jiaying Lin, Rynson W. H. Lau

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments AAAI 2025 (Extented Version). Project Page: https://youjunzhao.github.io/HCMA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13652 2025-03-11 cs.CL 83%

LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models

Yizheng Sun, Yanze Xin, Hao Li, Jingyuan Sun, Chenghua Lin, Riza Batista-Navarro

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06063 2025-03-11 cs.CV 83%

Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

Junyan Lin, Haoran Chen, Yue Fan, Yingqi Fan, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03202 2025-03-06 cs.CV 83%

Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings

Sneh Pillai

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00591 2025-03-04 cs.CV 83%

AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models

Sohan Patnaik, Rishabh Jain, Balaji Krishnamurthy, Mausoom Sarkar

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted for publication in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12164 2025-03-04 cs.LG cs.AI 83%

GAMED: Knowledge Adaptive Multi-Experts Decoupling for Multimodal Fake News Detection

Lingzhi Shen, Yunfei Long, Xiaohao Cai, Imran Razzak, Guanming Chen, Kang Liu, Shoaib Jameel

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16786 2025-03-03 cs.CV 83%

SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding

Liangtao Shi, Ting Liu, Xiantao Hu, Yue Hu, Quanjun Yin, Richang Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20008 2025-02-28 cs.CV 83%

Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up

Lang Huang, Qiyu Wu, Zhongtao Miao, Toshihiko Yamasaki

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18481 2025-02-27 cs.IR cs.AI 83%

MDE: Modality Discrimination Enhancement for Multi-modal Recommendation

Hang Zhou, Yucheng Wang, Huijing Zhan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17425 2025-02-25 cs.CV cs.LG 83%

Introducing Visual Perception Token into Multimodal Large Language Model

Runpeng Yu, Xinyin Ma, Xinchao Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13624 2025-02-20 cs.CV 83%

CardiacMamba: A Multimodal RGB-RF Fusion Framework with State Space Models for Remote Physiological Measurement

Zheng Wu, Yiping Xie, Bo Zhao, Jiguang He, Fei Luo, Ning Deng, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10419 2025-02-19 cs.NE cs.AI cs.LG 83%

A Hybrid Swarm Intelligence Approach for Optimizing Multimodal Large Language Models Deployment in Edge-Cloud-based Federated Learning Environments

Gaith Rjouba, Hanae Elmekki, Saidul Islam, Jamal Bentahar, Rachida Dssouli

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10167 2025-02-18 cs.CV eess.SP 83%

X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing

Xinyan Chen, Jianfei Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01943 2025-02-12 cs.CV 83%

DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

Jinda Lu, Junkang Wu, Jinghan Li, Xiaojun Jia, Shuo Wang, YiFan Zhang, Junfeng Fang, Xiang Wang, Xiangnan He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏