arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2311.10887 2025-07-29 cs.CV cs.AI 73%

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

Zhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li, Longlong Jing, Liang Yang, Bing Li

机构 * Clemson University(克莱姆森大学) Michigan State University(密歇根州立大学) Johns Hopkins University(约翰霍普金斯大学) The City University of New York(纽约城市大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11129 2025-07-18 cs.CV cs.AI cs.LG 73%

MMOne: Representing Multiple Modalities in One Scene

Zhifeng Gu, Bing Wang

机构 * Spatial Intelligence Group, The Hong Kong Polytechnic University(香港理工大学空间智能组)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11661 2025-07-17 cs.CL cs.AI 73%

Partitioner Guided Modal Learning Framework

Guimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu, Hasti Seifi

机构 * Guangdong University of Technology(广东工业大学) University of Copenhagen(哥本哈根大学) Nanjing University(南京大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tencent(腾讯) Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments acm multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04291 2025-06-23 cs.CL cs.CV 73%

Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models

Saketh Bachu, Erfan Shayegani, Rohit Lal, Trishna Chakraborty, Arindam Dutta, Chengyu Song, Yue Dong, Nael Abu-Ghazaleh, Amit K. Roy-Chowdhury

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICML 2025 as a spotlight poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16933 2025-06-05 cs.LG cs.CL cs.CV 73%

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Project page and codes: \url{https://ml-gsai.github.io/LLaDA-V-demo/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14562 2025-05-21 cs.SD cs.MM eess.AS 73%

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Parthasaarathy Sudarsanam, Irene Martín-Morató, Tuomas Virtanen

机构 * Audio Research Group, Tampere University(塔尔皮莱大学音频研究组)

专题命中 多模态训练与对齐 :multimodal(abstract);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Accepted to European Signal Processing Conference (EUSIPCO 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14246 2025-05-21 cs.CV cs.AI 73%

Visual Agentic Reinforcement Fine-Tuning

Ziyu Liu, Yuhang Zang, Yushan Zou, Zijian Liang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments project url: https://github.com/Liuziyu77/Visual-RFT/tree/main/Visual-ARFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16515 2025-05-21 cs.CV cs.AI cs.IR 73%

Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval

Delong Liu, Haiwen Li, Zhaohui Hou, Zhicheng Zhao, Fei Su, Yuan Dong

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) SenseTime(商汤科技) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Key Laboratory of Intereactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部交互技术与体验系统重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01950 2025-05-06 cs.CV cs.AI 73%

Segment Any RGB-Thermal Model with Language-aided Distillation

Dong Xing, Xianxun Zhu, Wei Zhou, Qika Lin, Hang Yang, Yuqing Wang

机构 * Changchun Institute of Optics, Fine Mechanicsand Physics, Chinese Academy of Sciences University(长春光学精密机械与物理研究所,中国科学院大学) School of Computing, Macquarie University(麦考瑞大学计算机学院) School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学 Saw Swee Hock 公共卫生学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: text overlap with arXiv:2412.04220 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01255 2025-05-05 cs.CL cs.IR cs.MM 73%

PREMISE: Matching-based Prediction for Accurate Review Recommendation

Wei Han, Hui Chen, Soujanya Poria

机构 * Singapore University of Technology and Design(新加坡科技设计大学) National University of Singapore(国立新加坡大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.MM

Comments 19 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04821 2025-04-30 eess.IV cs.AI cs.CV 73%

RGB-Thermal Infrared Fusion for Robust Depth Estimation in Complex Environments

Zelin Meng, Takanori Fukao

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09445 2025-04-02 cs.CV cs.AI 73%

Astrea: A MOE-based Visual Understanding Model with Progressive Alignment

Xiaoda Yang, JunYu Lu, Hongshun Qiu, Sijing Li, Hao Li, Shengpeng Ji, Xudong Tang, Jiayang Xu, Jiaqi Duan, Ziyue Jiang, Cong Lin, Sihang Cai, Zejian Xie, Zhuoyang Song, Songxin Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06405 2025-04-02 cs.SD cs.AI eess.AS 73%

Heterogeneous bimodal attention fusion for speech emotion recognition

Jiachen Luo, Huy Phan, Lin Wang, Joshua Reiss

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21307 2025-03-28 cs.CV cs.AI 73%

InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Dongchen Lu, Yuyao Sun, Zilu Zhang, Leping Huang, Jianliang Zeng, Mao Shu, Huo Cao

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10603 2025-03-26 cs.CV cs.AI 73%

Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition

Jun Yu, Lingsi Zhu, Yanjun Chi, Yunxiang Zhang, Yang Zheng, Yongqi Wang, Xilong Lu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00804 2025-03-04 cs.CV cs.AI 73%

DELST: Dual Entailment Learning for Hyperbolic Image-Gene Pretraining in Spatial Transcriptomics

Xulin Chen, Junzhou Huang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12431 2025-02-27 cs.LG cs.AI cs.CL 73%

Modality Interactive Mixture-of-Experts for Fake News Detection

Yifan Liu, Yaokun Liu, Zelin Li, Ruichen Yao, Yang Zhang, Dong Wang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by the Proceedings of the ACM Web Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19201 2025-02-03 cs.CL cs.AI cs.LG 73%

Efficient Reasoning with Hidden Thinking

Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, Jiuxiang Gu

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06546 2025-01-14 cs.CV cs.AI 73%

Natural Language Supervision for Low-light Image Enhancement

Jiahui Tang, Kaihua Zhou, Zhijian Luo, Yueen Hou

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17981 2024-12-20 cs.CV cs.AI 73%

From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information

Qirui Jiao, Daoyuan Chen, Yilun Huang, Yaliang Li, Ying Shen

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 32 pages, 22 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08771 2024-12-13 cs.CV cs.AI cs.LG 73%

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information

Ke Wang, Hong Xuan

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03854 2024-11-25 cs.LG cs.CL cs.CV 73%

Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity

Zitao Shuai, Chenwei Wu, Zhengxu Tang, Liyue Shen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13193 2024-06-21 cs.LG cs.AI cs.CL physics.chem-ph 73%

PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes

He Cao, Yanjun Shao, Zhiyuan Liu, Zijing Liu, Xiangru Tang, Yuan Yao, Yu Li

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05163 2024-06-21 cs.CV cs.AI 73%

MISS: A Generative Pretraining and Finetuning Approach for Med-VQA

Jiawei Chen, Dingkang Yang, Yue Jiang, Yuxuan Lei, Lihua Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments ICANN, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01460 2024-06-05 cs.CV cs.AI 73%

MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization

Yu Zhang, Qi Zhang, Zixuan Gong, Yiwei Shi, Yepeng Liu, Duoqian Miao, Yang Liu, Ke Liu, Kun Yi, Wei Fan, Liang Hu, Changwei Wang

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13530 2024-04-23 cs.CV cs.CL cs.LG 73%

Listen Then See: Video Alignment with Speaker Attention

Aviral Agrawal, Carlos Mateo Samudio Lezcano, Iqui Balam Heredia-Marin, Prabhdeep Singh Sethi

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13039 2024-04-22 cs.CV cs.CL 73%

LaPA: Latent Prompt Assist Model For Medical Visual Question Answering

Tiancheng Gu, Kaicheng Yang, Dongnan Liu, Weidong Cai

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments 10 pages, 4 figures, Accepted by CVPRW2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05225 2024-04-09 cs.CV cs.CL 73%

LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding

Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng, Zhi Yu, Cong Yao

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16530 2024-03-26 cs.CV cs.AI 73%

An Intermediate Fusion ViT Enables Efficient Text-Image Alignment in Diffusion Models

Zizhao Hu, Shaochong Jia, Mohammad Rostami

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18340 2024-03-26 cs.CL cs.AI 73%

UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web

Yibo Yan, Haomin Wen, Siru Zhong, Wei Chen, Haodong Chen, Qingsong Wen, Roger Zimmermann, Yuxuan Liang

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by The Web Conference 2024

详情

展开后加载摘要…

URL PDF HTML 收藏