arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2410.16853 2024-10-23 cs.CV cs.IR 83%

Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching

Xiang Ma, Xuemei Li, Lexin Fang, Caiming Zhang

专题命中 多模态训练与对齐 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15517 2024-10-22 cs.CL 83%

SceneGraMMi: Scene Graph-boosted Hybrid-fusion for Multi-Modal Misinformation Veracity Prediction

Swarang Joshi, Siddharth Mavani, Joel Alex, Arnav Negi, Rahul Mishra, Ponnurangam Kumaraguru

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15015 2024-10-22 cs.CV 83%

MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection

Yue Zhan, Zhihong Zeng, Haijun Liu, Xiaoheng Tan, Yinli Tian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14944 2024-10-22 cs.CV 83%

Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding

Yi Liu, Chengxin Li, Shoukun Xu, Jungong Han

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21659 2024-10-18 cs.CL 83%

Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models

Yue Xu, Xiuyuan Qi, Zhan Qin, Wenjie Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 12 pages, 9 figures, EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11358 2024-10-16 cs.CV 83%

SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection

Shuhan Dong, Yunsong Li, Weiying Xie, Jiaqing Zhang, Jiayuan Tian, Danian Yang, Jie Lei

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09572 2024-10-16 cs.CV 83%

Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ECCV2024 (Project Page: https://gyhdog99.github.io/projects/ecso/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16723 2024-09-27 cs.CV 83%

EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models

Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15936 2024-09-25 cs.CY cs.CV cs.HC 83%

DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection

Jiaxin Ye, Junping Zhang, Hongming Shan

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02392 2024-08-29 cs.CV 83%

TokenPacker: Efficient Visual Projector for Multimodal LLM

Wentong Li, Yuqian Yuan, Jian Liu, Dongqi Tang, Song Wang, Jie Qin, Jianke Zhu, Lei Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 16 pages, Codes:https://github.com/CircleRadon/TokenPacker

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12867 2024-08-26 cs.CV 83%

Semantic Alignment for Multimodal Large Language Models

Tao Wu, Mengze Li, Jingyuan Chen, Wei Ji, Wang Lin, Jinyang Gao, Kun Kuang, Zhou Zhao, Fei Wu

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10703 2024-08-21 cs.CV 83%

Large Language Models for Multimodal Deformable Image Registration

Mingrui Ma, Weijie Wang, Jie Ning, Jianfeng He, Nicu Sebe, Bruno Lepri

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10086 2024-08-20 cs.AI 83%

ARMADA: Attribute-Based Multimodal Data Augmentation

Xiaomeng Jin, Jeonghwan Kim, Yu Zhou, Kuan-Hao Huang, Te-Lin Wu, Nanyun Peng, Heng Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05019 2024-08-12 cs.CV 83%

Instruction Tuning-free Visual Token Complement for Multimodal LLMs

Dongsheng Wang, Jiequan Cui, Miaoge Li, Wang Lin, Bo Chen, Hanwang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ECCV2024 (20pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01800 2024-08-06 cs.CV 83%

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, Maosong Sun

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01003 2024-08-05 cs.AI 83%

Piculet: Specialized Models-Guided Hallucination Decrease for MultiModal Large Language Models

Kohou Wang, Xiang Liu, Zhaoxiang Liu, Kai Wang, Shiguo Lian

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19414 2024-07-30 cs.AI 83%

Appformer: A Novel Framework for Mobile App Usage Prediction Leveraging Progressive Multi-Modal Data Fusion and Feature Extraction

Chuike Sun, Junzhou Chen, Yue Zhao, Hao Han, Ruihai Jing, Guang Tan, Di Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18999 2024-07-30 cs.CV cs.LG 83%

Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models

Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin, Wenjun Zeng

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04727 2024-07-29 cs.LG cond-mat.soft cs.AI 83%

MMPolymer: A Multimodal Multitask Pretraining Framework for Polymer Property Prediction

Fanmeng Wang, Wentao Guo, Minjie Cheng, Shen Yuan, Hongteng Xu, Zhifeng Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted by the 33rd ACM International Conference on Information and Knowledge Management (CIKM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16714 2024-07-25 cs.LG cs.AI 83%

Masked Graph Learning with Recurrent Alignment for Multimodal Emotion Recognition in Conversation

Tao Meng, Fuchen Zhang, Yuntao Shou, Hongen Shao, Wei Ai, Keqin Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16168 2024-07-24 cs.CL 83%

Progressively Modality Freezing for Multi-Modal Entity Alignment

Yani Huang, Xuefeng Zhang, Richong Zhang, Junfan Chen, Jaein Kim

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 13pages, 8 figures, Accepted by ACL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15237 2024-07-23 cs.CL 83%

Two eyes, Two views, and finally, One summary! Towards Multi-modal Multi-tasking Knowledge-Infused Medical Dialogue Summarization

Anisha Saha, Abhisek Tiwari, Sai Ruthvik, Sriparna Saha

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02059 2024-07-23 cs.IR cs.CV 83%

IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT

Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Jie Wang, Joemon M. Jose

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted by SIGIR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11895 2024-07-17 cs.CV 83%

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Zehan Wang, Ziang Zhang, Hang Zhang, Luping Liu, Rongjie Huang, Xize Cheng, Hengshuang Zhao, Zhou Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Homepage is http://omnibind.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08786 2024-07-11 cs.CV 83%

Incorporating Clinical Guidelines through Adapting Multi-modal Large Language Model for Prostate Cancer PI-RADS Scoring

Tiantian Zhang, Manxi Lin, Hongda Guo, Xiaofan Zhang, Ka Fung Peter Chiu, Aasa Feragen, Qi Dou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06309 2024-07-10 cs.CY cs.AI 83%

Multimodal Chain-of-Thought Reasoning via ChatGPT to Protect Children from Age-Inappropriate Apps

Chuanbo Hu, Bin Liu, Minglei Yin, Yilu Zhou, Xin Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05540 2024-07-09 cs.CV 83%

GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation

Chenxin Li, Xinyu Liu, Cheng Wang, Yifan Liu, Weihao Yu, Jing Shao, Yixuan Yuan

专题命中 多模态训练与对齐 :omni-modal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00589 2024-07-04 cs.CV 83%

Merlin:Empowering Multimodal LLMs with Foresight Minds

En Yu, Liang Zhao, Yana Wei, Jinrong Yang, Dongming Wu, Lingyu Kong, Haoran Wei, Tiancai Wang, Zheng Ge, Xiangyu Zhang, Wenbing Tao

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ECCV2024. Project page: https://ahnsun.github.io/merlin

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17679 2024-06-26 cs.CV 83%

Local-to-Global Cross-Modal Attention-Aware Fusion for HSI-X Semantic Segmentation

Xuming Zhang, Naoto Yokoya, Xingfa Gu, Qingjiu Tian, Lorenzo Bruzzone

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13876 2024-06-19 cs.CV 83%

Multimodal Transformer Using Cross-Channel attention for Object Detection in Remote Sensing Images

Bissmella Bahaduri, Zuheng Ming, Fangchen Feng, Anissa Mokraou

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by ICIP2024

详情

展开后加载摘要…

URL PDF HTML 收藏