arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2207.07027 2023-03-03 eess.IV cs.CV cs.LG 83%

MedFuse: Multi-modal fusion with clinical time-series data and chest X-ray images

Nasir Hayat, Krzysztof J. Geras, Farah E. Shamout

专题命中 多模态训练与对齐 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14264 2023-03-01 cs.RO cs.CV 83%

RGB-D Grasp Detection via Depth Guided Learning with Cross-modal Attention

Ran Qin, Haoxiang Ma, Boyang Gao, Di Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted at ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08052 2023-02-17 cs.CV 83%

Hierarchical Cross-modal Transformer for RGB-D Salient Object Detection

Hao Chen, Feihong Shen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02363 2023-02-17 cs.CV 83%

CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection

Youwei Pang, Xiaoqi Zhao, Lihe Zhang, Huchuan Lu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by TIP-2023. Add more details and update the weight illustration

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09174 2023-01-24 cs.CV cs.HC cs.LG 83%

MATT: Multimodal Attention Level Estimation for e-learning Platforms

Roberto Daza, Luis F. Gomez, Aythami Morales, Julian Fierrez, Ruben Tolosana, Ruth Cobos, Javier Ortega-Garcia

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Preprint of the paper presented to the Workshop on Artificial Intelligence for Education (AI4EDU) of AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04856 2023-01-13 cs.CL cs.LG stat.ML 83%

Multimodal Deep Learning

Cem Akkus, Luyang Chu, Vladana Djakovic, Steffen Jauch-Walser, Philipp Koch, Giacomo Loss, Christopher Marquardt, Marco Moldovan, Nadja Sauter, Maximilian Schneider, Rickmer Schulte, Karol Urbanczyk, Jann Goschenhofer, Christian Heumann, Rasmus Hvingelby, Daniel Schalk, Matthias Aßenmacher

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02445 2023-01-12 cs.AI cs.LG 83%

IMKGA-SM: Interpretable Multimodal Knowledge Graph Answer Prediction via Sequence Modeling

Yilin Wen, Biao Luo, Yuqian Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 12pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00135 2022-12-02 cs.CV 83%

Attention Bottlenecks for Multimodal Fusion

Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, Chen Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments Published at NeurIPS 2021. Note this version updates numbers due to a bug in the AudioSet mAP calculation in Table 1 (last row)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03521 2022-12-01 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

Good Visual Guidance Makes A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction

Xiang Chen, Ningyu Zhang, Lei Li, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, Huajun Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03126 2022-11-23 cs.MM cs.AI cs.CL cs.CV cs.LG 83%

DM$^2$S$^2$: Deep Multi-Modal Sequence Sets with Hierarchical Modality Attention

Shunsuke Kitada, Yuki Iwazaki, Riku Togashi, Hitoshi Iyatomi

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 3 figures. Accepted by IEEE Access on Nov. 3, 2022

Journal ref in IEEE Access, vol. 10, pp. 120023-120034, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.00999 2022-11-11 cs.CV 83%

Order embeddings and character-level convolutions for multimodal alignment

Jônatas Wehrmann, Anderson Mattjie, Rodrigo C. Barros

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 7 pages, 5 figures, submitted to Pattern Recognition Letters

Journal ref Pattern Recognition Letters, vol. 102, 15, January 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04331 2022-11-09 cs.CV 83%

Multi-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition

Hyeongju Choi, Apoorva Beedu, Harish Haresamudram, Irfan Essa

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06778 2022-11-01 cs.CV 83%

X-Align: Cross-Modal Cross-View Alignment for Bird's-Eye-View Segmentation

Shubhankar Borse, Marvin Klingner, Varun Ravi Kumar, Hong Cai, Abdulaziz Almuzairee, Senthil Yogamani, Fatih Porikli

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09615 2022-10-19 cs.CV 83%

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

Xin Li, Botian Shi, Yuenan Hou, Xingjiao Wu, Tianlong Ma, Yikang Li, Liang He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13304 2022-10-17 eess.IV cs.CV 83%

DXM-TransFuse U-net: Dual Cross-Modal Transformer Fusion U-net for Automated Nerve Identification

Baijun Xie, Gary Milam, Bo Ning, Jaepyeong Cha, Chung Hyuk Park

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Journal ref Computerized Medical Imaging and Graphics, 2022-07-01, Volume 99, Article 102090

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06716 2022-10-14 cs.CL 83%

Low-resource Neural Machine Translation with Cross-modal Alignment

Zhe Yang, Qingkai Fang, Yang Feng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CL

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10771 2022-08-24 cs.CV 83%

Learning an Efficient Multimodal Depth Completion Model

Dewang Hou, Yuanyuan Du, Kai Zhao, Yang Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments To appear in ECCV 2022 workshop and codes are available at https://github.com/dwHou/EMDC-PyTorch

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03305 2022-07-12 cs.AI 83%

Multimodal E-Commerce Product Classification Using Hierarchical Fusion

Tsegaye Misikir Tashu, Sara Fattouh, Peter Kiss, Tomas Horvath

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11826 2022-06-27 cs.CV 83%

Toward Clinically Assisted Colorectal Polyp Recognition via Structured Cross-modal Representation Consistency

Weijie Ma, Ye Zhu, Ruimao Zhang, Jie Yang, Yiwen Hu, Zhen Li, Li Xiang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Early Accepted by MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08256 2022-05-11 cs.MM 83%

A Deep Multi-Level Attentive network for Multimodal Sentiment Analysis

Ashima Yadav, Dinesh Kumar Vishwakarma

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.MM

Comments 11 pages, 7 figures

Journal ref ACM Transactions on Multimedia Computing, Communications, and Applications, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01818 2022-05-06 cs.LG cs.AI cs.CL cs.CV eess.AS 83%

i-Code: An Integrative and Composable Multimodal Learning Framework

Ziyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant, Dongdong Chen, Yu Shi, Yichong Xu, Yao Qian, Mei Gao, Yi-Ling Chen, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao, Lu Yuan, Takuya Yoshioka, Michael Zeng, Xuedong Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09848 2022-04-22 cs.CV 83%

Weakly Aligned Feature Fusion for Multimodal Object Detection

Lu Zhang, Zhiyong Liu, Xiangyu Zhu, Zhan Song, Xu Yang, Zhen Lei, Hong Qiao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Journal extension of the previous conference paper arXiv:1901.02645, see IEEE page https://ieeexplore.ieee.org/document/9523596

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09994 2022-03-21 cs.CL cs.LG 83%

Graph-Text Multi-Modal Pre-training for Medical Representation Learning

Sungjin Park, Seongsu Bae, Jiho Kim, Tackeun Kim, Edward Choi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments To appear in Proceedings of the Conference on Health, Inference, and Learning (CHIL 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03850 2022-03-09 cs.CL cs.PL cs.SE 83%

UniXcoder: Unified Cross-Modal Pre-training for Code Representation

Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, Jian Yin

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments Published in ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10454 2022-02-23 cs.LG cs.AI 83%

A Novel Anomaly Detection Method for Multimodal WSN Data Flow via a Dynamic Graph Neural Network

Qinghao Zhang, Miao Ye, Hongbing Qiu, Yong Wang, Xiaofang Deng

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09828 2022-01-25 cs.LG cs.CV 83%

MMLatch: Bottom-up Top-down Fusion for Multimodal Sentiment Analysis

Georgios Paraskevopoulos, Efthymios Georgiou, Alexandros Potamianos

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted, ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11405 2021-12-14 cs.CV cs.CL cs.LG cs.MM cs.SD eess.AS 83%

M2P2: Multimodal Persuasion Prediction using Adaptive Fusion

Chongyang Bai, Haipeng Chen, Srijan Kumar, Jure Leskovec, V. S. Subrahmanian

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments published in IEEE Trans. on Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.07502 2021-11-11 cs.LG cs.AI cs.CL cs.CV cs.MM 83%

MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Chen, Peter Wu, Michelle A. Lee, Yuke Zhu, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2021 Datasets and Benchmarks Track. Code: https://github.com/pliang279/MultiBench and Website: https://cmu-multicomp-lab.github.io/multibench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12673 2021-10-18 cs.CV 83%

Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal Data

Xuhui Jia, Kai Han, Yukun Zhu, Bradley Green

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12946 2021-09-28 cs.CV 83%

Fusion-GCN: Multimodal Action Recognition using Graph Convolutional Networks

Michael Duhme, Raphael Memmesheimer, Dietrich Paulus

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 18 pages, 6 figures, 3 tables, GCPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏