arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2209.13801 2022-09-29 cs.CV 79%

Translation, Scale and Rotation: Cross-Modal Alignment Meets RGB-Infrared Vehicle Detection

Maoxun Yuan, Yinyan Wang, Xingxing Wei

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09768 2022-09-21 cs.CL 79%

An Efficient End-to-End Transformer with Progressive Tri-modal Attention for Multi-modal Emotion Recognition

Yang Wu, Pai Peng, Zhenyu Zhang, Yanyan Zhao, Bing Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03765 2022-09-09 eess.SP cs.CV cs.HC cs.LG 79%

Self-Supervised Multimodal Fusion Transformer for Passive Activity Recognition

Armand K. Koupai, Mohammud J. Bocus, Raul Santos-Rodriguez, Robert J. Piechocki, Ryan McConville

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages, 7 figures, submitted to IET Wireless Sensor Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02368 2022-09-07 cs.CV 79%

Finger Multimodal Feature Fusion and Recognition Based on Channel Spatial Attention

Jian Guo, Jiaxiang Tu, Hengyi Ren, Chong Han, Lijuan Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01728 2022-09-07 cs.AI 79%

Features Fusion Framework for Multimodal Irregular Time-series Events

Peiwang Tang, Xianchao Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00979 2022-09-07 eess.IV cs.CV cs.LG 79%

Multimodal Information Fusion for Glaucoma and DR Classification

Yihao Li, Mostafa El Habib Daho, Pierre-Henri Conze, Hassan Al Hajj, Sophie Bonnin, Hugang Ren, Niranchana Manivannan, Stephanie Magazzeni, Ramin Tadayoni, Béatrice Cochener, Mathieu Lamard, Gwenolé Quellec

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted preprint for presentation at MICCAI-OMIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.06180 2022-08-19 cs.CV cs.RO 79%

Multi-modal Depression Estimation based on Sub-attentional Fusion

Ping-Cheng Wei, Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ECCV 2022 ACVR Workshop. Code is publicly available at https://github.com/PingCheng-Wei/DepressionEstimation

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05621 2022-08-12 cs.CV 79%

ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design

Xujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie, Zequn Jie, Chengwen Huang, Jianqing Peng, Xiaodan Liang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ACMMM22

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.03845 2022-07-27 cs.CV eess.IV 79%

Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks

Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 24 pages, 11 figures, submitted to IJRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02904 2022-07-27 cs.CV 79%

Multimodal Object Detection via Probabilistic Ensembling

Yi-Ting Chen, Jinghao Shi, Zelin Ye, Christoph Mertz, Deva Ramanan, Shu Kong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments camera-ready with supplement for ECCV2022 (oral presentation); open-source code at https://github.com/Jamie725/RGBT-detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08238 2022-07-22 cs.CV cs.LG 79%

AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification

Ammarah Farooq, Muhammad Awais, Josef Kittler, Syed Safwan Khalid

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments AAAI-2022 (Oral Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08721 2022-07-18 cs.CV 79%

Multimodal Token Fusion for Vision Transformers

Yikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang, Fuchun Sun, Yunhe Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09324 2022-07-07 cs.CL 79%

Supervised Visual Attention for Simultaneous Multimodal Machine Translation

Veneta Haralampieva, Ozan Caglayan, Lucia Specia

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Journal of Artificial Intelligence Research (JAIR)

Journal ref Journal of Artificial Intelligence Research 74 (2022) 1059-1089

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14699 2022-07-01 cs.CV eess.IV 79%

Fast computation of mutual information in the frequency domain with applications to global multimodal image alignment

Johan Öfverstedt, Joakim Lindblad, Nataša Sladoje

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures, 2 tables. The article is under consideration at Pattern Recognition Letters

Journal ref Pattern Recognition Letters, Vol. 159, pp. 196-203, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01681 2022-06-30 cs.CV 79%

SoloGAN: Multi-domain Multimodal Unpaired Image-to-Image Translation via a Single Generative Adversarial Network

Shihua Huang, Cheng He, Ran Cheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments pages 14, 15 figures

Journal ref IEEE Transactions on Artificial Intelligence 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.13256 2022-06-28 cs.MM 79%

A Topic-Attentive Transformer-based Model For Multimodal Depression Detection

Yanrong Guo, Chenyang Zhu, Shijie Hao, Richang Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12714 2022-06-28 cs.CV cs.CR cs.LG 79%

Defending Multimodal Fusion Models against Single-Source Adversaries

Karren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa, Zico Kolter

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12489 2022-06-22 eess.IV cs.CV 79%

Multi-modal and frequency-weighted tensor nuclear norm for hyperspectral image denoising

Xiaozhen Xie, Sheng Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This version modifies the list of authors, due to the changes of some authors' affiliations and the affiliations' requirements

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07981 2022-06-20 cs.CV 79%

Multi-scale Cooperative Multimodal Transformers for Multimodal Sentiment Analysis in Videos

Lianyang Ma, Yu Yao, Tao Liang, Tongliang Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05515 2022-06-15 cs.CL 79%

CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection

Zhen Li, Bing Xu, Conghui Zhu, Tiejun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Findings of NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04401 2022-06-10 cs.CV 79%

Cross-modal Local Shortest Path and Global Enhancement for Visible-Thermal Person Re-Identification

Xiaohong Wang, Chaoqi Li, Xiangcai Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12396 2022-05-26 cs.LG cs.CL 79%

Recipe2Vec: Multi-modal Recipe Representation Learning with Graph Neural Networks

Yijun Tian, Chuxu Zhang, Zhichun Guo, Yihong Ma, Ronald Metoyer, Nitesh V. Chawla

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05545 2022-05-12 eess.IV cs.CV cs.LG 79%

CNN-LSTM Based Multimodal MRI and Clinical Data Fusion for Predicting Functional Outcome in Stroke Patients

Nima Hatami, Tae-Hee Cho, Laura Mechtouff, Omer Faruk Eker, David Rousseau, Carole Frindel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 44th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04235 2022-05-10 q-bio.NC cs.AI cs.HC 79%

Measuring Cognitive Workload Using Multimodal Sensors

Niraj Hirachan, Anita Mathews, Julio Romero, Raul Fernandez Rojas

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01380 2022-05-04 cs.CV cs.LG eess.SP 79%

Deep Learning in Multimodal Remote Sensing Data Fusion: A Comprehensive Review

Jiaxin Li, Danfeng Hong, Lianru Gao, Jing Yao, Ke Zheng, Bing Zhang, Jocelyn Chanussot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13707 2022-05-02 cs.LG cs.AI 79%

Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing Modalities

Jiandian Zeng, Tianyi Liu, Jiantao Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by SIGIR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12541 2022-04-28 eess.IV cs.CV cs.LG 79%

Multi stain graph fusion for multimodal integration in pathology

Chaitanya Dwivedi, Shima Nofallah, Maryam Pouryahya, Janani Iyer, Kenneth Leidal, Chuhan Chung, Timothy Watkins, Andrew Billin, Robert Myers, John Abel, Ali Behrooz

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12185 2022-04-27 cs.CV 79%

TranSiam: Fusing Multimodal Visual Features Using Transformer for Medical Image Segmentation

Xuejian Li, Shiqiang Ma, Jijun Tang, Fei Guo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10760 2022-04-25 cs.CV 79%

iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition

Yixuan Wei, Yue Cao, Zheng Zhang, Zhuliang Yao, Zhenda Xie, Han Hu, Baining Guo

专题命中 多模态训练与对齐 :image-text(title,abstract);分类 cs.CV

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.06493 2022-04-22 cs.CV 79%

AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

Zehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang, Qinghong Jiang, Feng Zhao, Bolei Zhou, Hang Zhao

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.CV

Comments Accepted to IJCAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏