arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2111.12698 2022-04-20 cs.CV 79%

Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, Ehsan Elhamifar

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12367 2022-04-19 cs.CV 79%

Transformer-based Multimodal Information Fusion for Facial Expression Analysis

Wei Zhang, Feng Qiu, Suzhen Wang, Hao Zeng, Zhimeng Zhang, Rudong An, Bowen Ma, Yu Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03249 2022-04-13 cs.CV 79%

Multimodal Colored Point Cloud to Image Alignment

Noam Rotstein, Amit Bracha, Ron Kimmel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05603 2022-04-05 eess.IV cs.CV 79%

Multi-Modal MRI Reconstruction Assisted with Spatial Alignment Network

Kai Xuan, Lei Xiang, Xiaoqian Huang, Lichi Zhang, Shu Liao, Dinggang Shen, Qian Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Final version, IEEE Transactions on Medical Imaging, code available at \url{https://github.com/woxuankai/SpatialAlignmentNetwork}

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09707 2022-03-29 cs.SE cs.AI 79%

M2TS: Multi-Scale Multi-Modal Approach Based on Transformer for Source Code Summarization

Yuexiu Gao, Chen Lyu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICPC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13411 2022-03-28 cs.RO cs.AI cs.LG cs.SY eess.SY 79%

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, Rogerio Bonatti

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11441 2022-03-23 cs.CV 79%

Multi-Modal Learning for AU Detection Based on Multi-Head Fused Transformers

Xiang Zhang, Lijun Yin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref FG 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08195 2022-03-17 cs.CV 79%

DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection

Yingwei Li, Adams Wei Yu, Tianjian Meng, Ben Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Bo Wu, Yifeng Lu, Denny Zhou, Quoc V. Le, Alan Yuille, Mingxing Tan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2022. 1st rank 3D detection method on Waymo Challenge Leaderboard: https://waymo.com/open/challenges/entry/?timestamp=1647356360224524&challenge=DETECTION_3D&emailId=5451f123-a0ea

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08055 2022-03-16 cs.CL 79%

Modular and Parameter-Efficient Multimodal Fusion with Prompting

Sheng Liang, Mengjie Zhao, Hinrich Schütze

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Findings of ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04586 2022-03-10 eess.IV cs.CV 79%

Multi-modal Brain Tumor Segmentation via Missing Modality Synthesis and Modality-level Attention Fusion

Ziqi Huang, Li Lin, Pujin Cheng, Linkai Peng, Xiaoying Tang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures, submitted to ICPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00510 2022-03-03 eess.SP cs.CV cs.HC cs.NI 79%

Multi-Modal Recurrent Fusion for Indoor Localization

Jianyuan Yu, Pu, Wang, Toshiaki Koike-Akino, Philip V. Orlik

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 5 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08080 2022-02-22 eess.IV cs.CV 79%

Deep multi-modal aggregation network for MR image reconstruction with auxiliary modality

Chun-Mei Feng, Huazhu Fu, Tianfei Zhou, Yong Xu, Ling Shao, David Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11641 2022-02-15 cs.CV eess.IV 79%

Cross-Modal Distillation for RGB-Depth Person Re-Identification

Frank Hafner, Amran Bhuiyan, Julian F. P. Kooij, Eric Granger

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Journal ref Computer Vision and Image Understanding, 103352 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04327 2022-02-10 cs.CV cs.IR 79%

Anchor Graph Structure Fusion Hashing for Cross-Modal Similarity Search

Lu Wang, Jie Yang, Masoumeh Zareapoor, Zhonglong Zheng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10274 2022-01-26 cs.CL 79%

Multi-channel Attentive Graph Convolutional Network With Sentiment Fusion For Multimodal Sentiment Analysis

Luwei Xiao, Xingjiao Wu, Wen Wu, Jing Yang, Liang He

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07520 2022-01-20 cs.CL 79%

CM3: A Causal Masked Multimodal Model of the Internet

Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05996 2022-01-19 cs.CR cs.CV 79%

Hardware Implementation of Multimodal Biometric using Fingerprint and Iris

Tariq M Khan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13782 2022-01-19 cs.LG cs.AI 79%

Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions

Anil Rahate, Rahee Walambe, Sheela Ramanna, Ketan Kotecha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments This is published in Information Fusion Journal and published copy can be downloaded from https://authors.elsevier.com/c/1eIVS5a7-Gls0Y for 50 days. https://www.sciencedirect.com/science/article/pii/S1566253521002530

Journal ref Information Fusion 81(2022) 203-239

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14740 2022-01-11 cs.CL 79%

LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments ACL 2021 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07800 2022-01-04 cs.CV 79%

Multi-modal Visual Place Recognition in Dynamics-Invariant Perception Space

Lin Wu, Teng Wang, Changyin Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Signal Processing Letters 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09379 2022-01-04 cs.CV cs.LG 79%

BM-NAS: Bilevel Multimodal Neural Architecture Search

Yihang Yin, Siyu Huang, Xiang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12792 2021-12-30 cs.LG cs.MM 79%

Understanding and Measuring Robustness of Multimodal Learning

Nishant Vishwamitra, Hongxin Hu, Ziming Zhao, Long Cheng, Feng Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12927 2021-12-28 cs.CV 79%

Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification

Zhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han, Jingyan Qin, Xu-Cheng Yin

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04537 2021-12-16 eess.IV cs.CV 79%

Multimodal Representation Learning via Maximization of Local Mutual Information

Ruizhi Liao, Daniel Moyer, Miriam Cha, Keegan Quigley, Seth Berkowitz, Steven Horng, Polina Golland, William M. Wells

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

Comments In Proceedings of International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2021

Journal ref In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 273-283. Springer, Cham, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01530 2021-12-13 cs.CV 79%

High-resolution Depth Maps Imaging via Attention-based Hierarchical Multi-modal Fusion

Zhiwei Zhong, Xianming Liu, Junjun Jiang, Debin Zhao, Zhiwen Chen, Xiangyang Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00775 2021-12-03 cs.CV 79%

Routing with Self-Attention for Multimodal Capsule Networks

Kevin Duarte, Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Samuel Thomas, Alexander Liu, David Harwath, James Glass, Hilde Kuehne, Mubarak Shah

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13361 2021-11-29 cs.LG cs.AI 79%

Geometric Multimodal Deep Learning with Multi-Scaled Graph Wavelet Convolutional Network

Maysam Behmanesh, Peyman Adibi, Mohammad Saeed Ehsani, Jocelyn Chanussot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11992 2021-11-29 cs.CV cs.LG 79%

Sparse Fusion for Multimodal Transformers

Yi Ding, Alex Rich, Mason Wang, Noah Stier, Matthew Turk, Pradeep Sen, Tobias Höllerer

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 4 figures, 5 tables, Yi Ding and Alex Rich contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.09624 2021-11-19 cs.CV 79%

IMFNet: Interpretable Multimodal Fusion for Point Cloud Registration

Xiaoshui Huang, Wentao Qu, Yifan Zuo, Yuming Fang, Xiaowei Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08703 2021-11-18 cs.CV cs.CR 79%

Benchmarking Quality-Dependent and Cost-Sensitive Score-Level Multimodal Biometric Fusion Algorithms

Norman Poh, Thirimachos Bourlai, Josef Kittler, Lorene Allano, Fernando Alonso-Fernandez, Onkar Ambekar, John Baker, Bernadette Dorizzi, Omolara Fatukasi, Julian Fierrez, Harald Ganster, Javier Ortega-Garcia, Donald Maurer, Albert Ali Salah, Tobias Scheidat, Claus Vielhauer

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Published at IEEE Transactions on Information Forensics and Security journal

详情

展开后加载摘要…

URL PDF HTML 收藏