arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2012.04977 2020-12-10 cs.CV 57%

Hateful Memes Detection via Complementary Visual and Linguistic Networks

Weibo Zhang, Guihua Liu, Zhuohua Li, Fuqing Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06985 2020-11-16 cs.RO cs.AI 57%

Robotic self-representation improves manipulation skills and transfer learning

Phuong D. H. Nguyen, Manfred Eppe, Stefan Wermter

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Submitted to IEEE Robotics and Automation Letters (RA-L) 2021 with International Conference on Robotics and Automation Conference Option (ICRA) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.15647 2020-11-03 eess.IV cs.CV 57%

Brain Tumor Segmentation Network Using Attention-based Fusion and Spatial Relationship Constraint

Chenyu Liu, Wangbin Ding, Lei Li, Zhen Zhang, Chenhao Pei, Liqin Huang, Xiahai Zhuang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08937 2020-09-04 cs.CV q-bio.GN q-bio.TO 57%

Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and Prognosis

Richard J. Chen, Ming Y. Lu, Jingwen Wang, Drew F. K. Williamson, Scott J. Rodig, Neal I. Lindeman, Faisal Mahmood

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Code and trained models are made available at: https://github.com/mahmoodlab/PathomicFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00784 2020-09-03 cs.CV 57%

CLOCs: Camera-LiDAR Object Candidates Fusion for 3D Object Detection

Su Pang, Daniel Morris, Hayder Radha

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05892 2020-09-01 cs.CV cs.LG eess.IV 57%

LGNN: A Context-aware Line Segment Detector

Quan Meng, Jiakai Zhang, Qiang Hu, Xuming He, Jingyi Yu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05658 2020-08-14 cs.CL 57%

Cognitive Representation Learning of Self-Media Online Article Quality

Yiru Wang, Shen Huang, Gongfu Li, Qiang Deng, Dongliang Liao, Pengda Si, Yujiu Yang, Jin Xu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

Comments Accepted at the Proceedings of the 28th ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.00450 2020-07-29 cs.LG cs.CV stat.ML 57%

Multi-Object Representation Learning with Iterative Variational Inference

Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Chris Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, Alexander Lerchner

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Journal ref ICML 2019 (PMLR 97:2424-2433)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.13732 2020-07-28 cs.CV 57%

Learning Lane Graph Representations for Motion Forecasting

Ming Liang, Bin Yang, Rui Hu, Yun Chen, Renjie Liao, Song Feng, Raquel Urtasun

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ECCV 2020 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.13538 2020-07-28 cs.CV eess.IV 57%

A Novel adaptive optimization of Dual-Tree Complex Wavelet Transform for Medical Image Fusion

T. Deepika, G. Karpaga Kannan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Conference on Computing Communication and Signal Processing. arXiv admin note: text overlap with arXiv:2007.11488

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.13135 2020-07-28 cs.CV eess.IV 57%

Contrastive Visual-Linguistic Pretraining

Lei Shi, Kai Shuang, Shijie Geng, Peng Su, Zhengkai Jiang, Peng Gao, Zuohui Fu, Gerard de Melo, Sen Su

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.12858 2020-07-28 cs.LG cs.CV stat.ML 57%

Modal Uncertainty Estimation via Discrete Latent Representation

Di Qiu, Lok Ming Lui

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.12072 2020-07-28 cs.CV cs.LG eess.IV 57%

TSIT: A Simple and Versatile Framework for Image-to-Image Translation

Liming Jiang, Changxu Zhang, Mingyang Huang, Chunxiao Liu, Jianping Shi, Chen Change Loy

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ECCV 2020 (Spotlight). Table 2 is updated. GitHub: https://github.com/EndlessSora/TSIT

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.09183 2020-07-21 cs.CV 57%

Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation

Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang, Wayne Wu, Chen Qian, Hongsheng Li, Gang Zeng

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.08763 2020-07-20 cs.CV 57%

AE-Net: Autonomous Evolution Image Fusion Method Inspired by Human Cognitive Mechanism

Aiqing Fang, Xinbo Zhao, Jiaqi Yang, Shihao Cao, Yanning Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.10664 2020-07-17 eess.IV cs.CV cs.LG stat.ML 57%

A review: Deep learning for medical image segmentation using multi-modality fusion

Tongxue Zhou, Su Ruan, Stéphane Canu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 26 pages, 8 figures

Journal ref Array, Volumes 3-4, September-December 2019, Article 100004

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.06793 2020-07-15 cs.CV cs.LG 57%

TCGM: An Information-Theoretic Framework for Semi-Supervised Multi-Modality Learning

Xinwei Sun, Yilun Xu, Peng Cao, Yuqing Kong, Lingjing Hu, Shanghang Zhang, Yizhou Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ECCV 2020 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02041 2020-07-07 cs.CV 57%

Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking

Pengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu, Xiaoyun Yang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10570 2020-06-30 cs.CV cs.RO eess.IV 57%

Real-time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-driving Images

Lei Sun, Kailun Yang, Xinxin Hu, Weijian Hu, Kaiwei Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Robotics and Automation Letters (RA-L); 8 Figures, 3 Tables; Code is available at https://github.com/AHupuJR/RFNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12201 2020-06-23 physics.app-ph cs.CV 57%

Thermal vulnerability detection in integrated electronic and photonic circuits using IR thermography

Bilal Hussain, Bushra Jalil, Maria Antonietta Pascali, Muhammad Imran, Giovanni Serafino, Davide Moroni, Paolo Ghelfi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 2020 Optical Society of America. One print or electronic copy may be made for personal use only. Systematic reproduction and distribution, duplication of any material in this paper for a fee or for commercial purposes, or modifications of the content of this paper are prohibited

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11481 2020-06-23 cs.CV 57%

Pseudo-LiDAR Point Cloud Interpolation Based on 3D Motion Representation and Spatial Supervision

Haojie Liu, Kang Liao, Chunyu Lin, Yao Zhao, Yulan Guo

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09917 2020-06-18 cs.CV cs.LG 57%

FISHING Net: Future Inference of Semantic Heatmaps In Grids

Noureldin Hendy, Cooper Sloan, Feng Tian, Pengfei Duan, Nick Charchut, Yuesong Xie, Chuang Wang, James Philbin

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08837 2020-06-09 cs.CV 57%

Learning Shared Cross-modality Representation Using Multispectral-LiDAR and Hyperspectral Data

Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Jian Kang, Xiao Xiang Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Journal ref IEEE Geoscience and Remote Sensing Letters, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08154 2020-05-22 cs.CV cs.LG 57%

Detailed 2D-3D Joint Representation for Human-Object Interaction

Yong-Lu Li, Xinpeng Liu, Han Lu, Shiyi Wang, Junqi Liu, Jiefeng Li, Cewu Lu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2020, supplementary materials included, code available:https://github.com/DirtyHarryLYL/DJ-RN

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00754 2020-05-06 cs.CV 57%

CoMoGCN: Coherent Motion Aware Trajectory Prediction with Graph Representation

Yuying Chen, Congcong Liu, Bertram Shi, Ming Liu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.09467 2020-04-21 cs.CV 57%

Disentangled Representation Learning in Cardiac Image Analysis

Agisilaos Chartsias, Thomas Joyce, Giorgos Papanastasiou, Michelle Williams, David Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Journal ref Medical Image Analysis 58 (2019) 101535

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11477 2020-04-16 cs.CV cs.LG stat.ML 57%

Learning a Directional Soft Lane Affordance Model for Road Scenes Using Self-Supervision

Robin Karlsson, Erik Sjoberg

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted for IEEE IV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07414 2020-04-09 cs.CV 57%

Potential Field: Interpretable and Unified Representation for Trajectory Prediction

Shan Su, Cheng Peng, Jianbo Shi, Chiho Choi

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02012 2020-03-26 cs.CV cs.NA math.NA 57%

Variational Osmosis for Non-linear Image Fusion

Simone Parisotto, Luca Calatroni, Aurélie Bugeau, Nicolas Papadakis, Carola-Bibiane Schönlieb

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.11559 2020-03-11 cs.CL 57%

Variational Self-attention Model for Sentence Representation

Qiang Zhang, Shangsong Liang, Emine Yilmaz

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:1511.06038 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏