arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2108.09916 2021-08-24 cs.CV 57%

PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose Estimation

Guangyuan Zhou, Huiqun Wang, Jiaxin Chen, Di Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08012 2021-08-19 cs.CV 57%

Multi-Anchor Active Domain Adaptation for Semantic Segmentation

Munan Ning, Donghuan Lu, Dong Wei, Cheng Bian, Chenglang Yuan, Shuang Yu, Kai Ma, Yefeng Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments ICCV 2021 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07979 2021-08-19 cs.CV eess.IV 57%

A New Bidirectional Unsupervised Domain Adaptation Segmentation Framework

Munan Ning, Cheng Bian, Dong Wei, Chenglang Yuan, Yaohua Wang, Yang Guo, Kai Ma, Yefeng Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments IPMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11872 2021-08-18 cs.CV 57%

Mixture-based Feature Space Learning for Few-shot Image Classification

Arman Afrasiyabi, Jean-François Lalonde, Christian Gagné

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04602 2021-08-11 cs.CV 57%

Joint Multi-Object Detection and Tracking with Camera-LiDAR Fusion for Autonomous Driving

Kemiao Huang, Qi Hao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments accepted by IROS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12964 2021-07-30 cs.CV cs.LG cs.SD eess.SP 57%

A Physiologically-Adapted Gold Standard for Arousal during Stress

Alice Baird, Lukas Stappen, Lukas Christ, Lea Schumann, Eva-Maria Meßner, Björn W. Schuller

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07853 2021-07-27 cs.CV 57%

G2DA: Geometry-Guided Dual-Alignment Learning for RGB-Infrared Person Re-Identification

Lin Wan, Zongyuan Sun, Qianyan Jing, Yehansen Chen, Lijing Lu, Zhihang Li

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04075 2021-07-20 cs.CV cs.LG cs.RO 57%

Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data Alignment

Xueying Shi, Yueming Jin, Qi Dou, Jing Qin, Pheng-Ann Heng

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted as a conference paper in IROS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01779 2021-07-07 cs.CV 57%

Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object Detection

Wenbo Zhang, Ge-Peng Ji, Zhuo Wang, Keren Fu, Qijun Zhao

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments accepted in ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12658 2021-06-25 cs.LG cs.AI 57%

Transformer-based unsupervised patient representation learning based on medical claims for risk stratification and analysis

Xianlong Zeng, Simon Lin, Chang Liu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.08617 2021-06-17 cs.CV 57%

CMF: Cascaded Multi-model Fusion for Referring Image Segmentation

Jianhua Yang, Yan Huang, Zhanyu Ma, Liang Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICIP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01797 2021-06-15 cs.CL 57%

TVDIM: Enhancing Image Self-Supervised Pretraining via Noisy Text Data

Pengda Qin, Yuhong Li, Kefeng Deng, Qiang Wu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13993 2021-05-31 eess.IV cs.CV cs.LG 57%

PTNet: A High-Resolution Infant MRI Synthesizer Based on Transformer

Xuzhe Zhang, Xinzi He, Jia Guo, Nabil Ettehadi, Natalie Aw, David Semanek, Jonathan Posner, Andrew Laine, Yun Wang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments arXiv Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08837 2021-05-20 cs.RO cs.CV 57%

Fusion-DHL: WiFi, IMU, and Floorplan Fusion for Dense History of Locations in Indoor Environments

Sachini Herath, Saghar Irandoust, Bowen Chen, Yiming Qian, Pyojin Kim, Yasutaka Furukawa

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments To be published in ICRA 2021. Code and data: https://github.com/Sachini/Fusion-DHL

Journal ref ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05708 2021-05-13 cs.CV 57%

Deep and Shallow Covariance Feature Quantization for 3D Facial Expression Recognition

Walid Hariri, Nadir Farah, Dinesh Kumar Vishwakarma

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.00490 2021-05-04 cs.CV 57%

Residual Enhanced Multi-Hypergraph Neural Network

Jing Huang, Xiaolin Huang, Jie Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ICIP 2021 submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06231 2021-04-21 eess.IV cs.CV 57%

Latent Correlation Representation Learning for Brain Tumor Segmentation with Missing MRI Modalities

Tongxue Zhou, Stéphane Canu, Pierre Vera, Su Ruan

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 12 pages, 10 figures, accepted by IEEE Transactions on Image Processing (8 April 2021). arXiv admin note: text overlap with arXiv:2003.08870, arXiv:2102.03111

Journal ref IEEE Transactions on Image Processing On page(s): 4263-4274 Print ISSN: 1057-7149 Online ISSN: 1941-0042

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08870 2021-04-21 eess.IV cs.CV cs.LG 57%

Brain tumor segmentation with missing modalities via latent multi-source correlation representation

Tongxue Zhou, Stéphane Canu, Pierre Vera, Su Ruan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 6 figures, accepted by MICCAI 2020. arXiv admin note: text overlap with arXiv:2102.03111

Journal ref MICCAI 2020 pp. 533-541

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09122 2021-04-20 cs.LG cs.AI 57%

Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning

Jie Ren, Yewen Li, Zihan Ding, Wei Pan, Hao Dong

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.05327 2021-04-15 cs.CV 57%

MinkLoc++: Lidar and Monocular Image Fusion for Place Recognition

Jacek Komorowski, Monika Wysoczanska, Tomasz Trzcinski

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12936 2021-04-07 cs.CV cs.RO 57%

6D Pose Estimation with Correlation Fusion

Yi Cheng, Hongyuan Zhu, Ying Sun, Cihan Acar, Wei Jing, Yan Wu, Liyuan Li, Cheston Tan, Joo-Hwee Lim

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICPR2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00652 2021-04-06 cs.CV 57%

Depth as Attention for Face Representation Learning

Hardik Uppal, Alireza Sepas-Moghaddam, Michael Greenspan, Ali Etemad

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 11 figures, Accepted to IEEE Transactions on Information Forensics and Security 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.16549 2021-03-31 cs.CV 57%

Deep Gaussian Processes for Few-Shot Segmentation

Joakim Johnander, Johan Edstedt, Martin Danelljan, Michael Felsberg, Fahad Shahbaz Khan

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04781 2021-03-23 cs.CV cs.LG eess.IV 57%

You Only Need Adversarial Supervision for Semantic Image Synthesis

Vadim Sushko, Edgar Schönfeld, Dan Zhang, Juergen Gall, Bernt Schiele, Anna Khoreva

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Published at ICLR 2021 (Main Conference). Code repository: https://github.com/boschresearch/OASIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.00359 2021-03-02 cs.CV cs.LG 57%

The Labeled Multiple Canonical Correlation Analysis for Information Fusion

Lei Gao, Rui Zhang, Lin Qi, Enqing Chen, Ling Guan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Multimedia, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.13313 2021-02-24 cs.CV cs.RO eess.IV 57%

Polarization-driven Semantic Segmentation via Efficient Attention-bridged Fusion

Kaite Xiang, Kailun Yang, Kaiwei Wang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by Optics Express. 18 pages, 16 figures, 3 tables, 9 equations. Code will be made publicly available at https://github.com/Katexiang/EAFNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05123 2021-02-03 cs.RO cs.CV 57%

Orientation Attentive Robotic Grasp Synthesis with Augmented Grasp Map Representation

Georgia Chalvatzaki, Nikolaos Gkanatsios, Petros Maragos, Jan Peters

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 7 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.12673 2021-01-19 cs.CV 57%

Semantic Implicit Neural Scene Representations With Semi-Supervised Training

Amit Kohli, Vincent Sitzmann, Gordon Wetzstein

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 3DV 2020 Camera Ready https://www.computationalimaging.org/publications/

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08226 2020-12-18 cs.CV 57%

Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation

Minsu Kim, Sunghun Joung, Seungryong Kim, JungIn Park, Ig-Jae Kim, Kwanghoon Sohn

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06843 2020-12-15 cs.CV 57%

Multi-Scale Cascading Network with Compact Feature Learning for RGB-Infrared Person Re-Identification

Can Zhang, Hong Liu, Wei Guo, Mang Ye

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures, ICPR2020 conference

详情

展开后加载摘要…

URL PDF HTML 收藏