arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2108.06281 2022-02-09 cs.CV 57%

Modal-Adaptive Gated Recoding Network for RGB-D Salient Object Detection

Jinchao Zhu, Xiaoyu Zhang, Xian Fang, Feng Dong, Qiu Yu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10147 2022-02-07 cs.CV 57%

TGFuse: An Infrared and Visible Image Fusion Approach Based on Transformer and Generative Adversarial Network

Dongyu Rao, Xiao-Jun Wu, Tianyang Xu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10656 2022-01-27 cs.CV 57%

MGA-VQA: Multi-Granularity Alignment for Visual Question Answering

Peixi Xiong, Yilin Shen, Hongxia Jin

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09574 2022-01-25 cs.CV 57%

Multi-Scale Iterative Refinement Network for RGB-D Salient Object Detection

Ze-yu Liu, Jian-wei Liu, Xin Zuo, Ming-fei Hu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 40 pages

Journal ref Engineering Applications of Artificial Intelligence(2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08673 2022-01-24 cs.CV 57%

Exploring Fusion Strategies for Accurate RGBT Visual Object Tracking

Zhangyong Tang, Tianyang Xu, Hui Li, Xiao-Jun Wu, Xuefeng Zhu, Josef Kittler

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 13 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07451 2022-01-20 cs.CV 57%

TransFuse: A Unified Transformer-based Image Fusion Framework using Self-supervised Learning

Linhao Qu, Shaolei Liu, Manning Wang, Shiman Li, Siqi Yin, Qin Qiao, Zhijian Song

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06784 2022-01-10 cs.CV cs.LG 57%

Exploring Adversarial Robustness of Multi-Sensor Perception Systems in Self Driving

James Tu, Huichen Li, Xinchen Yan, Mengye Ren, Yun Chen, Ming Liang, Eilyan Bitar, Ersin Yumer, Raquel Urtasun

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.10483 2021-12-21 cs.CV 57%

Fusion and Orthogonal Projection for Improved Face-Voice Association

Muhammad Saad Saeed, Muhammad Haris Khan, Shah Nawaz, Muhammad Haroon Yousaf, Alessio Del Bue

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13274 2021-12-21 cs.CV 57%

Adversarial Domain Adaptation with Prototype-Based Normalized Output Conditioner

Dapeng Hu, Jian Liang, Qibin Hou, Hanshu Yan, Yunpeng Chen

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Published at IEEE transactions on image processing 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01179 2021-12-15 cs.CV 57%

BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion

Zhaoqun Li, Xu Liang, Dandan Fan, Jinxing Li, David Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Extended version of ICONIP 2021 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09554 2021-12-10 cs.CV 57%

TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation

Haoyu Ma, Liangjian Chen, Deying Kong, Zhe Wang, Xingwei Liu, Hao Tang, Xiangyi Yan, Yusheng Xie, Shih-Yao Lin, Xiaohui Xie

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments BMVC 2021. Code is available at: https://github.com/HowieMa/TransFusion-Pose

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12482 2021-12-06 cs.CV 57%

Self-Supervised Pretraining for RGB-D Salient Object Detection

Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu, Xiang Ruan

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments This work was accepted by AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01522 2021-12-03 cs.CV 57%

Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks

Xizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu, Xiaogang Wang, Hongsheng Li, Xiaohua Wang, Jifeng Dai

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00219 2021-12-02 cs.CV cs.RO 57%

Scalable Primitives for Generalized Sensor Fusion in Autonomous Vehicles

Sammy Sidhu, Linda Wang, Tayyab Naseer, Ashish Malhotra, Jay Chia, Aayush Ahuja, Ella Rasmussen, Qiangui Huang, Ray Gao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Presented in Machine Learning for Autonomous Driving Workshop at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Sydney, Australia. 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14382 2021-12-02 cs.CV 57%

VPFNet: Improving 3D Object Detection with Virtual Point based LiDAR and Stereo Data Fusion

Hanqi Zhu, Jiajun Deng, Yu Zhang, Jianmin Ji, Qiuyu Mao, Houqiang Li, Yanyong Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15509 2021-12-01 cs.CV 57%

Regularized directional representations for medical image registration

Vincent Jaouen, Pierre-Henri Conze, Guillaume Dardenne, Julien Bert, Dimitris Visvikis

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12798 2021-11-29 cs.LG cs.CV 57%

Geometric Priors for Scientific Generative Models in Inertial Confinement Fusion

Ankita Shukla, Rushil Anirudh, Eugene Kur, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Brian K. Spears, Tammy Ma, Pavan Turaga

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 4 figures, Fourth Workshop on Machine Learning and the Physical Sciences, NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11862 2021-11-24 cs.CV cs.HC cs.LG 57%

Inferring User Facial Affect in Work-like Settings

Chaudhary Muhammad Aqdus Ilyas, Siyang Song, Hatice Gunes

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11075 2021-10-22 cs.RO cs.AI cs.HC 57%

Enabling a Social Robot to Process Social Cues to Detect when to Help a User

Jason R. Wilson, Phyo Thuta Aung, Isabelle Boucher

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Presented at AI-HRI symposium as part of AAAI-FSS 2021 (arXiv:2109.10836)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10490 2021-10-19 cs.CV cs.RO 57%

FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular Cameras

Anthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas, Jeffrey Hawke, Vijay Badrinarayanan, Roberto Cipolla, Alex Kendall

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04279 2021-10-11 eess.IV cs.CV cs.LG cs.NE q-bio.NC 57%

StairwayGraphNet for Inter- and Intra-modality Multi-resolution Brain Graph Alignment and Synthesis

Islem Mhiri, Mohamed Ali Mahjoub, Islem Rekik

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2107.06281

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03452 2021-10-08 eess.IV cs.CV cs.LG q-bio.NC 57%

Inter-Domain Alignment for Predicting High-Resolution Brain Networks Using Teacher-Student Learning

Basar Demir, Alaa Bessadok, Islem Rekik

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14557 2021-10-01 cs.CV 57%

Learning Deformable Image Registration from Optimization: Perspective, Modules, Bilevel Training and Beyond

Risheng Liu, Zi Li, Xin Fan, Chenying Zhao, Hao Huang, Zhongxuan Luo

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06363 2021-09-16 cs.CV cs.LG 57%

Sensor Adversarial Traits: Analyzing Robustness of 3D Object Detection Sensor Fusion Models

Won Park, Nan Liu, Qi Alfred Chen, Z. Morley Mao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Journal ref 2021 IEEE International Conference on Image Processing (ICIP), 2021, pp. 484-488

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03425 2021-09-09 cs.CV 57%

RGB-D Salient Object Detection with Ubiquitous Target Awareness

Yifan Zhao, Jiawei Zhao, Jia Li, Xiaowu Chen

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 15 pages, 13 figures, Accepted by IEEE Transactions on Image Processing (2021). arXiv admin note: text overlap with arXiv:2006.00269

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02227 2021-09-07 cs.CV 57%

Learning to Generate Scene Graph from Natural Language Supervision

Yiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu, Yin Li

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01078 2021-09-03 cs.CL 57%

Skim-Attention: Learning to Focus via Document Layout

Laura Nguyen, Thomas Scialom, Jacopo Staiano, Benjamin Piwowarski

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 15 pages, 6 figures, to be published in EMNLP 2021 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08608 2021-09-01 cs.CV 57%

DPANet: Depth Potentiality-Aware Gated Attention Network for RGB-D Salient Object Detection

Zuyao Chen, Runmin Cong, Qianqian Xu, Qingming Huang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12863 2021-08-31 cs.CV 57%

MBDF-Net: Multi-Branch Deep Fusion Network for 3D Object Detection

Xun Tan, Xingyu Chen, Guowei Zhang, Jishiyu Ding, Xuguang Lan

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.03289 2021-08-26 cs.CV 57%

Question-Agnostic Attention for Visual Question Answering

Moshiur R Farazi, Salman H Khan, Nick Barnes

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments To appear in the proceedings of International Conference on Pattern Recognition (ICPR) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏