arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2304.10759 2023-04-24 cs.CV cs.CL 62%

GeoLayoutLM: Geometric Pre-training for Visual Information Extraction

Chuwei Luo, Changxu Cheng, Qi Zheng, Cong Yao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2023 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07200 2023-04-18 cs.LG cs.AI cs.CV stat.ML 62%

Zero-Shot Compositional Policy Learning via Language Grounding

Tianshi Cao, Jingkang Wang, Yining Zhang, Sivabalan Manivasagam

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Extended version for ICLRW-2020 (BeTR-RL) paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02228 2023-04-04 eess.IV cs.CL cs.CV 62%

MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology

Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13518 2023-03-24 cs.CV cs.AI cs.LG 62%

Three ways to improve feature alignment for open vocabulary detection

Relja Arandjelović, Alex Andonian, Arthur Mensch, Olivier J. Hénaff, Jean-Baptiste Alayrac, Andrew Zisserman

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10421 2023-03-21 cs.CV cs.AI 62%

Mutilmodal Feature Extraction and Attention-based Fusion for Emotion Estimation in Videos

Tao Shu, Xinke Wang, Ruotong Wang, Chuang Chen, Yixin Zhang, Xiao Sun

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08129 2023-03-15 cs.CV cs.AI 62%

PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object Detection

Anthony Chen, Kevin Zhang, Renrui Zhang, Zihan Wang, Yuheng Lu, Yandong Guo, Shanghang Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR2023. Code is available at https://github.com/BLVLab/PiMAE

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04364 2023-03-09 cs.AI cs.CV cs.LG cs.RO 62%

Dynamic Scenario Representation Learning for Motion Forecasting with Heterogeneous Graph Convolutional Recurrent Networks

Xing Gao, Xiaogang Jia, Yikang Li, Hongkai Xiong

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02635 2023-03-07 cs.CV cs.MM 62%

VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning

Kang Chen, Xiangqian Wu

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14777 2023-03-01 cs.CV cs.AI 62%

VQA with Cascade of Self- and Co-Attention Blocks

Aakansha Mishra, Ashish Anand, Prithwijit Guha

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10951 2023-01-27 cs.CV cs.CL 62%

Cross Modal Global Local Representation Learning from Radiology Reports and X-Ray Chest Images

Nathan Hadjiyski, Ali Vosoughi, Axel Wismueller

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to Computer-Aided Diagnosis, SPIE Medical Imaging 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12392 2023-01-18 cs.AI cs.CL 62%

Emergent Communication through Metropolis-Hastings Naming Game with Deep Generative Models

Tadahiro Taniguchi, Yuto Yoshida, Akira Taniguchi, Yoshinobu Hagiwara

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04352 2023-01-12 cs.CV cs.AI cs.RO 62%

Graph based Environment Representation for Vision-and-Language Navigation in Continuous Environments

Ting Wang, Zongkai Wu, Feiyu Yao, Donglin Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00622 2023-01-03 cs.CV cs.AI 62%

Credible Remote Sensing Scene Classification Using Evidential Fusion on Aerial-Ground Dual-view Images

Kun Zhao, Qian Gao, Siyuan Hao, Jie Sun, Lijian Zhou

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 16 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14024 2022-12-08 cs.CV cs.AI cs.LG cs.RO 62%

Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer

Hao Shao, Letian Wang, RuoBing Chen, Hongsheng Li, Yu Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CoRL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15162 2022-11-29 cs.IR cs.CL cs.CV cs.LG 62%

Long-tail Cross Modal Hashing

Zijun Gao, Jun Wang, Guoxian Yu, Zhongmin Yan, Carlotta Domeniconi, Jinglin Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by the Thirty-Seventh AAAI Conference on Artificial Intelligence(AAAI2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14732 2022-11-29 cs.LG cs.AI cs.CV 62%

Deep representation learning: Fundamentals, Perspectives, Applications, and Open Challenges

Kourosh T. Baghaei, Amirreza Payandeh, Pooya Fayyazsanavi, Shahram Rahimi, Zhiqian Chen, Somayeh Bakhtiari Ramezani

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11153 2022-11-22 cs.LG cs.CL cs.CV 62%

Unifying Vision-Language Representation Space with Single-tower Transformer

Jiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim, Nojun Kwak

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI 2023, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10681 2022-11-22 cs.CV cs.AI 62%

Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning

Xiaocheng Lu, Ziming Liu, Song Guo, Jingcai Guo

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages included reference, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14163 2022-11-22 cs.CV cs.MM 62%

Multi-Granularity Cross-Modality Representation Learning for Named Entity Recognition on Social Media

Peipei Liu, Gaosheng Wang, Hong Li, Jie Liu, Yimo Ren, Hongsong Zhu, Limin Sun

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.MM

Comments We have reconducted experiments of the paper, but found that there were fatal errors in our datasets leading to the wrong results and analyses. Therefore, we have to withdraw the paper to ensure the authenticity of science. We are very sorry

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13462 2022-10-28 cs.LG cs.AI cs.CV 62%

Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data

Farida Mohsen, Hazrat Ali, Nady El Hajj, Zubair Shah

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted in Nature Scientific Reports. 20 pages

Journal ref Sci Rep 12, 17981 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13617 2022-10-27 cs.CL cs.AI 62%

Adapters for Enhanced Modeling of Multilingual Knowledge and Text

Yifan Hou, Wenxiang Jiao, Meizhen Liu, Carl Allen, Zhaopeng Tu, Mrinmaya Sachan

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CL、cs.AI

Comments Our code, models, and data (e.g., integration corpus and extended datasets) are available: https://github.com/yifan-h/Multilingual_Space

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12798 2022-10-25 cs.CL cs.AI cs.LG 62%

MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality Sequences

Wei Han, Hui Chen, Min-Yen Kan, Soujanya Poria

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted as a long paper at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10684 2022-10-20 cs.CL cs.AI 62%

Language Models Understand Us, Poorly

Jared Moore

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 1 figure, to be published in Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12729 2022-09-28 cs.CV cs.AI cs.LG cs.RO 62%

DeepFusion: A Robust and Modular 3D Object Detector for Lidars, Cameras and Radars

Florian Drews, Di Feng, Florian Faion, Lars Rosenbaum, Michael Ulrich, Claudius Gläser

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05993 2022-08-03 cs.CV cs.AI 62%

A new database of Houma Alliance Book ancient handwritten characters and classifier fusion approach

Xiaoyu Yuan, Zhibo Zhang, Yabo Sun, Zekai Xue, Xiuyan Shao, Xiaohua Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 7 figures, accepted by 14th International Conference on Graphics and Image Processing (ICGIP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05188 2022-07-04 cs.CL cs.SD eess.AS 62%

Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems

Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas, Hong-Kwang J. Kuo, Brian Kingsbury

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.14355 2022-06-30 cs.CV cs.CL cs.LG 62%

EBMs vs. CL: Exploring Self-Supervised Visual Pretraining for Visual Question Answering

Violetta Shevchenko, Ehsan Abbasnejad, Anthony Dick, Anton van den Hengel, Damien Teney

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.00100 2022-06-02 cs.CV cs.CL 62%

VALHALLA: Visual Hallucination for Machine Translation

Yi Li, Rameswar Panda, Yoon Kim, Chun-Fu Chen, Rogerio Feris, David Cox, Nuno Vasconcelos

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14095 2022-05-31 cs.CV cs.AI 62%

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li, Rongrong Ji, Chunhua Shen

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.13629 2022-05-30 cs.CV cs.AI cs.RO 62%

Deep Sensor Fusion with Pyramid Fusion Networks for 3D Semantic Segmentation

Hannah Schieber, Fabian Duerr, Torsten Schoen, Jürgen Beyerer

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments conditionally accepted at IEEE IV 2022, 7 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏