arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2308.04352 2023-08-09 cs.CV 57%

3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment

Ziyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng, Siyuan Huang, Qing Li

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11427 2023-08-08 cs.CV 57%

RGB-D And Thermal Sensor Fusion: A Systematic Literature Review

Martin Brenner, Napoleon H. Reyes, Teo Susnjak, Andre L. C. Barczak

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 34 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16142 2023-08-01 eess.IV cs.CV 57%

Implicit Neural Representation in Medical Imaging: A Comparative Survey

Amirali Molaei, Amirhossein Aminimehr, Armin Tavakoli, Amirhossein Kazerouni, Bobby Azad, Reza Azad, Dorit Merhof

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09749 2023-08-01 cs.CV 57%

Towards Robust Scene Text Image Super-resolution via Explicit Location Enhancement

Hang Guo, Tao Dai, Guanghao Meng, Shu-Tao Xia

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted as IJCAI2023 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13323 2023-07-26 cs.RO cs.AI 57%

Learning Autonomous Ultrasound via Latent Task Representation and Robotic Skills Adaptation

Xutian Deng, Junnan Jiang, Wen Cheng, Miao Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09731 2023-07-19 cs.CV eess.IV 57%

Semantic Labeling of High Resolution Images Using EfficientUNets and Transformers

Hasan AlMarzouqi, Lyes Saad Saoud

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14365 2023-07-17 cs.CV 57%

Image-Specific Information Suppression and Implicit Local Alignment for Text-based Person Search

Shuanglin Yan, Hao Tang, Liyan Zhang, Jinhui Tang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06771 2023-07-14 eess.IV cs.CV cs.LG 57%

Generalizing Supervised Deep Learning MRI Reconstruction to Multiple and Unseen Contrasts using Meta-Learning Hypernetworks

Sriprabha Ramanarayanan, Arun Palla, Keerthi Ram, Mohanasankar Sivaprakasam

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in Elsevier Applied Soft Computing Journal, 36 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05437 2023-06-28 cs.LG cs.AI 57%

One-step Multi-view Clustering with Diverse Representation

Xinhang Wan, Jiyuan Liu, Xinwang Liu, Siwei Wang, Yi Wen, Tianjiao Wan, Li Shen, En Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13380 2023-06-26 cs.CV 57%

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li, Zechuan Li, Jingwen Wang, Wei Miao, Wei Sun, Chen Chen

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Winner of CVPR2023 Long-form Video Understanding and Generation Challenge (Track 3)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10764 2023-06-21 cs.CV 57%

OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding

Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, Hao Su

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Project Website: https://colin97.github.io/OpenShape/

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01108 2023-06-21 eess.IV cs.CV 57%

Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis

Yonghao Li, Tao Zhou, Kelei He, Yi Zhou, Dinggang Shen

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 13 pages, 16 figures. This paper has been accepted by IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16697 2023-06-19 cs.CV cs.HC 57%

SGDraw: Scene Graph Drawing Interface Using Object-Oriented Representation

Tianyu Zhang, Xusheng Du, Chia-Ming Chang, Xi Yang, Haoran Xie

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 9 figures, video is https://youtu.be/acy0SNLfahg, accepted in HCI International 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06500 2023-06-16 cs.CV cs.LG 57%

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07344 2023-06-14 cs.RO cs.CV 57%

Towards a Robust Sensor Fusion Step for 3D Object Detection on Corrupted Data

Maciej K. Wozniak, Viktor Karefjards, Marko Thiel, Patric Jensfelt

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.05104 2023-06-06 cs.CV 57%

Improving Cytoarchitectonic Segmentation of Human Brain Areas with Self-supervised Siamese Networks

Hannah Spitzer, Kai Kiwitz, Katrin Amunts, Stefan Harmeling, Timo Dickscheid

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted at MICCAI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19924 2023-06-02 cs.CV 57%

Joint Adaptive Representations for Image-Language Learning

AJ Piergiovanni, Anelia Angelova

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments T4V Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00932 2023-05-26 cs.CV 57%

HypLiLoc: Towards Effective LiDAR Pose Regression with Hyperbolic Fusion

Sijie Wang, Qiyu Kang, Rui She, Wei Wang, Kai Zhao, Yang Song, Wee Peng Tay

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14969 2023-05-25 cs.CV 57%

MMNet: Multi-Mask Network for Referring Image Segmentation

Yichen Yan, Xingjian He, Wenxuan Wan, Jing Liu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13220 2023-05-23 cs.CV 57%

Fast Monocular Scene Reconstruction with Global-Sparse Local-Dense Grids

Wei Dong, Chris Choy, Charles Loop, Or Litany, Yuke Zhu, Anima Anandkumar

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02318 2023-05-23 cs.CV 57%

Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative Pretraining

Zekun Qi, Runpei Dong, Guofan Fan, Zheng Ge, Xiangyu Zhang, Kaisheng Ma, Li Yi

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08691 2023-05-23 cs.CV cs.RO 57%

Towards Long-Tailed 3D Detection

Neehar Peri, Achal Dave, Deva Ramanan, Shu Kong

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments This work has been accepted to the Conference on Robot Learning (CoRL) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09999 2023-05-18 cs.CV 57%

An Interactively Reinforced Paradigm for Joint Infrared-Visible Image Fusion and Saliency Object Detection

Di Wang, Jinyuan Liu, Risheng Liu, Xin Fan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09641 2023-05-17 cs.CV cs.GR cs.LG 57%

FitMe: Deep Photorealistic 3D Morphable Model Avatars

Alexandros Lattas, Stylianos Moschoglou, Stylianos Ploumpis, Baris Gecer, Jiankang Deng, Stefanos Zafeiriou

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2023, project page at https://lattas.github.io/fitme , 17 pages including supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14323 2023-05-12 cs.AI 57%

Pushing the Boundaries of Tractable Multiperspective Reasoning: A Deduction Calculus for Standpoint EL+

Lucía Gómez Álvarez, Sebastian Rudolph, Hannes Strass

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07133 2023-05-12 cs.CV 57%

CLIP-Lite: Information Efficient Visual Representation Learning with Language Supervision

Aman Shrivastava, Ramprasaath R. Selvaraju, Nikhil Naik, Vicente Ordonez

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05166 2023-05-11 cs.CL 57%

E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation

Cong Ma, Yaping Zhang, Mei Tu, Yang Zhao, Yu Zhou, Chengqing Zong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL

Comments Accepted at The 17th International Conference on Document Analysis and Recognition (ICDAR 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13187 2023-05-11 cs.AI 57%

Tractable Diversity: Scalable Multiperspective Ontology Management via Standpoint EL

Lucía Gómez Álvarez, Sebastian Rudolph, Hannes Strass

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Journal ref 32nd International Joint Conference on Artificial Intelligence, IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03041 2023-05-11 cs.LG cs.AI q-bio.QM 57%

Keeping it Simple: Language Models can learn Complex Molecular Distributions

Daniel Flam-Shepherd, Kevin Zhu, Alán Aspuru-Guzik

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Journal ref Nat Commun 13, 3293 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03314 2023-05-08 cs.CL 57%

Block the Label and Noise: An N-Gram Masked Speller for Chinese Spell Checking

Haiyun Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏