arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2311.13194 2023-12-18 cs.CV 57%

Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs

Yonghui Wang, Wengang Zhou, Hao Feng, Keyi Zhou, Houqiang Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08876 2023-12-15 cs.CV 57%

OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection

Hu Zhang, Jianhua Xu, Tao Tang, Haiyang Sun, Xin Yu, Zi Huang, Kaicheng Yu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08851 2023-12-15 cs.CV cs.CE cs.RO 57%

Achelous++: Power-Oriented Water-Surface Panoptic Perception Framework on Edge Devices based on Vision-Radar Fusion and Pruning of Heterogeneous Modalities

Runwei Guan, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Xiaohui Zhu, Limin Yu, Yong Yue, Jeremy Smith, Eng Gee Lim, Weiping Ding, Yutao Yue

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03661 2023-12-15 cs.CV 57%

Prompt-based Context- and Domain-aware Pretraining for Vision and Language Navigation

Ting Liu, Yue Hu, Wansen Wu, Youkai Wang, Kai Xu, Quanjun Yin

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12530 2023-12-12 eess.AS cs.SD 57%

Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio

Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 eess.AS

Comments Proceedings of Interspeech 2023; v4 version updates: correction of W2V2-base pretrained on 960-hour of LibriSpeech and number of families participated for LENA home recordings

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04891 2023-12-11 cs.CV 57%

Cross-BERT for Point Cloud Pretraining

Xin Li, Peng Li, Zeyong Wei, Zhe Zhu, Mingqiang Wei, Junhui Hou, Liangliang Nan, Jing Qin, Haoran Xie, Fu Lee Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04160 2023-12-08 cs.CV 57%

Text as Image: Learning Transferable Adapter for Multi-Label Classification

Xuelin Zhu, Jiuxin Cao, Jian liu, Dongqi Tang, Furong Xu, Weijia Liu, Jiawei Ge, Bo Liu, Qingpei Guo, Tianyi Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13355 2023-12-08 cs.CV 57%

SILC: Improving Vision Language Pretraining with Self-Distillation

Muhammad Ferjad Naeem, Yongqin Xian, Xiaohua Zhai, Lukas Hoyer, Luc Van Gool, Federico Tombari

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06791 2023-12-07 cs.CV 57%

InfMLLM: A Unified Framework for Visual-Language Tasks

Qiang Zhou, Zhibin Wang, Wei Chu, Yinghui Xu, Hao Li, Yuan Qi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 8

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01732 2023-12-05 cs.CV 57%

Likelihood-Aware Semantic Alignment for Full-Spectrum Out-of-Distribution Detection

Fan Lu, Kai Zhu, Kecheng Zheng, Wei Zhai, Yang Cao

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00874 2023-12-05 cs.CL 57%

Hi-ArG: Exploring the Integration of Hierarchical Argumentation Graphs in Language Pretraining

Jingcong Liang, Rong Ye, Meng Han, Qi Zhang, Ruofei Lai, Xinyu Zhang, Zhao Cao, Xuanjing Huang, Zhongyu Wei

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

Comments to be published in EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13653 2023-12-04 cs.CV 57%

RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search

Yang Bai, Min Cao, Daming Gao, Ziqiang Cao, Chen Chen, Zhenfeng Fan, Liqiang Nie, Min Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2023. Code is available at https://github.com/Flame-Chasers/RaSa

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13126 2023-11-23 cs.CL 57%

Towards Better Parameter-Efficient Fine-Tuning for Large Language Models: A Position Paper

Chengyu Wang, Junbing Yan, Wei Zhang, Jun Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13069 2023-11-23 cs.CV 57%

FuseNet: Self-Supervised Dual-Path Network for Medical Image Segmentation

Amirhossein Kazerouni, Sanaz Karimijafarbigloo, Reza Azad, Yury Velichko, Ulas Bagci, Dorit Merhof

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12602 2023-11-22 cs.CV cs.LG 57%

TouchSDF: A DeepSDF Approach for 3D Shape Reconstruction using Vision-Based Tactile Sensing

Mauro Comi, Yijiong Lin, Alex Church, Alessio Tonioni, Laurence Aitchison, Nathan F. Lepora

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10996 2023-11-22 cs.CV 57%

ProtoCLIP: Prototypical Contrastive Language Image Pretraining

Delong Chen, Zhao Wu, Fan Liu, Zaiquan Yang, Huaxi Huang, Ying Tan, Erjin Zhou

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Neural Networks and Learning Systems (TNNLS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07514 2023-11-14 cs.CV 57%

VGSG: Vision-Guided Semantic-Group Network for Text-based Person Search

Shuting He, Hao Luo, Wei Jiang, Xudong Jiang, Henghui Ding

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08789 2023-11-10 cs.RO cs.AI cs.LG 57%

PLEX: Making the Most of the Available Data for Robotic Manipulation Pretraining

Garrett Thomas, Ching-An Cheng, Ricky Loynd, Felipe Vieira Frujeri, Vibhav Vineet, Mihai Jalobeanu, Andrey Kolobov

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03620 2023-11-08 cs.CV 57%

FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Xinhao Xiang, Jiawei Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01308 2023-11-03 eess.IV cs.CV 57%

Hybrid-Fusion Transformer for Multisequence MRI

Jihoon Cho, Jinah Park

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19898 2023-11-01 cs.CV cs.LG 57%

MIST: Medical Image Segmentation Transformer with Convolutional Attention Mixing (CAM) Decoder

Md Motiur Rahman, Shiva Shokouhmand, Smriti Bhatt, Miad Faezipour

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 2 figures, 3 tables, accepted for publication in WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19743 2023-10-31 cs.LG cs.CV 57%

Tell Me What Is Good About This Property: Leveraging Reviews For Segment-Personalized Image Collection Summarization

Monika Wysoczanska, Moran Beladev, Karen Lastmann Assaraf, Fengjun Wang, Ofri Kleinfeld, Gil Amsalem, Hadas Harush Boker

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07288 2023-10-31 cs.CV 57%

Implicit Neural Feature Fusion Function for Multispectral and Hyperspectral Image Fusion

ShangQi Deng, RuoCheng Wu, Liang-Jian Deng, Ran Ran, Gemine Vivone

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06589 2023-10-27 cs.LG cs.AI 57%

Towards Better Generalization with Flexible Representation of Multi-Module Graph Neural Networks

Hyungeun Lee, Kijung Yoon

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Journal ref Transactions on Machine Learning Research (TMLR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16667 2023-10-26 cs.CV 57%

CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan, Xiaojuan Qi

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09347 2023-10-25 cs.CV cs.LG cs.RO 57%

Segment Any Point Cloud Sequences by Distilling Vision Foundation Models

Youquan Liu, Lingdong Kong, Jun Cen, Runnan Chen, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2023 (Spotlight); 37 pages, 16 figures, 15 tables; Code at https://github.com/youquanl/Segment-Any-Point-Cloud

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14158 2023-10-24 cs.CV 57%

Visual-Attribute Prompt Learning for Progressive Mild Cognitive Impairment Prediction

Luoyao Kang, Haifan Gong, Xiang Wan, Haofeng Li

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments MICCAI 2023, released code: https://github.com/lhaof/VAPL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14577 2023-10-19 cs.LG cs.CL 57%

Difference-Masking: Choosing What to Mask in Continued Pretraining

Alex Wilf, Syeda Nahida Akter, Leena Mathur, Paul Pu Liang, Sheryl Mathew, Mengrou Shou, Eric Nyberg, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11016 2023-10-18 cs.CL 57%

Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction

Chong Zhang, Ya Guo, Yi Tu, Huan Chen, Jinyang Tang, Huijia Zhu, Qi Zhang, Tao Gui

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Accepted as a long paper in the main conference of EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10844 2023-10-18 cs.CL cs.CR cs.LG 57%

Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, Nael Abu-Ghazaleh

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏