arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2210.10638 2022-11-07 cs.IR cs.AI cs.LG 57%

Digital Human Interactive Recommendation Decision-Making Based on Reinforcement Learning

Xiong Junwu, Xiaoyun Feng, YunZhou Shi, James Zhang, Zhongzhou Zhao, Wei Zhou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 1 table, the paper has been accepted and this is the final camera-ready for NeurIPS 2022 Workshop on Human in the Loop Learning, https://neurips-hill.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13430 2022-11-01 cs.CV cs.LG 57%

UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Janghyeon Lee, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim, Seung Hwan Kim, Honglak Lee, Junmo Kim

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Neural Information Processing Systems (NeurIPS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13848 2022-11-01 cs.AI cs.RO 57%

ProspectNet: Weighted Conditional Attention for Future Interaction Modeling in Behavior Prediction

Yutian Pang, Zehua Guo, Binnan Zhuang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13591 2022-10-28 cs.CV 57%

Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision

Tzu-Jui Julius Wang, Jorma Laaksonen, Tomas Langer, Heikki Arponen, Tom E. Bishop

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to WACV'23. Please find supplementary material at https://drive.google.com/file/d/1SmCBGsUgkYLAhmK83RZqY03bq4j3214p/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14427 2022-10-27 cs.CL 57%

ReSel: N-ary Relation Extraction from Scientific Text and Tables by Learning to Retrieve and Select

Yuchen Zhuang, Yinghao Li, Jerry Junyang Cheung, Yue Yu, Yingjun Mou, Xiang Chen, Le Song, Chao Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11204 2022-10-21 cs.CV 57%

PalGAN: Image Colorization with Palette Generative Adversarial Networks

Yi Wang, Menghan Xia, Lu Qi, Jing Shao, Yu Qiao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05507 2022-10-07 cs.CV 57%

TextMatcher: Cross-Attentional Neural Network to Compare Image and Text

Valentina Arrigoni, Luisa Repele, Dario Marino Saccavino

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at the 25th International Conference on Discovery Science 2022, 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04013 2022-10-05 eess.IV cs.CV 57%

Mutual Contrastive Low-rank Learning to Disentangle Whole Slide Image Representations for Glioma Grading

Lipei Zhang, Yiran Wei, Ying Fu, Stephen Price, Carola-Bibiane Schönlieb, Chao Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 14 pages, 4 figures, 3 tables; This paper was accepted by BMVC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14644 2022-09-30 cs.LG cs.CV 57%

Increasing Model Generalizability for Unsupervised Domain Adaptation

Mohammad Rostami

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Presented 2022 Conference on Lifelong Learning Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07285 2022-09-23 cs.CV 57%

X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval

Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, Rongrong Ji

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 13 pages, 6 figures, ACMMM22

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10304 2022-09-22 cs.CV 57%

I2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification

Muhammad Ferjad Naeem, Yongqin Xian, Luc Van Gool, Federico Tombari

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09433 2022-09-21 cs.CL 57%

Non-Linguistic Supervision for Contrastive Learning of Sentence Embeddings

Yiren Jian, Chongyang Gao, Soroush Vosoughi

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments Accepted to NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07996 2022-09-19 cs.RO cs.AI 57%

SoLo T-DIRL: Socially-Aware Dynamic Local Planner based on Trajectory-Ranked Deep Inverse Reinforcement Learning

Yifan Xu, Theodor Chakhachiro, Tribhi Kathuria, Maani Ghaffari

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01517 2022-09-07 cs.CV 57%

Joint Prediction of Meningioma Grade and Brain Invasion via Task-Aware Contrastive Learning

Tianling Liu, Wennan Liu, Lequan Yu, Liang Wan, Tong Han, Lei Zhu

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by MICCAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01152 2022-09-05 cs.CV 57%

A Novel Approach for Pill-Prescription Matching with GNN Assistance and Contrastive Learning

Trung Thanh Nguyen, Hoang Dang Nguyen, Thanh Hung Nguyen, Huy Hieu Pham, Ichiro Ide, Phi Le Nguyen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted for publication and presentation at the 19th Pacific Rim International Conference on Artificial Intelligence (PRICAI 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12238 2022-08-26 cs.HC cs.AI cs.LG 57%

Supervised Contrastive Learning for Affect Modelling

Kosmas Pinitas, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments This paper was accepted to ICMI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15087 2022-08-09 cs.LG cs.AI cs.SI 57%

MOOMIN: Deep Molecular Omics Network for Anti-Cancer Drug Combination Therapy

Benedek Rozemberczki, Anna Gogleva, Sebastian Nilsson, Gavin Edwards, Andriy Nikolov, Eliseo Papa

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01209 2022-08-02 cs.CV 57%

Learning Structural Representations for Recipe Generation and Food Retrieval

Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted at IEEE Transactions on Pattern Analysis and Machine Intelligence. arXiv admin note: substantial text overlap with arXiv:2009.00944

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13038 2022-07-27 cs.CV 57%

Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models

Robin Rombach, Andreas Blattmann, Björn Ommer

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09703 2022-07-21 cs.CV 57%

MaCLR: Motion-aware Contrastive Learning of Representations for Videos

Fanyi Xiao, Joseph Tighe, Davide Modolo

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.14204 2022-07-19 cs.CV 57%

Personalized Trajectory Prediction via Distribution Discrimination

Guangyi Chen, Junlong Li, Nuoxing Zhou, Liangliang Ren, Jiwen Lu

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2021. Code: https://github.com/CHENGY12/DisDis

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07869 2022-07-06 eess.SP cs.AI cs.LG 57%

Near out-of-distribution detection for low-resolution radar micro-Doppler signatures

Martin Bauw, Santiago Velasco-Forero, Jesus Angulo, Claude Adnet, Olivier Airiau

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00750 2022-07-05 cs.AI cs.IR 57%

GUIM -- General User and Item Embedding with Mixture of Representation in E-commerce

Chao Yang, Ru He, Fangquan Lin, Suoyuan Song, Jingqiao Zhang, Cheng Yang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12071 2022-06-27 cs.CV 57%

Contrastive Learning of Features between Images and LiDAR

Peng Jiang, Srikanth Saripalli

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments accepted in CASE2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10879 2022-06-23 cs.CV 57%

Symmetric Network with Spatial Relationship Modeling for Natural Language-based Vehicle Retrieval

Chuyang Zhao, Haobo Chen, Wenyuan Zhang, Junru Chen, Sipeng Zhang, Yadong Li, Boxun Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 8 pages, 3 figures, publised to CVPRW

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 3226-3233

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08172 2022-06-17 cs.CV 57%

RefCrowd: Grounding the Target in Crowd with Referring Expressions

Heqian Qiu, Hongliang Li, Taijin Zhao, Lanxiao Wang, Qingbo Wu, Fanman Meng

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04789 2022-06-08 cs.CV 57%

Look, Cast and Mold: Learning 3D Shape Manifold from Single-view Synthetic Data

Qianyu Feng, Yawei Luo, Keyang Luo, Yi Yang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments this work is no longer under development

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03891 2022-05-10 cs.CV 57%

Cross-lingual Adaptation for Recipe Retrieval with Mixup

Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Wing-Kwong Chan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICMR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08209 2022-05-10 cs.CV 57%

OMG: Observe Multiple Granularities for Natural Language-Based Vehicle Retrieval

Yunhao Du, Binyu Zhang, Xiangning Ruan, Fei Su, Zhicheng Zhao, Hong Chen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments CVPR 2022 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09869 2022-05-09 cs.CV 57%

Progressive Domain-Independent Feature Decomposition Network for Zero-Shot Sketch-Based Image Retrieval

Xinxun Xu, Muli Yang, Yanhua Yang, Hao Wang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏