arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2309.08760 2023-09-19 cs.CV 57%

Biased Attention: Do Vision Transformers Amplify Gender Bias More than Convolutional Neural Networks?

Abhishek Mandal, Susan Leavy, Suzanne Little

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08743 2023-09-19 cs.CV 57%

Active Learning for Fine-Grained Sketch-Based Image Retrieval

Himanshu Thakur, Soumitri Chattopadhyay

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted at BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05551 2023-09-12 cs.CV 57%

OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data

Giuseppe Cartella, Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments International Conference on Image Analysis and Processing (ICIAP) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05503 2023-09-12 cs.CL 57%

Long-Range Transformer Architectures for Document Understanding

Thibault Douzon, Stefan Duffner, Christophe Garcia, Jérémy Espinas

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments Conference: ICDAR 2023 Workshops on Document Analysis and Recognition

Journal ref Document Analysis and Recognition ICDAR 2023 Workshops pages 47 to 64

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01366 2023-09-06 cs.MM 57%

Target-Guided Composed Image Retrieval

Haokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei, Liqiang Nie

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

Journal ref ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01017 2023-09-06 cs.CV 57%

Contrastive Grouping with Transformer for Referring Image Segmentation

Jiajin Tang, Ge Zheng, Cheng Shi, Sibei Yang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00372 2023-09-04 eess.IV cs.CV 57%

On the Localization of Ultrasound Image Slices within Point Distribution Models

Lennart Bastian, Vincent Bürgin, Ha Young Kim, Alexander Baumann, Benjamin Busam, Mahdi Saleh, Nassir Navab

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments ShapeMI Workshop @ MICCAI 2023; 12 pages 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14126 2023-08-29 cs.CV 57%

Synergizing Contrastive Learning and Optimal Transport for 3D Point Cloud Domain Adaptation

Siddharth Katageri, Arkadipta De, Chaitanya Devaguptapu, VSSV Prasad, Charu Sharma, Manohar Kaul

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11752 2023-08-29 cs.CV 57%

CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose

Xu Zhang, Wen Wang, Zhe Chen, Yufei Xu, Jing Zhang, Dacheng Tao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11994 2023-08-24 cs.CV 57%

Progressive Feature Mining and External Knowledge-Assisted Text-Pedestrian Image Retrieval

Huafeng Li, Shedan Yang, Yafei Zhang, Dapeng Tao, Zhengtao Yu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08942 2023-08-24 cs.CV 57%

Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution

Zixiang Zhao, Jiangshe Zhang, Xiang Gu, Chengli Tan, Shuang Xu, Yulun Zhang, Radu Timofte, Luc Van Gool

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11485 2023-08-23 cs.CV 57%

Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features

Alberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto del Bimbo

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted in ACM Transactions on Multimedia Computing Communications and Applications (TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06309 2023-08-22 cs.CV 57%

UATVR: Uncertainty-Adaptive Text-Video Retrieval

Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Yuxin Song, Weiping Wang, Xiangbo Shu, Xiangyang Ji, Jingdong Wang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments To appear at ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09107 2023-08-21 cs.CV 57%

Hyperbolic Face Anti-Spoofing

Shuangpeng Han, Rizhao Cai, Yawen Cui, Zitong Yu, Yongjian Hu, Alex Kot

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07078 2023-08-15 cs.CV 57%

ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation

Chaohui Yu, Qiang Zhou, Zhibin Wang, Fan Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15872 2023-08-01 eess.IV cs.CV cs.LG 57%

Cross-dimensional transfer learning in medical image segmentation with deep learning

Hicham Messaoudi, Ahror Belaid, Douraied Ben Salem, Pierre-Henri Conze

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 30 pages, 12 figures, 6 tables, Accepted for publication in the Journal of Medical Image Analysis

Journal ref In Medical Image Analysis (Vol. 88, p. 102868). Elsevier BV (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.00223 2023-07-19 cs.CV 57%

Unsupervised Deep Cross-modality Spectral Hashing

Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen, Ngai-Man Cheung

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transaction on Image Processing (TIP) Add Acknowledgement

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01577 2023-07-06 cs.AI q-bio.NC 57%

Conceptual Cognitive Maps Formation with Neural Successor Networks and Word Embeddings

Paul Stoewer, Achim Schilling, Andreas Maier, Patrick Krauss

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06691 2023-06-13 cs.CV 57%

Self-Enhancement Improves Text-Image Retrieval in Foundation Visual-Language Models

Yuguang Yang, Yiming Wang, Shupeng Geng, Runqi Wang, Yimi Wang, Sheng Wu, Baochang Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06382 2023-06-05 cs.CV 57%

Differentiated Relevances Embedding for Group-based Referring Expression Comprehension

Fuhai Chen, Xuri Ge, Xiaoshuai Sun, Yue Gao, Jianzhuang Liu, Fufeng Chen, Wenjie Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19329 2023-06-01 cs.CV cs.IR cs.LG 57%

Mitigating Test-Time Bias for Fair Image Retrieval

Fanjie Kong, Shuai Yuan, Weituo Hao, Ricardo Henao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13500 2023-05-24 cs.CV 57%

Learning Emotion Representations from Verbal and Nonverbal Communication

Sitao Zhang, Yimu Pan, James Z. Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11503 2023-05-22 cs.CL 57%

A Topic-aware Summarization Framework with Different Modal Side Information

Xiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng, Qiang Yang, Qishen Zhang, Xin Gao, Xiangliang Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments SIGIR 2023, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06923 2023-05-12 cs.CV 57%

EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification

Souhail Bakkali, Ziheng Ming, Mickael Coustaty, Marçal Rusiñol

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at IJDAR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02265 2023-05-08 cs.CL 57%

A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text

Yunxin Li, Baotian Hu, Yuxin Ding, Lin Ma, Min Zhang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CL

Comments Accepted to ACL 2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02760 2023-05-05 cs.CV 57%

Multi-Modality Deep Network for JPEG Artifacts Reduction

Xuhao Jiang, Weimin Tan, Qing Lin, Chenxi Ma, Bo Yan, Liquan Shen

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 18 pages, 17 figures, accepted by IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00131 2023-05-02 cs.CV 57%

Regularizing Self-training for Unsupervised Domain Adaptation via Structural Constraints

Rajshekhar Das, Jonathan Francis, Sanket Vaibhav Mehta, Jean Oh, Emma Strubell, Jose Moura

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09520 2023-05-02 cs.CV 57%

Using Language to Extend to Unseen Domains

Lisa Dunlap, Clara Mohri, Devin Guillory, Han Zhang, Trevor Darrell, Joseph E. Gonzalez, Aditi Raghunathan, Anja Rohrbach

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04662 2023-04-24 cs.LG cs.AI 57%

Cognitively Inspired Learning of Incremental Drifting Concepts

Mohammad Rostami, Aram Galstyan

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 2023 International Joint Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07747 2023-04-18 cs.CV 57%

Language Guided Local Infiltration for Interactive Image Retrieval

Fuxiang Huang, Lei Zhang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments 10 pages, 9 figures, 4 tables, IEEE/CVF Conference on Computer Vision and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏