HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments 10 pages, 6 figures
Journal ref Proceedings of the 31st ACM International Conference on Multimedia. 2023: 4768-4777