Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
跨模态全模式细粒度对齐用于文本到图像人物检索
机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究所,电子科学与技术大学) ; University of Electronic Science and Technology of China(电子科学与技术大学) ; Sichuan Artificial Intelligence Research Institute(四川人工智能研究院)
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV
AI总结 本文提出FMFA框架,通过显式细粒度对齐和隐式关系推理实现文本到图像人物检索的高精度匹配。
Comments accepted by ACM Transactions on Multimedia Computing Communications and Applications in December 2025
Journal ref ACM Transactions on Multimedia Computing Communications and Applications, 2026