Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models
让我们进行BiFTA:用于视觉语言模型中细粒度文本-视觉对齐的双 refinement
机构 * School of Computing and Information Systems(计算机与信息系统学院) ; The University of Melbourne(墨尔本大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.AI
AI总结 BiFTA通过视图和描述细化方法提升视觉语言模型中细粒度文本-视觉对齐的零样本性能。
Comments 25 pages