LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
LLaVE: 大规模语言和视觉嵌入模型与基于难度加权的对比学习
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) ; Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)
AI总结 LLaVE通过基于难度加权的对比学习提升多模态嵌入模型性能,实现SOTA表现和强泛化能力。
Comments Accepted by Findings of EMNLP 2025