Are Multimodal Large Language Models Good Annotators for Image Tagging?
多模态大语言模型在图像标签任务中是好的标注者吗?
机构 * RIKEN Center for Advanced Intelligence Project, Japan(日本先进人工智能研究中心) ; The University of Tokyo, Japan(东京大学) ; Southeast University, China(东南大学) ; China University of Mining(中国矿业大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出TagLLM框架,通过候选生成和标签歧义消除方法,有效缩小多模态大语言模型生成标注与人工标注之间的差距,提升图像标签任务的效率和性能。