arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RA-CoA:基于检索增强属性链的无训练时尚图像描述生成

RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes

Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra

arXiv 2609.14100首次发表:更新:

发表机构

Mohamed Bin Zayed University of Artificial Intelligence; Indian Institute of Technology Jodhpur(穆罕默德·本·扎耶德人工智能大学; 印度理工学院焦特布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对时尚图像描述中通用VLM细粒度属性精度不足的问题,提出无训练框架RA-CoA,通过从知识库检索属性集并进行属性级推理,显著提升描述质量,METEOR平均提升26.3%。

AI 中文摘要

时尚图像描述生成(FIC)在提升电子商务平台的用户体验和商品搜索方面发挥着至关重要的作用。与自然场景图像描述不同,FIC需要细粒度的视觉推理和领域特定术语知识,以捕捉诸如领口和闭合类型、图形图案以及连衣裙轮廓等细微属性。此外,随着时尚库存随着新趋势、新风格和频繁出现的新词汇而快速演变,开发无训练的描述生成解决方案对于可扩展性和现实世界适应性至关重要。指令微调的视觉语言模型(VLM)凭借其强大的零样本能力和自然语言流畅性,为时尚图像描述生成提供了有前景的解决方案。然而,这些通用模型往往缺乏属性级别的覆盖和精度,并且容易产生幻觉或错误识别细粒度的时尚细节,使其不太适用于产品目录或个性化推荐等高保真应用。为解决这一问题,我们提出了RA-CoA(检索增强属性链),一种新颖的、无训练的框架,将时尚图像描述生成分解为两个可解释的阶段:(i)从产品知识库中检索相关属性集,以及(ii)进行属性级推理以生成最终描述。RA-CoA是一种模型无关的方法,可与冻结的VLM配合使用,在无需微调的情况下提高产品描述中细粒度属性的精度。在不同VLM模型家族和不同提示范式下进行的广泛评估表明,RA-CoA显著提高了描述质量,与零样本描述相比,METEOR分数平均提升了26.3%。我们公开了我们的代码。

英文摘要

Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image captioning, FIC requires fine-grained visual reasoning and knowledge of domain-specific terminology to capture subtle attributes such as neckline and closure types, graphic patterns, and dress silhouettes. Moreover, as fashion inventories evolve rapidly with new trends, styles, and frequently emerging vocabulary, developing training-free captioning solution becomes essential for scalability and real-world adaptability. Instruction-tuned vision-language models (VLMs) offer a promising solution to fashion image captioning dueto their strong zero-shot capabilities and natural language fluency. However, these general-purpose models often lack attribute-level coverage and precision, and tend to hallucinate or misidentify fine-grained fashion details, making them less suitable for high-fidelity applications like product cataloging or personalized recommendations. To address this, we propose RA-CoA (Retrieval-Augmented Chain-of-Attributes), a novel, training-free framework that disentangles fashion image captioning into two interpretable stages: (i) retrieval of relevant attribute sets from a product knowledge base, and (ii) attribute-level reasoning to generate the final caption. RA-CoA is a model-agnostic approach that works with frozen VLMs to improve fine-grained attribute precision in product captions without the need for fine-tuning. Extensive evaluations across diverse VLM model families under different prompting paradigms demonstrate that RA-CoA significantly improves caption quality, achieving an average gain of 26.3% METEOR score over zero-shot captioning. We make our code publicly available.

CommentsAccepted in TMLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑