arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13279cs.CVeess.IV

MODA通用属性套件:面向时尚属性提取的四轨评估基准

The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction

Arkid Mitra

首次发表
浏览论文内容

中文总结 AI 辅助

针对时尚属性提取评估不一致的问题,提出四轨基准MODA通用属性套件,通过分离问题、严格协议和逐轨验证,提供可靠评估并公开资源。

中文摘要 AI 辅助

时尚属性提取的评估方式不一致:结果以单一聚合数字报告,涵盖存在不同问题的图像类型;图像中不可见的字段被计为普通负例;数据集之间词汇不匹配的影响虽被承认但未加以衡量。我们发布了MODA通用属性套件,这是一个通过构造将这些问题分开的四轨基准。每条轨道(局部服装裁剪图、目录产品图像、适用性感知的全身照片和产品文本)都有自己固定的测试集、输入契约、评估指标和泄漏单元,且各轨道从不进行平均。该协议要求标签盲预测,在任何标签被打开之前对预测文件进行SHA-256承诺,使用故障关闭评分器,并在每条轨道的自然泄漏单元上进行10,000样本配对聚类自助法;晋升要求在每条轨道上获得正区间,而非有利的平均值。我们发布了评分器、分割构建器、我们自己的预测文件及其哈希(包括我们失败的运行),以及四条轨道中三条的MODA_NER(V)模型检查点。文本模型未分发;其基准已发布。我们报告了基线结果、六项未改善基线的干预措施以及证据的局限性。

英文摘要

Fashion attribute extraction is evaluated inconsistently: results are reported as single aggregate numbers across image types that pose different problems, fields that are not visible in an image are scored as ordinary negatives, and the effect of vocabulary mismatch between datasets is acknowledged but not measured. We release the MODA General Attribute Suite, a four-track benchmark that keeps these problems separate by construction. Each track (localized garment crops, catalogue product images, applicability-aware full-body photographs, and product text) carries its own frozen test set, input contract, metric, and leakage unit, and the tracks are never averaged. The protocol requires label-blind prediction, SHA-256 commitment of prediction files before any label is opened, a fail-closed scorer, and a 10,000-sample paired cluster bootstrap at each track's natural leakage unit; promotion requires a positive interval on every track rather than a favourable mean. We release the scorers, the split builders, our own prediction files with their hashes for every image-track system in the main results and for runs we lose, and the MODA_NER(V) model checkpoints for three of the four tracks. The text model is not distributed; its benchmark is. We report baseline results, six interventions that did not improve them, and the limitations of the evidence. On the garment-crop track we also measure how far the headline depends on the benchmark's own design: the order of the two strongest systems depends on whether fields are micro- or macro-averaged, two thirds of the cells the protocol judges carry no annotation, and disabling the applicability decision cuts attribute micro-F1 from 0.63 to 0.27, so most of that score rests on absence decisions the source labels cannot adjudicate. This version also corrects one figure in an external comparison that an inference preprocessing defect had affected.

发表机构

  • Hopit AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑