arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重标注图像-文本监督分布的匹配预算审计框架

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu

arXiv 2610.00952首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对重标注图像-文本监督分布难以审计的问题,提出匹配预算审计框架,在固定文本预算下报告五轴概况,并在多源语料库上验证其有效性,同时发布审计语料库与工件。

AI 中文摘要

重标注的图像-文本语料库现已成为文本到图像(T2I)训练的标准,视觉-语言模型(VLM)标注器用密集描述替代了稀疏的替代文本。重标注语料库是由有文档记录的标注策略($\pi$)、标注器($V_c$)和源语料库($C$)所诱导的监督分布。长度相关的代理指标无法捕捉标注语域伪影,且下游T2I基准将语料库与训练选择纠缠在一起,因此该分布在语料库规模上难以审计。我们引入了一个可复用的匹配预算审计框架,用于重标注监督分布 $D_{\pi,V_c,C}$:在固定的文本预算 $B = 64$ 下,它报告一个五轴概况,涵盖提示侧覆盖、图像条件忠实度和标注表面健康度,并以声称的可控基本单元(CBU)作为通用声称单元。我们在五个公共源语料库上对七对比较实例化了该框架。在四个跨语料库配对中,发布的表面在Qwen和Gemma评判器下将每个标注的支持CBU提高了 $+3.39$ 到 $+6.36$,并且在CC12M上,同一框架揭示了一个长-密边界,该边界在两个评判器和四个预算下均保持一致。我们发布了经过审计的多源重标注语料库(约490M)以及审计工件包。

英文摘要

Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($π$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{π,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.

Commentsinitial commit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑