arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynthAVE:通过大语言模型领域验证实现电子商务的可扩展合成标注

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

Andrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan, Victor Martinez, Marcello Federico

arXiv 2607.07469首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对电子商务属性提取微调大语言模型需大量标注数据、人工标注成本高的问题,提出SynthAVE基准及多LLM领域框架,通过多数投票验证合成标注,实现经济高效且质量与人工审核相当的大规模属性值提取验证。

AI 中文摘要

针对电子商务属性提取对大语言模型(LLMs)进行微调,需要涵盖数千种产品类型、属性以及多种语言的有代表性的标注数据。这种组合规模意味着数百万条标注,人工标注成本过高。近期工作虽展示了用LLMs生成合成标注,但工业规模部署需集成质量控制机制。我们提出SynthAVE,这是一个大规模的人工验证基准,涵盖229种产品类型、792个属性以及4种语言(西班牙语、法语、意大利语、德语)的12726种产品的属性值提取。为大规模验证合成标注,我们引入多LLM领域框架,样本由21种评判配置独立评估,最终标注通过多数投票确定。多数投票集合与人类专家的一致性达到Cohen's κ = 0.92(95.2%的一致性),个体评判者之间也有较高的模型间一致性(Fleiss' κ = 0.76)。这表明不同个体判断的多样模型能聚合为高度可靠的预测,实现大规模的经济高效验证,同时保持与人工审核的质量相当。

英文摘要

Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs, deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale benchmark for attribute-value verification spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where each sample is evaluated by 21 judge configurations (7 model families $\times$ 3 prompts), with final labels determined via majority voting; disagreements with the synthetic label are expert-adjudicated and agreement cases are audited on a stratified sample. The majority vote ensemble agrees with expert labels at Cohen's $κ= 0.92$ (95.0% agreement), while individual judges agree with one another only moderately (Fleiss' $κ= 0.76$)--by design, since we select for diversity. This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while concentrating expert effort on the cases where it changes the label. We estimate the resulting label quality at 97.9%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑