超越精确匹配:评估方法如何在基于大语言模型的产品属性提取中主导模型选择
Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
AI总结:
研究基于大语言模型的产品属性提取中评估方法的影响,通过在MAVE基准上对比两种模型的四种提示策略,用精确和模糊匹配评估预测并审计标签,发现评估方法产生的方差远超模型和提示策略选择,且基准有一定噪声率,得出评估方法和数据质量主导模型选择等结论。
AI中文摘要:
大语言模型已成为电子商务流程中结构化产品属性提取的默认选择,但从业者报告不同模型、数据集和提示策略的性能差异很大。本文进行了一项对照实证研究,在MAVE基准上比较了两种生产级大语言模型(GPT-4o-mini和Gemini 2.5 Flash)的四种提示策略——零样本、少样本、模式引导和定义增强。使用精确和模糊字符串匹配评估6400个属性级预测,并对真实标签进行严格噪声审计。发现评估方法产生的方差比模型选择大约大23倍,比提示策略选择大5倍。还确定MAVE基准对现代大语言模型输出的真实噪声率为23.2%。配对排列检验证实协议间F1差距非常显著(p<0.0001),协议间Cohen's kappa为0.769表明有实质性一致性。结论是对于生产属性提取管道,评估方法和数据质量主导模型选择和提示工程的影响。
英文摘要:
Large language models (LLMs) have become a default choice for structured product attribute extraction in e-commerce pipelines, with practitioners reporting widely varying performance across models, datasets, and prompting strategies. This paper presents a controlled empirical study comparing four prompting strategies -- zero-shot, few-shot, schema-guided, and definition-augmented -- across two production-grade LLMs (GPT-4o-mini and Gemini 2.5 Flash) on the MAVE benchmark. We evaluate 6,400 attribute-level predictions using both exact and fuzzy string matching, and conduct a rigorous noise audit of the ground truth labels. We formally decompose F1 variance across four experimental factors and find that evaluation methodology produces variance approximately 23 times larger than model choice and 5 times larger than prompting strategy choice. We further establish that the MAVE benchmark exhibits a 23.2% ground truth noise rate against modern LLM outputs. Paired permutation tests (B=10,000) confirm that the inter-protocol F1 gap is highly significant (p<0.0001) and Cohen's kappa of 0.769 between protocols indicates substantial agreement. We conclude that for production attribute extraction pipelines, evaluation methodology and data quality dominate the impact of model selection and prompt engineering.