AI 中文总结
本研究对比测试LLM与基于规则方法在两项电商数据质量任务中的表现,发现LLM在需背景知识的任务中更具优势,且判断高度一致,小型样本提示评估可能存在误导。
AI 中文摘要
大语言模型(LLM)已越来越多地被用于自动发现数据质量问题,但我们对这些判断的实际一致性知之甚少。本研究在零样本和少样本提示设置下,针对实体匹配和品牌错误标注两项电商数据质量任务,将某LLM与基于规则的基线方法、经人工验证的真实标注进行对比测试。在实体匹配任务中,使用Abt Buy基准数据集(含2194个标注对)时,简单基于规则的基线方法(F1值为0.950)的表现与LLM零样本提示(F1值为0.948)大致相当;此外,在小型验证样本上看似有效的少样本提示修订,却将全规模性能降至F1值0.914,这表明小型样本的提示评估可能具有误导性。在品牌错误标注检测任务中,使用注入了合成标注错误的500条亚马逊商品列表时,LLM的表现明显优于朴素基于规则的基线方法(F1值为0.833对比0.721),因为它能利用简单规则无法获取的品牌-商品关系背景知识。针对重复运行的一致性测试(200个对,温度设置为0.7时运行5次)显示,模型平均99.7%的时间与自身判断一致,99%的对在全部5次运行中给出相同答案;对这些运行结果采用多数投票仅将F1值提高0.005,但推理成本增加了4倍。这些结果表明,使用LLM相比传统方法的价值高度依赖于任务:当已存在强词汇信号时,LLM几乎无优势;但当任务需要背景知识时,LLM具有明显优势,同时在重复查询中保持高度一致性。
英文摘要
LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching and brand mislabeling, against rule based baselines and human verified ground truth, under both zero-shot and few-shot prompting. On entity matching while using the Abt Buy benchmark (2,194 labeled pairs), a simple rule based baseline (F1=0.950) performed about as well as LLM zero shot prompting (F1=0.948). Moreover, a few-shot prompt revision that looked effective on a small validation sample reduced full-scale performance to F1=0.914. This showed that small sample prompt evaluation can be misleading. On brand mislabeling detection, using 500 Amazon product listings with synthetically injected labeling errors, the LLM clearly outperformed a naive rule based baseline (F1=0.833 vs 0.721), because it could draw on background knowledge of brand product relationships that a simple rule could not access. Testing consistency across repeated runs (200 pairs, 5 runs at temperature 0.7) showed the model agreeing with itself 99.7% of the time on average, with 99% of pairs giving identical answers across all 5 runs. Using majority voting across these runs only improved F1 by 0.005, at 5 times the inference cost. These results suggest that the value of using an LLM over traditional methods depends heavily on the task. LLMs offer little advantage when strong lexical signals already exist, but a clear advantage when the task requires background knowledge, all while remaining highly consistent across repeated queries.
Comments6 pages, 4 figures