arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用两个置信门控本地LLM裁判策划商户匹配训练数据

Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges

Donghao Huang, Jinling Pei, Zhaoxia Wang

arXiv 2609.33878首次发表:更新:

发表机构

Mastercard; Singapore Management University(万事达卡; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出使用两个本地大语言模型裁判的置信门控一致性来策划商户匹配训练数据,通过分离阈值提高伪标签纯度,实验表明该方法在覆盖率和纯度上优于单一模型。

AI 中文摘要

商户匹配将嘈杂的支付描述符解析为检索到的商户实体,或返回无匹配结果。在策划训练标签时,一个关键挑战是区分教师弃权(不执行)与不存在可接受实体的证据:错误的无匹配标签会污染伪标签数据,而保守的标注会降低覆盖率。我们研究了两个本地大语言模型裁判之间的一致性是否能提高伪标签的可靠性。仅当裁判一致时保留标签,对选择和弃权(不执行)分别设置有序阈值,确保正负标签集不相交。对2,000条专家标注查询的回顾性重放表明,较高的选择阈值可提高正标签纯度,而较高的弃权(不执行)阈值会增加错误的无匹配标签。在阈值(0.86, 0.80)下,Muse Glimmer 30B和Gemma 4 31B共同标注了1,633条查询(覆盖率81.7%),纯度为96.88%;正负纯度分别为99.47%和93.38%。这比相同阈值下的任一组成模型高出两个百分点以上,且覆盖率更低。分裂半检查发现阈值选择乐观度仅为0.14个百分点。对称阈值0.86增加了40个错误的无匹配标签,而即使没有置信阈值,仍有46个错误弃权(不执行)持续存在。在五次匹配的模型内比较中,更高的推理努力并未带来明确的F0.5增益,且将中位延迟增加了1.8至5.0倍。这些结果促使对正负伪标签分别设置阈值和进行审计。该研究确立了标签纯度,而非学生效用;新鲜数据策划和学生微调仍是展示下游价值的必要条件。

英文摘要

Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.

Comments7 pages, 2 figures, 5 tables, Accepted for publication in 2026 IEEE International Conference on Data Mining Workshops (ICDMW), SENTIRE 2026 Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑