arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25502cs.DB

无泄漏的支出分类:一个评估框架及其在部署系统中的改变

Spend Classification Without Leakage: An Evaluation Harness and What It Changed in a Deployed System

Harshit Gupta

AI总结:

针对支出分类评估中的泄漏问题,构建了四种协议的无泄漏评估框架,发现泄漏掩盖真实差异,并揭示了文本分类器的准确率上限及部署中的实际接受率。

AI中文摘要:

将标准商品代码分配给自由文本采购行是企业支出分析的基础,但其报告的准确性无法被核查。已发表的研究使用专有数据或未记录协议下的私人样本;没有两项研究可相互比较,也没有公开基准。我们构建了一个包含四种协议的评估框架,使用来自美国两个州政府的126万条带标签采购行。企业会持续重复购买,因此随机划分会将匹配的商品文本置于两侧:60.7%和61.2%的测试行完全匹配,54.7%和60.6%逐字节匹配。在加利福尼亚州的订单上,最佳经典基线在重复文本上得分为56.0%,但在新文本上仅为31.9%,差距达24个百分点,且随分类法深度而扩大。嵌入检索将差距从23.8个百分点缩小到19.1个百分点;微调后的Transformer无法摆脱这一差距。在一个语料库上,有泄漏的评估无法区分检索与Transformer,而三种无泄漏协议显示Transformer领先2.3至3.6个百分点,因此有泄漏的划分掩盖了真实差异,而不仅仅是使两者都显得更好。我们从具有冲突代码的相同文本推导出纯文本分类器的上限:在一个语料库上商品准确率为79.4%,而在另一个上为98.4%,因此准确率无法跨数据集比较。对上限进行支出加权后,其反映金额字段的程度超过标签。在920,927行的部署中,审核者在新文本群体上接受了32.4%的建议;29.7%-35.2%的区间将基准率定在31.9%和34.8%附近。我们发布了该框架并已投入生产使用。

英文摘要:

Assigning a standard commodity code to a free-text purchase line underpins enterprise spend analytics, and its reported accuracy cannot be checked. Published results use proprietary data or private samples under undocumented protocols; no two compare and there is no public benchmark. We build a harness with four protocols over 1.26 million labelled purchase lines from two US state governments. Enterprises rebuy continuously, so random splits put matching item text on both sides: 60.7% and 61.2% of test rows, 54.7% and 60.6% byte for byte. On California orders the best classical baseline scores 56.0% on repeated text but 31.9% on novel text, a 24-point gap widening with taxonomy depth. Embedding retrieval shrinks the gap from 23.8 to 19.1 points; a fine-tuned transformer does not escape it. On one corpus leaky evaluation cannot separate retrieval from the transformer, while three leak-free protocols put the transformer 2.3 to 3.6 points ahead, so leaky splits hide real differences, not just flatter both. We derive an upper bound on text-only classifiers from identical text with conflicting codes: 79.4% commodity accuracy on one corpus against 98.4% on the other, so accuracy does not compare across datasets. Spend-weighting the bound reflects the amount field more than labels. In a deployment of 920,927 lines, reviewers accepted 32.4% of suggestions on the novel-text population; 29.7-35.2% brackets benchmark rates near 31.9% and 34.8%. We release the harness in production use.

补充信息

↑