arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00118cs.LG

99%准确率的隐藏代价:Telco客户流失基准的可信度审计

The Hidden Costs of 99% Accuracy: A Trustworthiness Audit of the Telco Customer Churn Benchmark

Soumyadeep Roy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计IBM Telco客户流失基准,揭示SMOTE泄漏、特征冗余、校准失效和成本阈值偏差等可信度问题,提出四部分报告清单以提升模型评估可靠性。

中文摘要 AI 辅助

针对IBM Telco客户流失基准(n=7,043)的客户流失预测,常规报告测试准确率超过95%,其中被引用最多的已发表研究报告准确率达99.01%。我们对该基准进行了四项可信度失效审计,这些失效在主导文献的以准确率和F1为中心的报告中不可见。第一,分割前SMOTE在十个分类器和十五个随机种子下将流失类F1提高了13.1个百分点(每个分类器的Wilcoxon p<10^-4);相同的泄漏流水线顺序配合类别加权则不产生提升,从而将效应隔离为SMOTE的几何构造所致。我们直接测量了机制:约36%的合成训练点是测试集实例的最近邻插值。第二,TotalCharges字段近似由tenure乘以MonthlyCharges决定(R2=0.999);移除它使准确率变化小于0.2个百分点,但TreeSHAP将其在平均绝对归因中排第九——这一模式实质性地破坏了基于SHAP的解释。我们提出R2>0.95的建模前诊断。第三,在15种子校准审计中,等渗回归是最强的默认方法;温度缩放对类别加权树集成(其预测概率分布呈双峰)失效。第四,成本最优决策阈值(在50美元留存优惠和24个月CLV代理下)比F1最优阈值低约5-10倍,每1000名客户节省约77,000美元。我们在伊朗电信流失(域内)和银行客户流失(跨域)上复现F1和F2:F1可泛化;F2仅在电信域内泛化。我们将这些发现综合为四部分报告清单——流水线披露、冗余诊断、校准审计和成本敏感阈值——并发布可复现的实现。

英文摘要

Customer churn prediction on the IBM Telco Customer Churn benchmark (n = 7,043) routinely reports test accuracies above 95%, with the most cited published study reporting 99.01%. We audit this benchmark for four trustworthiness failures invisible to the accuracy- and F1-centred reporting that dominates the literature. First, pre-split SMOTE inflates churn-class F1 by 13.1 percentage points across ten classifiers and fifteen seeds (Wilcoxon p < 10^-4 per classifier); the same leaky pipeline ordering paired with class weighting yields no inflation, isolating the effect to SMOTE's geometric construction. We measure the mechanism directly: approximately 36% of synthetic training points are nearest-neighbour interpolations of test-set instances. Second, the TotalCharges field is approximately determined by tenure multiplied by MonthlyCharges (R2 = 0.999); removing it changes accuracy by less than 0.2 percentage points, yet TreeSHAP ranks it ninth in mean absolute attribution - a pattern that materially corrupts SHAP-based interpretation. We propose an R2 > 0.95 pre-modelling diagnostic. Third, in a 15-seed calibration audit, isotonic regression is the strongest default; temperature scaling fails on class-weighted tree ensembles whose predicted-probability distribution is bimodal. Fourth, the cost-optimal decision threshold (under a 50 USD retention offer and 24-month CLV proxy) is approximately 5-10 times lower than the F1-optimal threshold, saving approximately 77,000 USD per 1,000 customers. We replicate F1 and F2 on Iranian Telecom Churn (within domain) and Bank Customer Churn (across domain): F1 generalises; F2 generalises only within telecom. We synthesise these findings into a four-component reporting checklist - pipeline disclosure, redundancy diagnostic, calibration audit, and cost-sensitive thresholds - and release a reproducible implementation.

↑