arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向电信客户流失预测的可解释人工智能:一种CRM集成框架

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

Sandeep Gaddamwar

arXiv 2608.26151首次发表:更新:

AI 中文总结

本文针对电信客户流失预测中模型不透明导致的CRM集成缺口,测试四种分类器并结合SHAP、LIME提供可解释性,提出四层CRM集成架构,预计可降低流失率3.3-5.3个百分点、节省19.9万-31.9万美元。

AI 中文摘要

用户流失是电信运营商面临的代价高昂且持续存在的挑战,成熟市场的月流失率约为1.9%,每年造成数十亿美元的收入损失。预测模型可准确标记高风险客户,但高性能的集成模型与非线性架构因不透明而常被排除在一线CRM工作流程之外:客户留存专员仅靠概率评分无法设计个性化干预措施,而不清楚用户为何处于风险中。本文解决这一缺口。我们在IBM Telco Customer Churn基准数据集(7043条记录、19个特征、26.5%的流失率,仅在训练分区通过SMOTE将其平衡至50%)上对四种分类器进行基准测试:Logistic Regression(逻辑回归)、Random Forest(随机森林)、XGBoost和LightGBM。其中Logistic Regression取得最高AUC-ROC(0.8411),LightGBM取得最高准确率(78.42%);四种模型的AUC均在0.011的区间内(0.831-0.841),5折交叉验证证实领先模型表现基本持平。解释分为两个粒度:全局SHAP排名确定 tenure(使用时长)、total charges(总费用)和month-to-month contract(月度合约)是主导流失信号;实例级SHAP与LIME分解则揭示每个预测背后的驱动因素。基于这些输出,我们提出一种四层CRM集成架构,将风险评分与归因向量转换为分层细分,将核心特征映射为结构化留存行动模板,并将活动结果导入重训练反馈循环。针对最高风险的五分之一用户,预计可将整体流失率降低3.3-5.3个百分点,每个活动周期预计节省19.9万-31.9万美元。

英文摘要

Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature markets eroding billions in revenue annually. Predictive models can flag at-risk customers accurately, yet they are routinely excluded from frontline CRM workflows because high-performing ensemble and non-linear architectures are opaque: a retention specialist cannot design a personalised intervention from a probability score alone, without knowing why a subscriber is at risk. This paper addresses that gap. We benchmark four classifiers--Logistic Regression, Random Forest, XGBoost, and LightGBM--on the IBM Telco Customer Churn benchmark (7,043 records; 19 features; 26.5% churn, balanced to 50% via SMOTE on the training partition only). Logistic Regression attains the strongest AUC-ROC (0.8411) and LightGBM the highest accuracy (78.42%); all four fall within a 0.011 AUC band (0.831--0.841), and 5-fold cross-validation confirms the leading models are effectively tied. Explanations are delivered at two granularities: a global SHAP ranking identifying tenure, total charges, and month-to-month contract as the dominant churn signals, and instance-level SHAP and LIME decompositions that expose the drivers behind each prediction. Building on these outputs, we introduce a four-layer CRM integration architecture that converts risk scores and attribution vectors into tiered segmentation, maps top features to structured retention-action templates, and routes campaign outcomes into a retraining feedback loop. Targeting the highest-risk quintile is projected to cut overall churn by 3.3--5.3 percentage points, preserving an estimated $199K--$319K per campaign cycle.

Comments10 pages, 7 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑