arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据驱动的电信营销优化:基于机器学习的客户流失预测与客户细分框架

Data-Driven Telecom Marketing Optimization: A Machine Learning-Based Churn Prediction and Customer Segmentation Framework

Nada Ali, Lina Ahmed, Tahani Abdalla Attia Gasmalla

arXiv 2607.10260首次发表:更新:

AI 中文总结

研究电信公司客户流失问题,提出集成机器学习流失预测、客户细分及定制营销和ROI策略的数据驱动框架,用特定数据集训练调整模型,划分客户聚类并设计策略,经实验验证该框架比单独流失预测能产生更优营销决策。

AI 中文摘要

客户流失是电信公司面临的重大挑战,直接侵蚀收入和长期客户关系。传统的留存计划依赖通用而非个性化激励,缺乏在高风险客户流失前识别他们的精准度。本文提出一个数据驱动的营销优化框架,集成基于机器学习的客户流失预测、结合流失风险与客户价值的客户细分,以及针对特定细分的定制营销和投资回报率(ROI)策略。使用含7043个客户和21个特征的IBM电信客户流失数据集,通过随机搜索、分层5折交叉验证、类别加权和F1分数驱动的决策阈值优化训练和调整了三个梯度提升集成模型(XGBoost、LightGBM和CatBoost),以应对73.4%对26.6%的类别不平衡。选择CatBoost作为部署模型,在留出的测试集上实现了77.68%的准确率、0.6366的F1分数、0.6553的PR AUC和0.8403的ROC AUC。通过K均值聚类对客户进行划分,经肘部方法验证并使用主成分分析可视化,分为高、中、低价值细分,与流失风险标签交叉制表以定义四个可操作的聚类。为每个聚类设计了特定细分的留存、追加销售和参与策略,一个理论ROI和CLV框架量化了所提议干预措施的财务影响。该流程在交互式Streamlit Web应用程序中运行,允许营销团队上传数据、按细分过滤、通过SHAP可视化流失驱动因素并下载自动细分报告。结果证实,将预测性客户流失建模与价值感知细分相结合,比单独进行客户流失预测能产生更具可操作性和盈利性的营销决策。

英文摘要

Customer churn is a major challenge for telecommunication companies, directly eroding revenue and long term customer relationships. Traditional retention programs rely on generic, not personalized incentives and lack the precision to identify high risk customers before they leave. This paper presents a data driven marketing optimization framework integrating machine learning based churn prediction, customer segmentation combining churn risk with customer value, and tailored, segment specific marketing and Return on Investment ROI strategies. Using the IBM Telco Customer Churn dataset with 7043 customers and 21 features, three gradient boosting ensembles, XGBoost, LightGBM, and CatBoost, were trained and tuned via randomized search with stratified 5 fold cross validation, class weighting, and F1 score driven decision threshold optimization to counter a class imbalance of 73.4% versus 26.6%. CatBoost was selected as the deployment model, achieving 77.68% accuracy, an F1 score of 0.6366, a PR AUC of 0.6553, and a ROC AUC of 0.8403 on the held out test set. Customers were partitioned with K Means clustering, validated via the Elbow method and visualized with Principal Component Analysis, into High, Medium, and Low Value segments, cross tabulated against churn risk labels to define four actionable clusters. Segment specific retention, upsell, and engagement strategies were designed for each cluster, and a theoretical ROI and CLV framework quantifies the financial impact of the proposed interventions. The pipeline was operationalized in an interactive Streamlit web application allowing marketing teams to upload data, filter by segment, visualize churn drivers via SHAP, and download automated segment reports. Results confirm that combining predictive churn modeling with value aware segmentation yields more actionable and profitable marketing decisions than churn prediction alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑