arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARRE:用于可解释流失处方的反事实行动检索与原因评估

CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription

Minjoo Kim, Sangjin Park, Seung Hwan Cho

arXiv 2609.09766首次发表:更新:

发表机构

Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CARRE提出三阶段框架,结合检索增强生成、成本感知反事实评分和LLM推理,在IBM流失数据集上显著优于SHAP基线,实现可解释的流失处方。

AI 中文摘要

流失模型通常能识别高风险客户,但无法指明应考虑哪种可行的保留行动,或为何该行动是合适的。我们提出了CARRE(反事实行动检索与原因评估),这是一个三阶段框架,结合了检索增强的候选生成、成本感知的反事实评分以及大语言模型(LLM)推理。CARRE检索预定义的保留行动目录,在显式特征变换下估计模型预测的流失风险变化,并为所选行动生成结构化的流失原因和基于客户档案的解释。在IBM Telco客户流失数据集上,在313个高风险测试案例中,CARRE实现的平均模型预测风险降低比普通SHAP基线高79.8%,比成本控制的SHAP+Cost基线高80.4%;其成本归一化效率比普通SHAP高10.5%。在136个案例的原因分层评估样本上,基于诊断的提示优化将弱标签一致性从79.4%提高到90.4%,且无辅助计划约束违规;由于同一批样本既用于错误诊断又用于重新评估,优化后的结果并非泛化能力的独立估计。对于使用优化前v2原因输出生成的135个解释,两位跨供应商的LLM评审员给出的平均评分在4.02到5.00分(满分5分)之间,尽管其中一位评审员在可操作性维度上达到饱和;确定性审计在66条可验证的档案声明中未发现矛盾。检索消融实验表明,在该数据集中,k=5在候选覆盖率和下游推理一致性之间提供了最佳的评估折中。这些结果展示了如何在原型流失处方流程中分离并联合评估检索、基于模型的反事实评分和语言生成。

英文摘要

Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three-stage framework that combines retrieval-augmented candidate generation, cost-aware counterfactual scoring, and large language model (LLM) reasoning. CARRE retrieves a predefined catalog of retention actions, estimates model-predicted churn-risk changes under explicit feature transformations, and generates a structured churn reason and a profile-grounded explanation for the selected action. On the IBM Telco Customer Churn dataset, CARRE achieves 79.8% greater mean model-predicted risk reduction than the plain SHAP baseline and 80.4% greater reduction than the cost-controlled SHAP+Cost baseline across 313 high-risk test cases; its cost-normalized efficiency is 10.5% higher than that of plain SHAP. On a 136-case reason-stratified evaluation sample, diagnosis-driven prompt refinement increases weak-label agreement from 79.4% to 90.4%, with no auxiliary-plan constraint violations; because the same sample was used for error diagnosis and re-evaluation, the post-refinement result is not an independent estimate of generalization. For 135 explanations generated using the pre-refinement v2 reason outputs, two cross-vendor LLM judges assign mean scores ranging from 4.02 to 5.00 out of 5, although one judge saturates on actionability, and a deterministic audit finds no contradictions among 66 verifiable profile claims. Retrieval ablations show that k=5 provides the best evaluated compromise between high candidate coverage and downstream reasoning agreement in this dataset. These results illustrate how retrieval, model-based counterfactual scoring, and language generation can be separated and jointly evaluated in a prototype churn-prescription pipeline.

Comments14pages, 1 figure, Accepted at Workshop on 5th End-to-End Customer Journey Optimization at the International Conference on Knowledge Discovery and Data Mining

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑