发表机构
BBVA(毕尔巴鄂比斯开银行)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个六步骤端到端因果机器学习流水线,处理连续治疗变量,包含正性违反检测、两阶段降维等方法,在金融债务催收合成数据上验证,比标准方法更高效且偏差更小。
AI 中文摘要
本文提出了一种端到端的因果机器学习(ML)流水线,专为具有连续治疗变量的实际应用而设计。该框架由六个连续步骤组成:降维、因果识别、正性假设违反处理、估计、反驳与评估以及策略优化。我们引入了现有因果ML工具包中尚不可用的实际贡献,具体包括:(1)一种在连续治疗设置中检测和量化正性违反的方法;(2)一种新颖的、可扩展的两阶段降维框架,专为高维数据的因果推断而设计;(3)将原本针对二元治疗设计的敏感性分析和估计方法适配到连续治疗空间;(4)将这些组件端到端集成到一个模块化、可复现的工作流程中。这些创新解决了因果推断中的实际挑战,这些挑战通常未在理论框架中涵盖,但在工业应用中经常遇到。该方法通过一个受真实世界金融债务催收用例启发的合成数据集进行了验证,但其设计可应用于不同行业的类似问题。结果表明,与标准方法相比,所提出的方法在处理连续治疗和高维数据的问题时提供了更具计算效率的方法,并产生了偏差更小的估计。我们提供了一个功能完整的GitHub仓库,包含文档化代码和编号笔记本,以确保可复现性和实际实施。所提出的流水线旨在弥合学术方法与行业实际应用之间的差距,在金融部门等因果ML可能高度有益的行业背景下尤其如此。
英文摘要
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assumption violation handling, estimation, refutation and evaluation, and policy optimization. We introduce practical contributions not currently available in existing causal ML toolkits, specifically: (1) a method for detecting and quantifying positivity violations in continuous treatment settings (2) a novel, scalable two-stage dimensionality reduction framework tailored for causal inference with high-dimensional data; (3) the adaptation of sensitivity analysis and estimation methods originally designed for binary treatments to the continuous treatment space and (4) an end-to-end integration of these components into a modular, reproducible workflow. These innovations address real-world challenges in causal inference that are often not covered in theoretical frameworks but frequently encountered in industrial applications. The methodology is validated with a synthetic dataset inspired in a real-world financial debt collection use case, however its design can be applied to analogous problems across different industries. Results demonstrate that the proposed methodology offers a more computationally efficient approach and produces less biased estimates compared to standard methods for problems with continuous treatment and high-dimensional data. A fully functional GitHub repository with documented code and numbered notebooks is made available ensuring reproducibility and practical implementation. The pipeline presented is intended to contribute to closing the gap between academic approaches and practical application in industry contexts where causal ML can be highly beneficial such as the financial sector.
CommentsOral presentation at the 3rd Workshop on Causal Inference and Machine Learning in Practice, KDD 2025, Toronto. Code: https://github.com/javiermoralh/causal-pipeline