发表机构
Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ADORE框架,利用一阶和二阶导数统一全局与局部可解释性,通过随机SVD和动态稀疏检测提升效率,在表格、文本和图像数据上优于LIME和SHAP。
AI 中文摘要
复杂机器学习模型的可解释性至关重要,尤其是在医疗和金融等现实世界的高风险领域。然而,现有的事后可解释性方法存在固有局限性:分析过程碎片化、建模非线性特征交互的能力不足、计算效率低下,以及过度依赖特定模型架构。为解决这些挑战,本文提出了一种新颖方法——自适应导数阶随机解释(ADORE)——该方法利用一阶和二阶导数来适应非线性模型的复杂性,同时能够在统一的分析框架内有效捕获特征-样本交互。ADORE将全局特征重要性与局部样本贡献相结合,通过捕获幅度和方向来精确量化特征影响,并识别影响模型决策的关键样本。此外,它通过随机奇异值分解(SVD)和动态稀疏性检测实现了计算效率,使其可扩展到大规模高维数据集。在表格、文本和图像三种数据模态上的实验表明,ADORE在处理复杂交互和计算效率方面优于LIME和SHAP等现有方法,同时提供详细且可靠的解释。为促进采用和可复现性,ADORE已作为开源Python包发布,托管在GitHub上,使研究人员和从业者能够轻松地将我们的方法适配并应用于其特定任务、模型和数据集。
英文摘要
The interpretability of complex machine learning models is of paramount importance, especially in real-world high-stakes domains such as healthcare and finance. However, existing post-hoc interpretability methods suffer from inherent limitations: fragmented analytical processes, inadequate capacity to model nonlinear feature interactions, computational inefficiencies, and over-reliance on specific model architectures. To address these challenges, this paper provides a novel method - Adaptive Derivative-Ordered Random Explanation (ADORE) - that leverages first- and second-order derivatives to accommodate nonlinear model complexities, while enabling effective capture of feature-sample interactions within a unified analytical framework. ADORE integrates global feature importance with local sample contributions, precisely quantifying feature impact by capturing both magnitude and direction, and identifying critical samples influencing model decisions. Furthermore, it achieves computational efficiency through randomized singular value decomposition (SVD) and dynamic sparsity detection, making it scalable to large, high-dimensional datasets. Experiments across three data modalities - tabular, text, and image - demonstrate that ADORE outperforms existing methods such as LIME and SHAP in handling complex interactions and computational efficiency, while providing detailed and reliable explanations. To facilitate adoption and reproducibility, ADORE has been released as an open-source Python package, hosted on GitHub, enabling researchers and practitioners to readily adapt and apply our approach to their specific tasks, models, and datasets.