arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可审计分类器模型学习:源不相交树集成

Learning Auditable Classifier Models: Source-Disjoint Tree Ensembles

Srikumar Krishnamoorthy

arXiv 2608.15725首次发表:更新:

发表机构

Indian Institute of Management Ahmedabad(印度管理研究所艾哈迈达巴德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出残差模式树集成(RPTE)三阶段学习方法,实现可审计性,在12个临床二分类基准上性能与对比模型相当,且审计复杂度显著低于XGBoost、EBM。

AI 中文摘要

临床及受监管场景中的预测模型必须兼具准确性与完全可审计性。树集成在表格数据上表现出较强的准确性,但它们的序列提升将结构发现与系数估计耦合在一起,使得紧凑的逐预测审计变得困难。可解释替代方案则施加了限制表达性的结构约束:广义加性模型通常将交互作用限制为成对项,而事后规则提取器会产生重叠规则,阻碍紧凑解释。我们引入残差模式树集成(Residual Pattern Tree Ensemble, RPTE),这是一种三阶段学习方法,基于三个关键原则:受限特征预算、源不相交性和独立系数估计。第一阶段构建有监督符号特征词汇表;第二阶段在源不相交性约束下生长浅层树,其中每个原始变量最多分配给一棵树,且仅保留发现的树结构;第三阶段在叶区域指示符上求解单个ℓ₁正则化逻辑回归,得到联合最优的稀疏系数。该学习方法确保每个预测可分解为命名的、不重叠规则贡献的代数和,从设计上实现完全可审计性。通过在12个临床领域二分类基准上进行重复分层5折交叉验证的实证评估显示,RPTE与调优后的黑箱集成及可解释基线相比具有竞争力;相对于XGBoost,RPTE将模型检查单元减少9倍至87倍,且在全部12个数据集上保持比EBM更低的审计复杂度;RuleFit在其规则数量较少的3个数据集上需要相当或更少的检查单元,但不具备源不相交性保证。源代码可在[this https URL]获取。

英文摘要

Predictive models in clinical and regulated settings must be accurate and fully auditable. Tree ensembles deliver strong accuracy on tabular data, but their sequential boosting couples structure discovery with coefficient estimation, making compact per-prediction auditing difficult. Interpretable alternatives impose structural constraints that limit expressiveness: generalized additive models typically restrict interactions to pairwise terms and post-hoc rule extractors produce overlapping rules that hinder compact interpretation. We introduce Residual Pattern Tree Ensemble (RPTE), a three-stage learning approach, that is built on three key principles: bounded feature budget, source disjointness, and separate coefficient estimation. Stage~1 builds a supervised symbolic feature vocabulary. Stage~2 grows shallow trees under a source-disjointness constraint, where each raw variable is allocated to at most one tree, and retains only the discovered tree structures. Stage~3 solves a single $\ell_1$-regularized logistic regression over leaf-region indicators, yielding jointly optimal sparse coefficients. This learning approach ensures that every prediction decomposes into an algebraic sum of named, non-overlapping rule contributions, enabling full auditability by design. Empirical evaluation on twelve clinical-domain binary classification benchmarks using repeated stratified 5-fold cross-validation shows that RPTE performs competitively against tuned opaque ensembles and interpretable baselines. RPTE reduces model inspection units by 9$\times$ to 87$\times$ relative to XGBoost and maintains lower audit complexity than EBM on all 12 datasets. RuleFit requires comparable or fewer inspection units on three datasets where its rule count is small, but without source-disjointness guarantees. The source code is available at \href{https://github.com/srikumar2050/hugiml-core}{this https URL}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑