arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模型无关的混合效应与多层模型影响力异常值检测

Model-Agnostic Influential Outlier Detection for Mixed Effects and Multi-Level Models

Colin C Jones, David A Campbell, Yan Liu

arXiv 2610.01720首次发表:更新:

发表机构

Institute for Data Science, Carleton University; School of Computer Science, Carleton University; Department of Psychology, Carleton University(卡尔顿大学数据科学研究所; 卡尔顿大学计算机科学学院; 卡尔顿大学心理系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对混合效应模型提出模型无关的影响力异常值检测方法,结合SHAP值与残差并利用归一化流,在多种模型中验证其优势与局限。

AI 中文摘要

影响力异常值检测是针对聚类数据上的混合效应模型而开发的。影响力异常值度量被定义为SHapley加性解释(SHAP)值与模型残差的组合,两者均经过测度变换。基于先前工作展示的归一化流在将任意分布映射到灵活的基础分布以进行统计推断方面的适用性,所构建的归一化流允许包含上下文信息,并为模型评估提供拟合优度诊断。在构建中使用SHAP值摆脱了特定于模型的工具,而是提供逐点的模型无关影响力异常值。在包括线性模型、随机森林和梯度提升树在内的多种模型中,考察了该方法的优势与局限性。

英文摘要

Influential Outlier Detection is developed for mixed-effects models on clustered data. The Influential Outlier Metric is defined as a combination of SHapley Additive exPlanantion (SHAP) values and model residuals, both of which undergo a change of measure transformation. Building on previous work showcasing the suitability of using Normalizing flows to map arbitrary distributions to a flexible base distribution for statistical inference, the Normalizing Flows are constructed to allows contextual information and also provide a goodness of fit diagnostic for model evaluation. The use of SHAP values in the construction moves away from model specific tools and instead provides point-wise model agnostic influential outlier. The advantages and limitations of this approach are examined in several models including the linear model, the random forest, and gradient-boosted trees.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑