AI 中文总结
本文提出自解释分段树(SEST),一种面向商业分析的KPI条件分段框架,通过递归子空间划分实现节点级解释,其解释保留源单位与类别标签,将树简化为KPI极值分段,构建成本随总体规模呈二次或几何衰减。
AI 中文摘要
面对指标变动的业务用户需要了解数据哪部分发生了变动以及原因。现有数据解释方法通常返回谓词:用于隔离责任记录的属性-值条件的合取式。谓词是精确的,可直接作为筛选器执行,但它们描述的是轴对齐区域,可能无法紧凑捕获由连续趋势组合定义的分段。本文提出自解释分段树(Self-Explaining Segment Trees, SEST),其解释是针对指定关键绩效指标(Key Performance Indicator, KPI)相关性所选特征子空间中的多元聚类。SEST通过决策树代理的Shapley属性为每个KPI选择一次子空间,在递归划分总体时,通过混合模型轮廓搜索在每个节点独立选择分支因子,并为每个节点附加双重解释载荷:数值特征的标准化效应量以及用户指定维度上的类型依赖贡献轮廓。这些解释由未转换数据计算,因此呈现值保留源单位和类别标签。姿态层将任意树深度简化为其极致的KPI抑制和KPI放大分段。我们建立终止条件和由深度限制与最小分段大小决定的节点数边界,并将每棵树的构建成本表征为退化情况下与总体规模成二次关系,平衡情况下随深度呈几何衰减。这是一篇架构与方法论论文;我们未报告预测准确率或验证结果,将结果验证留待未来工作。
英文摘要
Business users confronted with a moving metric need to know which part of their data moved and why. Existing data-explanation methods typically return predicates: conjunctions of attribute-value conditions that isolate responsible records. Predicates are exact and directly executable as filters, but they describe axis-aligned regions and may not compactly capture segments defined by combinations of continuous tendencies. This paper presents Self-Explaining Segment Trees (SEST), an architecture in which an explanation is a multivariate cluster in a feature subspace selected for relevance to a designated key performance indicator (KPI). SEST selects the subspace once per KPI using Shapley attributions over a decision-tree surrogate, recursively partitions the population while choosing the branching factor independently at each node through mixture-model silhouette search, and attaches to every node a dual explanation payload: standardized effect sizes over numeric features and type-dependent contribution profiles over user-designated dimensions. These explanations are computed from untransformed data so surfaced values retain source units and category labels. A stance layer reduces any depth of the tree to its extremal KPI-suppressing and KPI-amplifying segments. We establish termination and a node-count bound determined by the depth limit and minimum segment size, and characterize per-tree construction cost as quadratic in population size in the degenerate case and geometrically decaying across depth in the balanced case. This is an architecture and methodology paper; we report no predictive-accuracy or validation results and leave outcome validation to future work.
Comments14 pages, 3 figures, 3 tables