AI 中文总结
本文研究四种构造性分类器生长方法,发现将叶节点变为门以加深软决策树时,新门梯度恒为零而无法学习,并提出随机扰动修复;其他方法各有收益,如拟合残差减小网络、分裂最大误差叶节点提升稀疏性。
AI 中文摘要
构造性分类器在训练过程中增加结构:向树添加一层,向隐藏层添加一个单元,在叶节点进行分裂。本文在固定协议下,针对树结构和构造性模型,探究四种此类生长决策各自的实际收益,并对其中一种无收益的决策给出精确诊断和修复方案。诊断涉及软化决策树最自然的加深方式:将每个叶节点变为一个门,其两个子节点继承父节点的类别分布,从而保持函数不变。我证明这会使每个新门的梯度恒为零,且当门处于1/2时,两个子节点梯度相同,因此新增层永远无法学习。与Net2Net通过噪声打破对称性或分裂最陡下降通过二阶信息逃离鞍点不同,这里的一阶信息并非微弱而是完全缺失。在三次五折交叉验证中,与从头训练的相同深度相比,该构造在Iris上损失19.6个准确率点,在Wine上损失19.1个,在Digits上损失55.6个。修复方案是对子节点施加小幅随机扰动,其大小影响甚微。实用规则是测试中的一行代码:添加参数后,断言其梯度非零。其他三种决策各有其收益。在安装新隐藏单元前将其拟合到残差误差,能在每个数据集上获得更小的网络,但并非更准确,且在Digits上显著牺牲准确率。分裂预期误差最大的叶节点能获得稀疏性,在3.7次分裂时达到0.885,而完整深度六的树需63次分裂,但在更难的问题上损失4.3个点。在节点获得更表达性的分裂前要求统计显著性则无收益:树变得更大且更不准确。论文中每个数字均来自测量脚本。
英文摘要
Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions buys, measured under one protocol on 24 datasets, and gives an exact diagnosis and fix for the one that buys nothing. The diagnosis concerns the natural way to deepen a soft decision tree: turn every leaf into a gate whose two children inherit the parent's class distribution, so the function is unchanged. I prove that this leaves the gradient of every new gate identically zero and with the gate at 1/2, gives the two children identical gradients, so under this construction the added level can never learn. That predicts where it costs: nothing on two-class problems, where two leaves already suffice and a great deal where more classes need more leaves. Measured, the cost is -0.2 points over 13 binary datasets and 40.2 points over 7 multi-class ones, and it tracks the number of classes (Spearman 0.64), not the number of features (0.03). The fix is any perturbation of the children. Neither its size nor its direction matters: a residual-directed initialisation changes accuracy by +0.20 points against noise. The other three decisions each buy one thing. Fitting a new hidden unit to the residual before installing it buys a smaller network but not a more accurate one. Splitting one leaf at a time buys sparsity; it lost accuracy until I found that the split started its two children identical and untrained, the same defect in another place; with the children inheriting the parent and the symmetry broken, per-leaf growth comes within 2.3 points of a complete tree using 23% of its splits. Requiring statistical significance before a node gets a more expressive split buys nothing. Every number comes from the measurement scripts.
Commentsv2: twenty OpenML datasets added, the per-leaf growth defect found and fixed, claims scoped to the construction. 13 pages, 7 tables. Code and measurement scripts: https://doi.org/10.5281/zenodo.22718897