arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

四边形损失:可加性作为密集神经网络的一种可测量行为

The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks

Antonio Di Cecco

arXiv 2607.20201首次发表:更新:

发表机构

Università “G. d’Annunzio” Chieti–Pescara(基耶蒂 - 佩斯卡拉“加布里埃莱·邓南遮”大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过引入四边形损失,将可加性作为密集神经网络的可测量行为。该损失能提升小数据集的准确性与可加性,还可在线观察。文中对比了通往精确可加性的多种途径,揭示了不同数据模式下排名反转及无声失败模式等情况。

AI 中文摘要

加法模型通过禁止特征交互来实现可解释性,这一约束在神经网络实例中通过架构来执行。我们引入了四边形损失,这是一种可微惩罚,它将可加性视为一种可测量行为:对交换一个坐标的训练点对进行二阶混合差分,当且仅当该坐标不携带交互时该差分消失,对分段线性网络仍然具有信息性,并且在期望上等于干预性Shapley - GAM的每个坐标的交互质量。该损失将可加性变成一个刻度盘——大多数学习到的交互几乎可以免费去除,在小数据集上适度的惩罚可以同时提高准确性和可加性——并变成一个在线可观察量:其每个特征的投降曲线表明,在不同种子和数据集上,预正则化交互幅度几乎无法预测正则化模型保留了什么,这削弱了事后交互排名。我们将通往精确可加性的途径与四边形损失进行比较,这些途径包括结构掩码、行为惩罚(可选地结晶为精确结构)、权重衰减、反向拟合、共享部分模型和袋装增强树桩:在结构主导权重空间约束之前约束行为,在不同数据模式下排名会反转,并且收敛途径在形状函数本身方面达成一致。我们记录的三种无声失败模式具有相同的剖析:在悄然使前提条件无效的设置中引入保证。

英文摘要

Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of training points swapping one coordinate, which vanishes if and only if the coordinate carries no interaction, remains informative for piecewise-linear networks, and equals in expectation the per-coordinate interaction mass of the interventional Shapley-GAM. The loss turns additivity into a dial - most learned interactions prove removable almost for free, and on small datasets a moderate penalty improves accuracy and additivity simultaneously - and into an online observable: its per-feature surrender curves show, across seeds and datasets, that pre-regularization interaction magnitude barely predicts what a regularized model retains, undermining post-hoc interaction rankings. Against this instrument we compare routes to exact additivity, spanning structural masks, behavioral penalties (optionally crystallized into exact structure), weight decay, backfitting, the shared-section model, and bagged boosted stumps: constraining behavior before structure dominates weight-space constraints, rankings reverse between data regimes, and converging routes agree on the shape functions themselves. Three silent failure modes we document share one anatomy: guarantees imported into settings that quietly void their preconditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑