arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05196cs.LGq-fin.RMstat.ML

在匹配可容许性下测量学习的单调时间聚合

Measuring Learned Monotone Temporal Aggregation at Matched Admissibility

Yew Lee Tan

AI总结:

本研究在固定可容许性下比较学习与手工单调时间聚合,发现学习通道在可分离族内可替代工程方法,但存在结构性失败,且无单一现有方法全面占优。

AI中文摘要:

风险监管对评分施加方向性约束;我们采用其严格的逐输入形式——评分在每个暴露输入上单调非递减——作为规范性承诺。已部署的流水线(由单调手工聚合特征输入符号约束梯度提升)通过组合已满足该约束,因此约束与无约束的比较为现有方法免费获得的保证定价。我们改为在两侧固定可容许性,并衡量学习聚合的价值。我们的工具是一个循环网络,其状态为经典风险统计量(带学习变换的指数加权移动平均和高水位标记),在构造上对每个输入和每个MC-dropout样本均单调。通过函数回归,核心发现是一个包含边界:学习的单调通道复现了手工统计量的几何加权可分离族,每个成员一个通道,Spearman ρ ≥ 0.996,以可测量的上限逼近窗口统计量,并在连续性和时间定位上失败(ρ = 0.924 和 0.628),两者均为结构性的,以及在暴露下限(0.829)上失败,这是一个可学习性边界。一个显式的可容许基修复了每个失败(秩相关 1.000)。在可分离族内或附近,学习聚合和工程聚合是可替代的,并且在完整样本量和指定容量下,学习通道在统计上从不落后。其优势是现有方法特定的:一个固定的网格在其未覆盖的衰减区域付出高达 0.019 AUC 的代价(学习通道在每个扫描点上保持在最强工程消费者的 0.004 以内);最高维比较器在数据稀缺时退化最快;超出训练支持范围,网格输入的树集成评分趋于平坦,而严格递增的头部保持排序。没有单一现有方法在所有三个轴上被支配。

英文摘要:

Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-constrained gradient boosting -- already satisfy it by composition, so constrained-versus-unconstrained comparisons price a guarantee the incumbent has for free. We instead hold admissibility fixed on both sides and measure what learning the aggregation is worth. Our instrument is a recurrent network whose state is classical risk statistics (an exponentially weighted moving average and a high-water mark with learned transforms), monotone by construction in every input and per MC-dropout sample. The central finding, by functional regression, is a subsumption boundary: a learned monotone channel reproduces the geometrically weighted separable family of hand-crafted statistics, one channel per member, to Spearman $ρ\ge 0.996$, approximates window statistics with measurable ceilings, and fails at consecutivity ($ρ= 0.924$) and time localization (0.628), both structural, and at the exposure floor (0.829), a learnability boundary. One explicit admissible basis repairs each failure (rank correlation 1.000). In or near the separable family, learned and engineered aggregation are substitutes, and the learned channel is never statistically behind at full sample size and specified capacity. Its advantages are incumbent-specific: a committed grid pays up to 0.019 AUC in decay regions it leaves uncovered (the learned channel stays within 0.004 of the strongest engineered consumer at every swept point); the highest-dimensional comparator degrades fastest with scarce data; and beyond the training support, grid-fed tree-ensemble scores go flat while a strictly increasing head keeps ranking. No single incumbent is dominated on all three axes.

补充信息

↑