arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优化决策而非预测:平滑净收益作为训练目标的探索

Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective

Koen M. F. Gorgels, Lasai Barreñada, Maarten van Smeden, Ben Van Calster, Ewout W. Steyerberg, Wouter A. C. van Amsterdam

arXiv 2609.12752首次发表:更新:

发表机构

University Medical Center Utrecht; Cooperatie VGZ; KU Leuven(乌得勒支大学医学中心; VGZ合作社; 荷语鲁汶大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探索平滑净收益作为训练目标,发现其在逻辑回归中略有收益,但对灵活模型无益,不能替代负对数似然训练。

AI 中文摘要

目的:预测模型通常使用诸如伯努利负对数似然(NLL)等目标进行训练,尽管下游临床决策可能依赖于特定的风险阈值。我们引入了平滑净收益($\sigma$NB),这是净收益的一种可微近似,旨在使模型训练与特定阈值的临床效用对齐。材料与方法:我们评估了$\sigma$NB作为逻辑回归、广义加性模型(GAMs)和XGBoost的训练目标,并使用了三种Hessian实现。实验使用了Framingham心血管风险数据集和44个TabZilla数据集,共包含72个数据集-阈值组合。结果:在Framingham数据集中,$\sigma$NB训练并未持续改善净收益。在TabZilla基准测试中,逻辑回归的平均标准化净收益从使用NLL时的0.5669增加到使用$\sigma$NB时的0.5765(平均差异0.0096,95%置信区间-0.0001至0.0193)。对于GAMs,平均标准化净收益从0.5921下降到0.5625(平均差异-0.0296,95%置信区间-0.0721至0.0129)。对于XGBoost,NLL达到了0.6745,而$\sigma$NB各实现的值为0.6723至0.6735。在逻辑回归中,$\sigma$NB的收益与XGBoost相对于NLL训练的逻辑回归的性能优势呈正相关。讨论:$\sigma$NB的效果依赖于上下文,收益主要集中在逻辑回归中,对更灵活的模型类别几乎没有益处。这表明,当模型灵活性有限、留有更大改进空间时,决策聚焦优化可能最为有用。结论:我们的结果不支持$\sigma$NB作为NLL训练的通用替代方案,但支持在传统基于似然的训练可能无法充分捕捉决策相关结构的环境中进一步研究决策聚焦目标。

英文摘要

Objective Prediction models are commonly trained using objectives such as Bernoulli negative log-likelihood (NLL), although downstream clinical decisions may depend on specific risk thresholds. We introduce Smooth Net Benefit ($σ$NB), a differentiable approximation of Net Benefit designed to align model training with threshold-specific clinical utility. Materials and Methods We evaluated $σ$NB as a training objective for logistic regression, generalized additive models (GAMs), and XGBoost with three Hessian implementations. Experiments used the Framingham cardiovascular risk dataset and 44 TabZilla datasets comprising 72 dataset-threshold combinations. Results $σ$NB training did not consistently improve Net Benefit in Framingham. Across the TabZilla benchmark, mean standardized Net Benefit for logistic regression increased from 0.5669 with NLL to 0.5765 with $σ$NB (mean difference 0.0096, 95% CI -0.0001 to 0.0193). For GAMs, mean standardized Net Benefit decreased from 0.5921 to 0.5625 (mean difference -0.0296, 95% CI -0.0721 to 0.0129). For XGBoost, NLL achieved 0.6745 compared with 0.6723--0.6735 across $σ$NB implementations. In logistic regression, $σ$NB gains were positively associated with the performance advantage of XGBoost over NLL-trained logistic regression. Discussion The effect of $σ$NB was context dependent, with modest gains concentrated in logistic regression and little benefit for more flexible model classes. This suggests that decision-focused optimization may be most useful when limited model flexibility leaves greater scope for improvement. Conclusion Our results do not support $σ$NB as a general replacement for NLL training, but support further investigation of decision-focused objectives in settings where conventional likelihood-based training may not adequately capture decision-relevant structure.

Comments22 pages, 5 figures, 2 tables. Code and supplementary data available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑