缩放定律、表格数据与精算费率厘定模型
Scaling Laws, Tabular Data and Actuarial Ratemaking Models
浏览论文内容
中文总结 AI 辅助
该研究探究精算费率厘定场景下的缩放规律,对比不同模型发现TabM的数据缩放能力更强,Transformer需额外归纳偏差才能实现有效参数缩放,为该场景的模型选择提供定量指导。
中文摘要 AI 辅助
现代深度学习中的缩放定律描述了保留损失如何随模型容量、训练数据量和计算资源的增加而改善,通常遵循幂律趋势。我们研究在精算费率厘定中是否会出现类似的缩放规律,精算费率厘定的数据为表格型、异质性强且存在噪声,而广义线性模型(GLMs)等经典模型仍是强大的基准。我们使用真实世界的汽车保险业务组合,训练不同系列的模型,覆盖不断增加的训练数据比例和多个随机种子,评估样本外泊松偏差——这是一种用于泊松计数预测的基于似然的损失,值越低表示保留拟合效果越好。我们发现所有模型系列都会随数据增加而改善,但缩放指数差异显著:TabM相比纯监督式表格Transformer和标准多层感知机(MLP)基准,表现出明显更强的数据缩放能力。Transformer变体显示出较弱的参数缩放能力,除非通过额外的归纳偏差(TabM式适配或自监督)进行增强。这些结果为按数据场景进行模型选择提供了定量指导,并表明精算表格任务的有效缩放依赖于架构和损失函数目标设计,单纯增大Transformer规模带来的增益有限。
英文摘要
Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.
发表机构
- insureai
- Bayes Business School(贝叶斯商学院)
机构由 AI 辅助整理,请以论文原文为准。