arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40120cs.LGmath.OC

基于分组风险集和更优LogSumExp速率的大规模Cox回归

Scalable Cox Regression via Grouped Risk Sets and Sharper LogSumExp Rates

Elizaveta Iashchinskaia, Egor Gladin

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模Cox回归的计算挑战,提出基于分组风险集和softplus替代的随机优化方法,获得更优收敛速率,并保持渐近分布匹配,实验验证性能优越。

中文摘要 AI 辅助

受大规模Cox回归计算挑战的启发,我们研究了在大规模集合上随机最小化LogSumExp目标的问题。小批量归一化器估计通常会产生有偏梯度。我们转而使用一种softplus替代函数,该函数为每个归一化器引入一个辅助标量,并允许无偏的单样本梯度。对于光滑凸LogSumExp目标,我们证明了$O(T^{-1/2})$的平均目标界,改进了之前的$T^{-1/4}$分析。在原始变量上加入强凸正则化器后,我们还获得了$\tilde{O}(T^{-1})$的最后迭代平方误差率,而无需辅助变量具有强凸性。对于Cox回归,归一化器定义在嵌套风险集上。我们利用这一结构,将相邻失败分组,并为每组共享一个辅助变量。由此产生的压缩目标具有统一的得分和曲率界,可控制分组和softplus近似带来的误差。结合一般优化结果,这些界给出了相对于完整Cox解的均方速率$T^{-4/5}$(忽略对数因子)。压缩估计器也与完整估计器的渐近分布相匹配。在合成和真实生存数据集上,风险集缓慢递减的实验表明,相对于随机基线,其性能优越。

英文摘要

Motivated by the computational challenges of large-scale Cox regression, we study stochastic minimization of LogSumExp objectives over large sets. Mini-batch normalizer estimates generally yield biased gradients. We instead use a softplus surrogate that introduces one auxiliary scalar per normalizer and admits unbiased single-sample gradients. For smooth convex LogSumExp objectives, we prove an $O(T^{-1/2})$ averaged objective bound, improving the previous $T^{-1/4}$ analysis. With a strongly convex regularizer on the original variable, we also obtain a last-iterate squared-error rate of $\widetilde{O}(T^{-1})$ without strong convexity in the auxiliary variables. For Cox regression, the normalizers are defined over nested risk sets. We exploit this structure by grouping neighboring failures and sharing one auxiliary variable per group. The resulting compressed objective admits uniform score and curvature bounds that control the errors from grouping and softplus approximation. Together with the general optimization result, these bounds give a mean-square rate of $T^{-4/5}$, up to logarithmic factors, relative to the full Cox solution. The compressed estimator also matches the full estimator's asymptotic distribution. Experiments on synthetic and real survival datasets with slowly decreasing risk sets show a favorable performance relative to stochastic baselines.

发表机构

  • HSE University(高等经济大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑