发表机构
CrowdStrike, Inc.; Univ. of Maryland, Baltimore County(CrowdStrike公司; 马里兰大学巴尔的摩县分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文测试多种L1逻辑回归求解器,发现旧方法优于新提出的“最先进”方案,并证明基于次梯度的LBFGS简单基线经微调后高效且易扩展,适合生产使用。
AI 中文摘要
带有$L_1$范数惩罚的线性模型在高维($d > 1,000,000$)任务中仍处于最先进水平,为解决现实世界的工业问题提供了一种直接的方法。尽管它们在工业界被广泛使用且实用,但许多$L_1$求解器对于一般用途并不有效,速度慢得令人望而却步,并且在并行化方面效果不佳。这使得它们难以在大型工业规模语料库的MLOps流水线中进行训练。在这项工作中,我们测试了文献中提出的几种“最先进”的解决方案,发现目前较旧的方法在一般用途上远优于这些方案。我们还为学术界提出了几项建议,以开展避免错误地过度自信结果的研究,这些结果可能会阻碍向生产使用的过渡。同样令人惊讶的是,我们发现一种新的简单基线方法,即在次梯度上使用LBFGS,经过微小的调整后非常有效,尽管在文献中因其理论上的不收敛而被忽视。在实践中,我们发现它是一种更易于支持、更易于扩展的生产使用方法。
英文摘要
Linear models with an $L_1$-norm penalty remain state-of-the-art for high-dimensional ($d > 1,000,000$) tasks, offering a straightforward method for solving real-world industry problems. Despite their widespread use in industry and utility, many $L_1$ solvers are not effective for general use, are prohibitively slow, and are ineffective in parallelization. This makes them difficult to train in an MLOps pipeline on large industry-scale corpora. In this work, we test several proposed ``state-of-the-art'' solutions from the literature and find that older methods are currently far superior for general use. We also identify several recommendations for academics to perform research that avoids erroneously overconfident results, which can prevent the transition to production use. Equally surprising, we find that a new and simple baseline, using LBFGS on a sub-gradient, is highly effective with minor tweaks, despite being dismissed in the literature for theoretical non-convergence. In practice, we find it is an easier-to-support and easier-to-scale method for production use.
CommentsTo appear in The 13th IEEE International Conference on Data Science and Advanced Analytics (DSAA 2026)