发表机构
İnönü University(伊诺努大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Lambert惩罚,通过限制保留比例的对数收缩实现稀疏回归,在预测与变量选择间平衡,模拟和QSAR应用中表现竞争性。
AI 中文摘要
稀疏回归必须在预测准确性与可复现的变量选择之间取得平衡。我们通过限制标量分数在变量进入后保留比例的增长速度,引入了Lambert惩罚。在此尺度等变约束下最大化保留比例,可得到从精确零到恒等尾部的对数过渡。该有界惩罚具有固定形状,并通过仅使用训练集的五折交叉验证选择单一惩罚参数。我们推导了显式标量更新、尖锐的弱凸性界、多元唯一性的充分条件以及精确循环坐标下降的条件收敛性。对于固定维度,全局最小化器在渐近调参带上与oracle估计量一致等价。高斯模拟将Lambert与LASSO、岭回归、弹性网、MCP和SCAD进行比较。Lambert在强等信号下偏好预测,并且通常比SCAD改善支持恢复;MCP通常提供更好的支持恢复,而凸方法在弱信号或强相关性下仍具竞争力。在QSAR毒性应用中,重复嵌套交叉验证给出了具有竞争性的预测和紧凑、稳定的模型,尽管与MCP和SCAD的误差差异较小。实证设计遵循探索性开发;独立确认仍然必要。
英文摘要
Sparse regression must balance prediction accuracy with reproducible variable selection. We introduce the Lambert penalty by limiting how quickly the retained fraction of a scalar score increases after variable entry. Maximizing retention under this scale-equivariant constraint yields a logarithmic transition from exact zero to an identity tail. The bounded penalty has a fixed shape and a single penalty parameter selected by training-only five-fold cross-validation. We derive an explicit scalar update, a sharp weak-convexity bound, a sufficient condition for multivariate uniqueness, and conditional convergence of exact cyclic coordinate descent. For fixed dimension, global minimizers are uniformly equivalent to the oracle estimator over an asymptotic tuning band. Gaussian simulations compare Lambert with LASSO, ridge, elastic net, MCP, and SCAD. Lambert favors prediction for strong equal signals and usually improves support recovery over SCAD; MCP often gives better support recovery, and convex methods remain competitive with weak signals or strong correlation. In a QSAR toxicity application, repeated nested cross-validation gives competitive prediction with compact, stable models, although error differences from MCP and SCAD are small. The empirical design followed exploratory development; independent confirmation remains necessary.
Comments33 pages including a 8-page technical supplement; 2 figures, 7 tables