发表机构
School of Mathematical Sciences, Soochow University; Digital innovation research center, Duke Kunshan University(苏州大学数学科学学院; 昆山杜克大学数字创新研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模表格学习,提出TabLoRA方法,通过共享通用主干和引入特定预测器的低秩适配,实现参数高效的神经集成学习,在相同资源约束下平衡了预测性能与实际效率,提升了神经集成学习的可行性。
AI 中文摘要
表格学习仍以梯度提升决策树(GBDT)为主,尽管近期深度学习方法竞争力渐强。将深度表格模型应用于大规模数据集仍具挑战,因样本量、特征维度或目标类别多会带来巨大计算成本。我们提出TabLoRA,一种用于大规模表格学习的参数高效可训练神经集成方法。它跨预测器共享通用主干并引入特定预测器的低秩适配,实现无完全参数复制的集成式预测。在基准测试中,与GBDT方法和近期深度学习基线相比,TabLoRA在相同资源约束下实现了预测性能与实际效率的良好平衡。内存分析和消融研究表明该设计提高了神经集成学习的可行性,同时保留了完全集成的诸多优点。
英文摘要
Tabular learning is still dominated by gradient-boosted decision trees (GBDTs), while recent deep learning approaches have become increasingly competitive. However, applying deep tabular models to large-scale datasets remains challenging, as large sample sizes, high feature dimensionality, or many target classes can introduce substantial computational cost. We propose TabLoRA, a parameter-efficient trainable neural ensemble for large-scale tabular learning. Instead of using fully independent ensemble backbones, TabLoRA shares a common backbone across predictors and introduces predictor-specific low-rank adaptations, enabling ensemble-style prediction without full parameter duplication. Across benchmarks, TabLoRA achieves a favorable balance between predictive performance and practical efficiency compared with GBDT methods and recent deep learning baselines under the same resource constraints. Memory analysis and ablation studies further show that the proposed design improves the feasibility of neural ensemble learning while preserving much of the benefit of full ensembles.