arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

联合稀疏迁移学习用于高维多输出回归

Joint-Sparse Transfer Learning for High-Dimensional Multi-Output Regression

Sunwoo Lim, Mladen Kolar

arXiv 2609.30879首次发表:更新:

发表机构

Marshall School of Business, University of Southern California; Mohamed bin Zayed University of Artificial Intelligence(南加州大学马歇尔商学院; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维多输出回归,提出联合稀疏迁移学习框架,结合融合与去偏估计器,平衡源域信息与偏差,理论误差界及实验验证其有效性。

AI 中文摘要

多任务线性模型可以通过利用响应之间共享的结构来改进估计和预测,从而在逐任务拟合与完全合并之间取得平衡。然而,在许多应用中,目标是在数据受限的目标域中进行估计,同时存在数据丰富但异质的源域。从这些源域借用信息可以提高效率,但可能引入偏差。我们开发了一种用于高维多输出回归的联合稀疏迁移学习框架,该框架将跨响应的共享预测变量结构与源-目标相似性相结合。该框架产生两个互补的估计器:一个融合估计器,聚合联合拟合的域特定系数;以及一个基于目标的去偏估计器,调整源引起的偏移。我们的误差界表明,迁移如何增加可用信息,以及跨响应共享预测变量如何降低选择成本。它们还揭示了一个权衡:融合估计器受益于更大的源域,但如果源偏移指向相似方向,则可能保留偏差,而去偏则将这种偏差换取由较小的目标样本控制的额外估计误差。与极小极大下界的比较识别了边界匹配到对数因子的机制、匹配仍未解决的机制,以及投影到基于目标的凸集以弥合差距的机制。模拟实验和跨细胞类型的单细胞RNA及表面蛋白谱分析支持了该理论。

英文摘要

Multitask linear models can improve estimation and prediction by exploiting structure shared across responses, bridging taskwise fitting and complete pooling. In many applications, however, the objective is estimation in a data-limited target domain, while data-rich but heterogeneous source domains are available. Borrowing from these sources can improve efficiency but introduce bias. We develop a joint-sparse transfer-learning framework for high-dimensional multi-output regression that combines shared predictor structure across responses with source-target similarity. The framework yields two complementary estimators: a fused estimator that aggregates jointly fitted domain-specific coefficients and a target-based debiased estimator that adjusts for source-induced shifts. Our error bounds show how transfer increases the available information and sharing predictors across responses reduces selection costs. They also reveal a tradeoff: the fused estimator benefits from larger sources but may retain bias if source shifts point in similar directions, whereas debiasing trades this bias for additional estimation error governed by the smaller target sample. Comparison with a minimax lower bound identifies regimes where the bounds match up to logarithmic factors, where matching remains unresolved, and where projection onto a target-based convex set closes the gap. Simulations and an analysis of single-cell RNA and surface-protein profiles across cell types support the theory.

Comments69 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑