arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

样本加权的端到端迹范数几何用于多任务学习

Sample-Weighted End-to-End Trace-Norm Geometry for Multitask Learning

Mahdi Mohammadigohari

arXiv 2609.29520首次发表:更新:

发表机构

Free University of Bozen–Bolzano(博尔扎诺自由大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多任务学习,提出样本加权端到端迹范数几何,推导精确Rademacher复杂度,实验表明加权联合核正则化在252次比较中平均超额降低0.00764,优于未加权方法。

AI 中文摘要

多任务模型将共享表示与任务特定输出相结合,但泛化界通常分别控制这两个组成部分。这种分离乘积可能丢弃相对方向和相消信息,并且即使在所表示的预测器不变的情况下,也可能在中间坐标的等价变换下发生变化。我们转而研究从任务系数到输入空间预测器的端到端映射的样本量加权迹范数。对于其固定半径类,我们推导出了精确的经验Rademacher复杂度。同一量通过消除表示作用后的正定任务协方差来表征,并且在有限维中间空间中,通过在所有等价的可逆重构因式分解上优化分离乘积来表征。显式构造展示了无界的方向和因式分解间隙,以及用于相消线性层的指数深度间隙。作为几何应用,有限对一的Lipschitz共享映射产生一个由重数和局部方向畸变确定的精确Sobolev任务Gram矩阵。我们在两个协议锁定的未见套件中评估相应的凸正则化器。在252次配对保留比较中,加权联合核正则化相比未加权核正则化将平均总体超额降低了0.00764,分层自助法95%区间为[0.00465, 0.01110]。正确的任务计数相比偏移计数将平均和最少采样四分位超额分别改善了0.01072和0.02847;所有15个不平衡秩套件单元均为正,平衡效应为零。加权联合核还优于加权Frobenius和独立岭回归。与未加权核的最少采样四分位比较仍未解决,界定了而非反驳平均优势。所有七个预先声明的门均通过。

英文摘要

Multitask models combine a shared representation with task-specific outputs, but generalization bounds often control the two components separately. Such products can discard relative orientation and cancellation and can change under equivalent transformations of intermediate coordinates even when the represented predictors are unchanged. We study instead the sample-size-weighted trace norm of the end-to-end map from task coefficients to input-space predictors. For its fixed-radius class, we derive the exact empirical Rademacher complexity. The same quantity is characterized by eliminating a positive-definite task covariance after the representation acts and, in finite-dimensional intermediate spaces, by optimizing the separated product over all equivalent invertible refactorizations. Explicit constructions show unbounded orientation and factorization gaps and an exponential depth gap for cancelling linear layers. As a geometric application, finite-to-one Lipschitz shared maps yield an exact Sobolev task Gram matrix determined by multiplicity and local directional distortion. We evaluate the corresponding convex regularizer in two protocol-locked unseen suites. Across 252 paired held-out comparisons, weighted joint nuclear regularization improves average population excess over unweighted nuclear regularization by 0.00764, with a stratified-bootstrap 95% interval [0.00465, 0.01110]. Correct task counts improve average and least-sampled-quartile excess over shifted counts by 0.01072 and 0.02847; all 15 imbalanced rank-suite cells are positive and the balanced effect is zero. Weighted joint nuclear also outperforms weighted Frobenius and independent ridge. The least-sampled-quartile comparison with unweighted nuclear remains unresolved, delimiting rather than contradicting the average advantage. All seven predeclared gates pass.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑