arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

渐进式适应中的最优训练时间缩放

Optimal Training-Time Scaling in Gradual Adaptation

Zonghuan Xu, Krishna Harish

arXiv 2608.04927首次发表:更新:

发表机构

Fudan University; Lawrence E. Elkins High School(复旦大学; 劳伦斯·E·埃尔金斯高中)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对过参数化线性回归任务,推导得出渐进式适应中最优单任务训练时间缩放为Θ(N⁻¹),实验验证了路径划分越精细应减少单任务训练的结论。

AI 中文摘要

在渐进式适应中,随着中间任务数量的增加,每个任务的训练时间应如何变化?我们针对过参数化线性回归任务研究该问题,此类任务平滑变化且共享零损失解。当任务数为N、每个任务训练时间为s_N时,若Ns_N→τ,最终学习进展会收敛到连续曲线。对于小τ,极限进展为Θ(τ);对于大τ,极限进展为Θ(τ⁻¹),因此过短或过长的训练都会导致进展甚微。由此可得,最优单任务训练时间缩放为s_N^⋆=Θ(N⁻¹),等价于Ns_N^⋆=Θ(1)。在渐进旋转MNIST和自然年鉴时间偏移上的实验结果表明,随着路径划分更精细,减少单任务训练的结论成立。

英文摘要

In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overparameterized linear regression tasks that change smoothly and share a zero-loss solution. With $N$ tasks and training time $s_N$ on each, the final learning progress converges to a continuum curve when $Ns_N\toτ$. The limiting progress is $Θ(τ)$ for small $τ$ and $Θ(τ^{-1})$ for large $τ$, so both very short and very long training produce little progress. It follows that optimal per-task training times scale as $s_N^\star=Θ(N^{-1})$, equivalently $Ns_N^\star=Θ(1)$. Experiments on gradually rotated MNIST and a natural Yearbook time shift are consistent with less per-task training as the path is divided more finely.

Comments26 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑