arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32298cs.LG

刷新还是实现?漂移模型中的计算分配

Refresh or Realize? Compute Allocation in Drifting Models

Sipeng Chen, Xu Zheng, Shibo Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探讨漂移模型训练中计算分配问题,发现重新计算场比更深度拟合目标更有效,并揭示其内在原因。

中文摘要 AI 辅助

漂移模型通过每次迭代重新计算有限样本漂移场,并朝着漂移目标执行优化器步骤来训练单步生成器。该场指示生成样本应如何移动,但步骤是在所有样本共享的参数中进行的,因此网络实际做出的运动不一定与给定的运动相匹配。这留下了一个基本的训练问题:额外的计算应该用于更紧密地拟合当前目标,还是用于重新计算场?我们在ImageNet 256x256上研究这个问题。将目标固定k个优化器步骤并测量实现的位移,我们发现更深的拟合确实使网络更接近冻结的目标,并且在其取得任何净进展之前所需的步骤数从训练早期的大约十六步下降到后期的一步。当额外的步骤免费时,k=2也降低了FID。一旦需要付费,结果就会反转:在近似匹配的实测墙钟时间下,将预算花在新场上比更深的拟合产生更低的FID,在两个训练种子上都是如此。目标本身显示了为什么新场如此有价值。重新绘制有限支撑会使其方向旋转远超过参数更新(余弦约0.3-0.6对约0.95),并且在场空间中最优的校正并不在FID上可靠地优于无参数校正。对于漂移模型,良好地拟合每个目标和良好地分配计算是不同的目标。

英文摘要

Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.

补充信息

↑