发表机构
Westlake University; Hangzhou Normal University(西湖大学; 杭州师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TDAction,通过软终端控制选择传输目标,优化训练动力学以实现一步生成,在ImageNet上无需蒸馏达到FID低于1.1。
AI 中文摘要
一步生成模型通过迭代的训练时传输构建静态生成器。现有的传输目标主要评估分布运动,尽管神经生成器需要通过共享参数更新联合实现所需的样本位移。训练时构建提出了一个问题:\n\emph{一旦训练成为构建最终一步映射的迭代过程,应该优化什么:下一步的分布移动,还是有限生成器学习最终映射的路径?}\n为了解决这个问题,我们引入了训练动力学动作(TDAction),它根据局部共享参数实现成本选择传输目标,同时保持规定的分布进展水平。我们将该成本表述为软终端控制问题,并推导出闭式批切线动作到价值(Batch Tangent Action-to-Go),该值考虑了参数努力和终端不匹配。该准则捕捉了独立成对成本忽略的跨样本交互;在各向同性移动性下,该准则与确定性平衡耦合的二次欧几里得分配一致。随机切线探针提供了一种低秩实现,构建共享的分离目标,而无需增加推理时轨迹。受控研究考察了生成器几何、传输选择和实现的局部动作之间的关系。在ImageNet $256\times256$上,TDAction在无需蒸馏的情况下实现了低于$1.1$的FID。
英文摘要
One-step generative models construct a static generator through iterative training-time transport. Existing transport objectives primarily assess distributional motion, although a neural generator needs to realize the requested sample displacements jointly through shared parameter updates. The training-time construction raises the question: \emph{once training becomes the iterative process that constructs the final one-step map, what to optimize: the next distributional move, or the route by which the finite generator learns the final map?} To address the question, we introduce \textbf{T}raining \textbf{D}ynamics \textbf{A}ction (\textbf{TDAction}), which selects transport targets according to local shared-parameter realization cost while retaining a prescribed level of distributional progress. We formulate the cost as a soft-terminal control problem and derive a closed-form Batch Tangent Action-to-Go value that accounts for parameter effort and terminal mismatch. The criterion captures cross-sample interactions omitted by independent pairwise costs; under isotropic mobility, the criterion agrees with quadratic Euclidean assignment for deterministic balanced couplings. Randomized tangent probes provide a low-rank implementation that constructs shared detached targets without adding an inference-time trajectory. Controlled studies examine the relationship between generator geometry, transport selection, and realized local action. On ImageNet $256\times256$, TDAction attains an FID below $1.1$ without distillation.