AI 中文总结
StraightDP利用整流流的几何异质性,通过释放类条件矩结合DP-SGD,在强差分隐私约束下提升文本条件生成模型的下游准确率与生成质量,且可迁移到SD3-medium。
AI 中文摘要
文本条件生成模型的差分隐私(DP)训练在强隐私约束下会出现效用断崖。我们通过整流流的几何重新审视该问题:在噪声与数据的线性插值过程中,噪声端的贝叶斯最优速度主要由少数类条件矩决定,而越靠近数据端,样本特异性结构的影响越大。StraightDP全程利用这种异质性:将一小部分预算一次性释放白化后的类条件矩,用于蒸馏到权重或在采样时注入;剩余预算则通过预声明的DP-SGD用于数据端,超出矩的覆盖范围。在MNIST上,当ε=1时,仅释放的矩就能达到0.76的下游准确率,生成类似原型的样本,FID为237,而均匀DP-SGD仅达到0.21。基于该释放构建的 pipeline 在公共潜在空间中达到0.81的准确率,FID为56。约束多模态主干的每token流范数不会改变预训练损失,但在极端噪声像素空间场景中提升了下游准确率,且随着隐私增强,其准确率效果单调提升。释放的矩还可迁移到冻结的SD3-medium,采样时注入的效果以一小部分预算击败DP-LoRA训练。
英文摘要
Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. We revisit this problem through the geometry of rectified flows: along the straight interpolation between noise and data, the Bayes-optimal velocity is governed to leading order at the noise end by a few class-conditional moments, and increasingly sample-specific structure matters toward the data end. StraightDP exploits this heterogeneity end to end. A small budget share releases whitened class-conditional moments once, to be distilled into the weights or injected at sampling time. The rest is spent by pre-declared DP-SGD toward the data end, beyond the moments' reach. At $\varepsilon=1$ on MNIST, the released moments alone already attain $0.76$ downstream accuracy with prototype-like samples and an FID of $237$, and uniform DP-SGD attains $0.21$. The pipeline built on the release reaches $0.81$ accuracy at FID $56$ in a public latent space. Constraining per-token stream norms of the multimodal backbone leaves the pretraining loss unchanged yet improves downstream accuracy in the extreme-noise pixel-space regime, and its accuracy effect becomes monotonically more favorable as privacy strengthens. The released moments also port to frozen SD3-medium, where sampling-time injection beats DP-LoRA training at a fraction of the budget.