arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40030cs.LGcs.AI

Fenchel Tilting: 加权校正用于生成模型的高效微调

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

发表机构应用人工智能研究所 · 穆罕默德·本·扎耶德人工智能大学
查看机构详情
  • Applied AI Institute(应用人工智能研究所)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov

首次发表
浏览论文内容

中文总结 AI 辅助

提出Fenchel倾斜流控制(FTFC),通过解耦效用优化与生成模型拟合,利用Fenchel对偶和加权去噪实现高效微调,在图像和分子生成基准上优于基线且效率提升高达20倍。

中文摘要 AI 辅助

将预训练生成模型适应于以效用函数表达的任意偏好,是奖励对齐、引导设计和约束满足的基础,从而实现多种应用。现有的微调方法在通用性和计算成本之间进行权衡:它们要么限制支持的偏好类别以保持优化简单,要么以牺牲效率为代价保留通用性。我们引入了Fenchel倾斜流控制(FTFC),它将效用优化与生成模型拟合解耦。FTFC首先通过联合拟合预训练样本上的有效奖励和密度比权重来优化目标分布。该方法结合了效用的变分结构与Fenchel对偶性,支持一般的$f$-散度惩罚,这些惩罚决定了奖励如何转化为分布校正权重。这些权重随后被冻结,并用于在单阶段的加权去噪或流匹配中修改扩散或流模型,而无需对采样轨迹进行微分。我们在适当条件下建立了凹效用的精确对偶性,并表明加权拟合能够为给定效用重现最优目标分布。在图像和分子生成基准上,FTFC在各种偏好函数上优于基线,同时效率提高高达$20\ imes$。所提出的方法能够在无需复杂优化的情况下实现超越期望奖励最大化的适应,同时与基线相比,对更一般的效用函数类别保持鲁棒性。

英文摘要

Adapting a pretrained generative model to an arbitrary preference expressed as a utility function underlies reward alignment, guided design, and constraint satisfaction, enabling diverse applications. Existing fine-tuning methods trade off generality against computational cost: they either restrict the family class of supported preferences to keep optimization simple or preserve generality at the expense of efficiency. We introduce Fenchel Tilt Flow Control (FTFC), which decouples utility optimization from generative-model fitting. FTFC first optimizes for a target distribution by jointly fitting an effective reward and density-ratio weights on pretrained samples. Method combines the utility's variational structure with Fenchel duality, supporting general $f$-divergence penalties that determine how rewards are transformed into an distribution-correction weights. These weights are then frozen and used to modify a diffusion or flow model in a single stage of importance-weighted denoising or flow matching, without differentiating through sampling trajectories. We establish exact duality for concave utilities under suitable conditions and show that weighted fitting reproduces the optimal target distribution for a given utility. Across image and molecule generation benchmarks, FTFC improves over baselines on diverse preference functions, while also being up to $20\times$ more efficient. roposed method enables adaptation beyond expected-reward maximization without complex optimization, while preserving robustness for more general class of the utility functions compared to baselines.

↑