发表机构
HKUST; HKUST(GZ); Tencent(香港科技大学; 香港科技大学(广州); 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大规模视频扩散模型下游任务微调计算成本高的问题,提出MagicPrompt轻量级框架。采用注意力嵌入提示调整和双空间奖励反馈优化,在可训练参数不到1%时达竞争性能,显著降低训练成本。
AI 中文摘要
大规模视频扩散模型(VDMs)具有强大的生成性能,但对下游任务进行完全微调会产生高昂的计算成本。现有参数高效微调(PEFT)方法在十亿规模模型上存在两个关键缺陷:仍需要大量可训练参数,且基于奖励的训练在条件引导任务中存在噪声诱导的优化不稳定性。我们提出了MagicPrompt,这是一个轻量级框架,实现了极高的参数效率和稳定的奖励优化。它首先采用注意力嵌入提示调整,通过轻量级软提示引导生成,参数数量少几个数量级,同时保留预训练知识。还引入了双空间奖励反馈优化,使用自监督潜在目标改进条件引导奖励训练。实验表明,MagicPrompt在可训练参数不到1%的情况下达到了有竞争力的性能,并显著降低了训练成本。
英文摘要
Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational costs. Existing parameter-efficient fine-tuning (PEFT) methods have two critical flaws on billion-scale models: they still require substantial trainable parameters, and reward-based training suffers from noise-induced optimization instability in condition-guided tasks. We propose MagicPrompt, a lightweight framework that achieves extreme parameter efficiency and stable reward optimization. It first adopts Attention-Embedded Prompt Tuning, which steers generation via lightweight soft prompts with orders of magnitude fewer parameters while preserving pre-trained knowledge. It further introduces Dual-Space Reward Feedback Optimization, which uses self-supervised latent objectives to improve condition-guided reward training. Experiments show MagicPrompt reaches competitive performance with less than 1% trainable parameters and notably reduces training costs.