arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35189cs.CVcs.LG

G$^3$-LoRA:利用梯度引导的分组LoRA组织奖励加权视频数据

G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Jia Song, Wenhow Li, Lichen Bai, Bada Ye, Zeke Xie

AI总结:

针对异构奖励加权视频数据后训练中的更新冲突问题,提出G$^3$-LoRA方法,通过梯度兼容性聚类分组训练LoRA专家并合并蒸馏,在Wan2.1和CogVideoX上优于基线,验证了梯度兼容性作为数据组织诊断的有效性。

AI中文摘要:

对异构奖励加权数据上的基础视频模型进行后训练通常假设所有数据类别会引发兼容的更新。当类别对应不同的技能、领域或评估维度时,这一假设是脆弱的。我们在文本到视频的后训练中研究此问题,其中VBench2.0维度定义数据桶,外部多模态奖励管道分配样本权重。我们提出G$^3$-LoRA(梯度引导的分组LoRA),一种数据组织程序,它探测由奖励加权视频样本引发的类别级梯度,移除共享的全局更新方向,根据残差梯度兼容性对类别进行聚类,训练组特定的LoRA专家,并通过权重合并随后从专家进行策略内蒸馏将它们整合为一个适配器。我们通过将奖励加权流匹配视为速度场回归来激励此程序:不兼容的奖励维度可能在重叠的噪声潜在区域中偏好不同的去噪方向,导致共享LoRA训练平均化能力。在Wan2.1-T2V-1.3B-Diffusers上,合并的分组适配器在匹配的VBench2.0评估上优于基础模型、联合奖励加权LoRA基线以及使用相同管道训练的随机、语义和原始梯度分区;独立评估者同意此结论,并且在CogVideoX-2B上分组避免了联合训练的负迁移。增益并不均匀:合并压缩了最大的专家增益,蒸馏恢复了部分损失,并且相机运动和几个局部质量维度仍然具有挑战性。综合来看,这些结果表明梯度兼容性可以作为组织奖励加权视频后训练数据的实用诊断工具。

英文摘要:

Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.

补充信息

↑