arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MSR:视频生成的多主体参考

MSR: Multiple Subject Reference for Video Generation

Guannan Li, Jiaji Chen, Jingyuan Liao, Yu Geng, Baolan Qiu

arXiv 2609.18393首次发表:更新:

发表机构

Licon Studio(Licon Studio)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MSR,一种基于槽感知条件化的视频生成方案,通过独立编码参考图像并添加槽嵌入与时间偏移,减少多参考混淆,并发布权重与推理流程。

AI 中文摘要

在多个图像条件下生成视频,需要在保留外观的同时,将每个参考图像与其预期角色关联起来。我们提出了MSR(多主体参考),一种用于基于LTX的视频生成的槽感知条件化方案。每个参考图像被独立编码为静态片段,并由单独的潜在令牌组表示。一个紧凑的傅里叶特征多层感知器添加了数值槽嵌入,而依赖槽的时间偏移修改了该组的旋转坐标。参考组被前置到噪声目标令牌之前,并在仅目标流的匹配训练期间作为干净上下文。我们通过低秩适应实现该方案,并发布了所得权重和推理工作流。定性示例展示了在写实和风格化场景中包含不同角色和参考环境的组合。开发观察表明,与早期连续参考基线相比,参考混淆有所减少,但相似服装、复杂衣物和视角变化仍具挑战性。我们描述了条件化机制、保留的训练配置以及所发布系统的观察到的优势和局限性。一个补充的音频参考实验在保持视觉参数冻结的情况下添加了语音条件化。

英文摘要

Conditioning a video generator on multiple images requires preserving appearance while associating each reference with its intended role. We present MSR (Multiple Subject Reference), a slot-aware conditioning scheme for LTX-based video generation. Each reference image is independently encoded as a static clip and represented by a separate latent-token group. A compact Fourier-feature multilayer perceptron adds a numeric slot embedding, while slot-dependent temporal offsets modify the group's rotary coordinates. The reference groups are prepended to noisy target tokens and serve as clean context during target-only flow-matching training. We implement this scheme through low-rank adaptation and release the resulting weights and inference workflows. Qualitative examples demonstrate compositions containing distinct characters and referenced environments in realistic and stylized scenes. Development observations suggest reduced reference confusion relative to an earlier continuous-reference baseline, while similar clothing, complex garments, and viewpoint changes remain challenging. We describe the conditioning mechanism, the retained training configuration, and the observed strengths and limitations of the released system. A supplementary audio-reference experiment adds voice conditioning while keeping the visual parameters frozen.

Comments11 pages, 4 figures. Model weights and inference workflows are publicly available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑