arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14681cs.CV

ReBind:通过具有显式参考关系的结构化指令进行多参考视频编辑

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

发表机构香港科技大学 · 克林团队 · 南京大学
另 3 家 · 查看机构详情
  • HKUST(香港科技大学)
  • Kling Team(克林团队)
  • NJU(南京大学)
  • UCAS(中国科学院大学)
  • PKU(北京大学)
  • HKU(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinyu Liu, Shihao Li, Weihong Lin, Xinlong Chen, Yang Shi, Yujin Han, Yiyang Cai, Yanghao Wang, Ruibin Yuan, Yuanxing Zhang, Pengfei Wan, Wenhan Luo, Yike Guo

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多参考图像条件视频编辑中现有方法难以协调多视觉源信息的问题,提出ReBind框架,通过带嵌入参考令牌的语义指令及两阶段渐进方案学习建立显式绑定,实现轻量级适配,在指令质量和性能上表现出色。

中文摘要 AI 辅助

近期基于扩散的视频生成模型在多参考图像条件视频编辑方面取得显著进展,但现有方法难以准确协调来自多个视觉源的信息。现有编辑指令缺乏显式参考关系,多数多模态大语言模型无法可靠生成。为此提出ReBind框架,引入带嵌入参考令牌的语义指令作为中间表示。开发了ReBind-Instruct学习建立显式绑定,还开发了ReBind-Edit实现轻量级适配。实验表明ReBind在指令质量上大幅超越通用多模态大语言模型,在开源方法中取得最优性能。

英文摘要

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordinate information from multiple visual sources accurately. We identify a critical deficiency in existing approaches. Existing editing instructions lack explicit reference relationships, and most multimodal large language models (MLLMs) cannot generate them reliably. To address this problem, we propose ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing. Our key insight is embedding reference tokens at semantic positions to eliminate ambiguity and establish precise bindings between visual attributes and their sources. We develop ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships. We further develop ReBind-Edit, which enables lightweight adaptation of text-to-video models to coordinate multiple references by binding visual attributes to their designated sources. Extensive experiments demonstrate that ReBind substantially outperforms general-purpose MLLMs in instruction quality and achieves state-of-the-art performance among open-source methods on reference image conditioned video editing. Our project webpage: https://rebind-mrv2v.github.io/.

补充信息

↑