arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从密集预测到视觉编辑:用于统一图像与视频创作的结构化监督

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

Zhefan Rao, Bin Zou, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Haoxuan Che, Qifeng Chen

arXiv 2608.14740首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Celia Research HK; City University of Hong Kong(香港科技大学; Celia Research HK; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于深度与表面法向量预测的结构化视觉监督框架,共享MMDiT主干实现统一图像视频创作,提升了相关基准任务分数,证明密集监督对下游创作的结构知识迁移价值。

AI 中文摘要

统一图像和视频创作要求模型遵循多样指令,同时保留视觉上下文的身份、几何与时间结构。然而,仅语义条件和仅创作训练并未明确监督精确、时间一致编辑所需的局部结构。因此,我们将深度和表面法向量预测构建为图像形式的去噪目标,在同一创作界面内使用这些密集任务作为结构化视觉监督。我们的框架将语义解释与空间对齐的视觉注入解耦,同时在所有任务间共享一个多模态扩散Transformer(MMDiT)主干。互注意力上下文(MCA)、配对视频数据构建流程以及渐进式训练课程,随后将学习到的结构线索连接到时间局部编辑和参考条件创作。单个检查点在报告的统一系统比较中获得最高总体分数(4.15);添加密集监督将OpenVE总体分数从3.98提升至4.06,局部添加分数从3.92提升至4.18。这些结果支持一个有明确边界的结论:面向感知的密集监督可将有用的结构知识迁移到下游创作,尤其是编辑局部性与保留性;我们不主张其作为独立密集预测器的优越性。

英文摘要

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editing. We therefore formulate depth and surface-normal prediction as image-form denoising targets, using these dense tasks as structured visual supervision within the same creation interface. Our framework decouples semantic interpretation from spatially aligned visual injection while sharing one multimodal diffusion transformer (MMDiT) backbone across all tasks. Mutual Context Attention (MCA), a paired-video data-construction procedure, and a progressive training curriculum then connect the learned structural cues to temporally localized editing and reference-conditioned creation. A single checkpoint obtains the highest overall score in the reported comparison of unified systems (4.15); adding dense supervision improves OpenVE Overall from 3.98 to 4.06 and Local Add from 3.92 to 4.18. These results support a deliberately bounded conclusion: perception-oriented dense supervision transfers useful structural knowledge to downstream creation, especially editing locality and preservation; we do not claim superiority as a standalone dense predictor.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑