arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08460cs.CV

InstructionCrafter:生成一致且高保真的视觉指令

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions

  • University of Tsukuba(筑波大学)

机构由 AI 辅助整理,请以论文原文为准。

Shun Okamoto, Satoshi Iizuka, Kazuhiro Fukui

AI总结:

本研究提出基于扩散的InstructionCrafter框架,通过空间冻结训练和两种适配器优化,实现了文本到分步视觉指令的生成,在多基准数据集上达成了最优性能且减少了噪声等问题,代码和模型将公开。

AI中文摘要:

给定文本任务指令,生成作为图像序列的分步视觉指令需同时满足多项属性,具体包括步骤忠实性、跨图像一致性以及单帧视觉质量。现有文本到图像生成方法很少能同时满足这三项属性,原因在于独立采样会破坏一致性、在低质量视频上微调会降低单帧质量,以及冻结的骨干网络缺乏对多步骤的理解。在本研究中,我们提出InstructionCrafter,这是一个基于扩散的框架,其核心思想是通过(1)空间冻结训练和(2)指令感知适配器,将时间和指令对齐的优化与单帧视觉质量分离开来。InstructionCrafter构建于预训练的视频扩散骨干网络之上,它冻结控制单帧细节的空间层,仅更新时间和文本条件通路以学习指令语义和步骤间关系,这既保留了单帧质量的生成先验,又比全微调减少了约50%的可训练参数。我们还引入了两个轻量适配器,以增强模型对指令上下文的理解:一致适配器(Consistent Adapter)聚合整个指令序列和相邻步骤的文本线索,以保持帧间对象身份和属性的一致性;上下文感知时间适配器(Context-Aware Temporal Adapter)将交叉注意力输出转换为时间自注意力的偏置,明确传播帧间关系。在两个基准数据集上的大量实验表明,InstructionCrafter在步骤忠实性、跨图像一致性和单帧视觉质量方面实现了最先进的整体性能,同时显著减少了噪声、模糊和虚假字幕。我们的代码和训练模型将公开提供。

英文摘要:

Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple properties, specifically step faithfulness, cross-image consistency, and per-frame visual quality. Existing text-to-image generation approaches rarely meet all three properties, owing to independent sampling that breaks consistency, finetuning on low-quality video that degrades per-frame quality, and frozen backbones that lack multi-step understanding. In this work, we propose InstructionCrafter, a diffusion-based framework with the key idea of separating the optimization of temporal and instructional alignment from per-frame visual quality via (1) spatial-freeze training and (2) instruction-aware adapters. Built on a pretrained video diffusion backbone, InstructionCrafter freezes the spatial layers that control per-frame detail and updates only temporal and text-conditioning pathways to learn instruction semantics and inter-step relations, which preserves the generative prior for per-frame quality and reduces trainable parameters by about 50 percent compared with full finetuning. We also introduce two lightweight adapters that enhance the model's understanding of instructional context. The Consistent Adapter aggregates textual cues from the entire instruction sequence and from neighboring steps to keep object identity and attributes consistent across frames, and the Context-Aware Temporal Adapter converts cross-attention outputs into biases for temporal self-attention, explicitly propagating inter-frame relations. Extensive experiments on two benchmark datasets demonstrate state-of-the-art overall performance on step faithfulness, cross-image consistency, and per-frame visual quality while significantly reducing noise, blur, and spurious subtitles. Our code and trained models will be publicly available.

↑