发表机构
DGIST; Baidu, Inc.; KAIST(大邱科技研究院; 百度公司; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无需组合目标监督的CRAFT框架,通过LoRA适配器微调参考感知MMDiT,仅用1万张参考样本,在XVerseBench上实现图像主题个性化的最先进性能,且可迁移至其他骨干网络。
AI 中文摘要
主题驱动的图像个性化是指生成在新场景中保留一个或多个参考主题身份的新图像,这是现代视觉内容创作的基础能力。当前该领域主要由通用方法主导,这类方法会在数十万至数百万对(参考图像、组合目标)示例上微调预训练的多模态扩散Transformer(MMDiT),其中每个组合目标是主题在新场景中的合成图像。生成此类组合目标需要成本高昂的多阶段策划流程,包括基于大语言模型(LLM)的提示词生成、基于文本到图像(T2I)的组合目标合成、参考主题提取、基于视觉语言模型(VLM)的质量过滤以及对应关系标注,且这类方法会紧密绑定特定的目标合成器和策划选择。本文提出CRAFT(Constrained Reward via Attention Fine-Tuning,注意力微调约束奖励),这是一种单步的奖励反馈学习(ReFL)框架,它通过LoRA适配器微调预训练的参考感知MMDiT,使用仅包含1万张参考图像和主题掩码的紧凑参考数据集,无需组合目标监督。CRAFT实现了“关注何处”原则:注意力级奖励将噪声标记和短语标记的注意力与正确的参考主题对齐,生成的每个主题的注意力掩码会控制像素级身份奖励,使图像空间监督与学习到的注意力路径保持一致。将其应用于FLUX.2-klein-9B时,CRAFT在XVerseBench上达到了最先进的性能,且未使用任何组合目标监督,仅需1万张仅参考样本,而此前的通用方法需要15万至超过200万对组合目标。该方法的相同方案可迁移至其他参考感知骨干网络,持续提升性能。项目页面:this https URL。
英文摘要
Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene. Producing such targets demands a costly multi-stage curation pipeline---LLM-based prompt generation, T2I-based composed-target synthesis, reference-subject extraction, VLM-based quality filtering, and correspondence labeling---and tightly couples each method to a particular target synthesizer and curation choice. We introduce \emph{CRAFT} (Constrained Reward via Attention Fine-Tuning), a single-step ReFL framework that fine-tunes a pre-trained \emph{reference-aware} MMDiT via LoRA adapters using a compact reference-only data construction---$10$K reference images and subject masks, with no composed-target supervision. CRAFT realizes a \emph{Where to look} principle: attention-level rewards align noise- and phrase-token attention with the correct reference subject, and the resulting per-subject attention masks gate a pixel-level identity reward to keep image-space supervision consistent with the learned attention routing. Applied to FLUX.2-klein-9B, CRAFT achieves state-of-the-art performance on XVerseBench \rev{while using no composed-target supervision---only $10$K reference-only samples, whereas prior generalized methods require $150$K to over $2$M composed-target pairs}. The same recipe transfers to other reference-aware backbones, consistently improving performance. Project page: https://jihun999.github.io/projects/CRAFT/.
Comments20 pages, 8 figures, ACM SIGGRAPH Asia 2026