arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06490cs.CV

InsertFuse:面向多类别参考引导图像插入的统一框架

InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion

Guangzhao Li, Qingyan Wei, Huayu Zheng, Yige Zheng, Chaoyang Zhang, Jie Yang, Yunan Ding, Yan Tai, Siqi Luo, Xiaohong Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出InsertFuse统一框架,通过解耦类别专属学习与跨类别整合,结合IOPD、TAGC等技术提升多类别参考引导图像插入的空间控制与生成质量,在基准测试中取得最优性能。

中文摘要 AI 辅助

我们提出InsertFuse,这是一个面向多类别参考引导图像插入的统一框架。其核心思路是将类别专属知识学习与跨类别能力整合解耦。InsertFuse首先为不同插入类别训练专属专家模型,随后引入插入在线蒸馏(IOPD)将这些专家的能力整合至单个学生模型中。通过在学生模型访问的状态下查询匹配的专家,IOPD既能保留类别专属的插入行为,又能缓解直接联合训练引发的跨类别干扰。为提升空间控制能力,我们提出令牌对齐几何条件(TAGC),该方法将掩码衍生的几何线索映射至视觉令牌网格;还提出区域平衡流匹配,该方法对插入区域内外的预测误差分别进行归一化,以避免背景主导及尺度依赖的监督问题。我们进一步引入参考分类器引导(Reference CFG),在固定场景与几何条件下隔离并强化视觉参考带来的引导,IOPD则将此增强的监督信号迁移至统一学生模型。在公开AnyInsertion基准及我们的多类别测试集上开展的大量实验,在多数指标上展现出了当前最优性能,表明其在各类插入类别中具备出色的参考保真度与生成质量。

英文摘要

We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise learning from cross-category capability consolidation. InsertFuse first trains specialized experts for different insertion categories and then introduces Insertion On-Policy Distillation (IOPD) to consolidate their capabilities into a single student. By querying the matched expert at states visited by the student, IOPD preserves category-specific insertion behavior while mitigating the cross-category interference caused by direct joint training. To improve spatial control, we propose Token-Aligned Geometry Conditioning (TAGC), which maps mask-derived geometric cues to the visual token grid, and Region-Balanced Flow Matching, which separately normalizes prediction errors inside and outside the insertion region to prevent background-dominated and scale-dependent supervision. We further introduce Reference CFG to isolate and strengthen the guidance induced by the visual reference under fixed scene and geometry conditions, with IOPD transferring this enhanced supervision into the unified student. Extensive experiments on the public AnyInsertion benchmark and our multi-category test set demonstrate state-of-the-art performance on most metrics, showing strong reference fidelity and generation quality across diverse insertion categories.

补充信息

↑