V-Engram:用于模块化文本到图像个性化的触发索引外部记忆
V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
浏览论文内容
中文总结 AI 辅助
提出V-Engram,一种触发索引外部记忆机制,为Stable Diffusion 3.5实现模块化文本到图像个性化,通过分离记忆与主干适应,实现无需合并更新的多概念访问,匹配DreamBooth-LoRA保真度并支持组合。
中文摘要 AI 辅助
预训练文本到图像模型包含广泛的视觉知识,然而它们无法仅从少量参考中可靠地获取或细化特定视觉身份,同时保持组合控制。令牌嵌入方法紧凑但往往对身份拟合不足,而基于适配器的方法通过持续权重更新提高保真度,但存储成本高且在概念组合时可能产生干扰。我们引入V-Engram,一种用于Stable Diffusion 3.5的触发索引外部记忆机制。每个概念被分配一个显式触发词,该触发词检索概念特定记忆,其门控方向作为相对残差进入冻结的文本编码器和MMDiT上下文状态。将这种记忆与主干适应分离,使得无需合并模型更新即可实现提示选择和多概念访问。实验表明,V-Engram在整体主体保真度上广泛匹配DreamBooth-LoRA,同时在上下文主体保留等设置中显示出优势。提示匹配加载仅检索匹配条目,减少单概念查询的额外适应状态加载。定性结果进一步展示了配对触发组合和同类分离,而未注册条目的提示保留冻结模型的基行为。总之,这些结果确立了触发索引记忆作为添加针对性视觉证据而无需重写生成器的模块化接口。
英文摘要
Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.
发表机构
- AutoArk-AI
机构由 AI 辅助整理,请以论文原文为准。