AutoRef:面向智能体多参考图像生成的流程优化
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
查看机构详情
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对多参考图像生成中流程设计困难的问题,提出AutoRef方法,用编码智能体自动优化流程代码,在保持模型冻结下显著提升性能并具备跨设置泛化能力。
中文摘要 AI 辅助
最近的图像生成模型可以接受多张参考图像作为输入,并将它们组合成一张新图像。然而,多参考图像生成仍然具有挑战性:模型可能会遗漏或重复参考图像中的主体,或者生成多个主体看起来不自然拼接的图像。近期工作提出了图像生成智能体,它结合了图像生成模型、推理模型和一个流程(harness),流程是一个可执行程序,规定了参考图像如何被解释、生成如何执行、输出如何被诊断以及最终图像如何被选择。然而,在多参考生成中,参考图像扮演不同角色,输出必须同时满足许多标准,例如对每个参考的保真度和整体图像的自然度,因此流程的许多部分都可以改进,从参考图像的处理方式到输出的诊断方式。这使得难以预测哪些更改会提升性能以及提升多少,也使得好的流程难以手工设计;事实上,人工编写的流程在性能上差异很大。因此,我们提出AutoRef,它在保持两个模型冻结的同时自动优化流程:一个编码智能体迭代地重写流程代码。AutoRef将反馈用于提出建议的任务与用于选择候选的任务分开,并继续从选择任务上排名靠前的流程束(beam)中进行搜索。通过这一过程,我们发现了AutoRef-Harness,它将开放权重的FLUX.2 [klein] 4B在MultiBanana基准的留出四参考任务上的得分从5.72提升到7.37,达到或超过包括Nano Banana Pro和GPT-Image-1.5在内的专有模型。无需重新优化,当生成器、参考数量、基准、评估器或推理模型与搜索中使用的不同时,同一流程也能提升结果。
英文摘要
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.