arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoGen-Harness:学习在何处以及如何演化图像生成外部系统

EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses

Jiabin Luo, Yinan Liu, Chunlei Meng, Yufei Guo

arXiv 2610.00383首次发表:更新:

AI 中文总结

提出EvoGen-Harness框架及Trace方法,通过失败归因引导多职责外部系统演化,在多个基准上显著提升冻结文本到图像生成器的性能。

AI 中文摘要

现代文本到图像(T2I)系统可以通过在冻结的生成器周围调整外部系统来改进,而无需修改生成器参数。然而,现有方法通常优化预定义维度,如提示、路由或工作流,限制了可纠正生成失败的空间。允许多个生成器外部职责进行演化提供了更广泛的适应空间,但引入了新挑战:视觉反馈揭示了失败之处,但并未指明持久演化应发生在何处,以及如何高效地探索该空间。我们提出EvoGen-Harness,一个生成器无关的多职责图像生成外部系统演化框架,以及Trace(轨迹相对归因与协调演化)。Trace跨随机执行聚合证据,使用失败归因作为搜索先验以聚焦候选更新,并逐步重新归因残余失败以协调跨职责的演化,而No-Patch和保留验证防止不必要或有害的更新。在GenEval2、T2I-CompBench++和WISE上,EvoGen-Harness分别比最强评估基线提高了+0.2633、+0.0720和+0.0752,同时实现了87.9-91.4%的归因召回率、94.8%的No-Patch准确率和仅1.9%的回归。这些结果表明,归因引导的多职责演化可以显著增强冻结的T2I系统,超越单维度适应。

英文摘要

Modern text-to-image (T2I) systems can be improved without modifying generator parameters by adapting the external system around frozen generators. However, existing approaches typically optimize a predefined dimension, such as prompts, routing, or workflows, restricting the space in which generation failures can be corrected. Allowing multiple generator-external responsibilities to evolve provides a broader adaptation space, but introduces a new challenge: visual feedback reveals what failed, but not where persistent evolution should occur or how this space should be explored efficiently. We introduce EvoGen-Harness, a generator-agnostic framework for multi-responsibility image-generation harness evolution, together with Trace (Trajectory-Relative Attribution and Coordinated Evolution). Trace aggregates evidence across stochastic executions, uses failure attribution as a search prior to focus candidate updates, and progressively re-attributes residual failures to coordinate evolution across responsibilities, while No-Patch and held-out validation prevent unnecessary or harmful updates. Across GenEval2, T2I-CompBench++, and WISE, EvoGen-Harness improves over the strongest evaluated baselines by +0.2633, +0.0720, and +0.0752, respectively, while achieving 87.9-91.4% attribution recall, 94.8% No-Patch accuracy, and only 1.9% regression. These results demonstrate that attribution-guided multi-responsibility evolution can substantially enhance frozen T2I systems beyond single-dimension adaptation.

Comments23 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑