arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越试错:面向图像到视频一致性的智能体优化

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson

arXiv 2608.12290首次发表:更新:

发表机构

Google Cloud; Google DeepMind(谷歌云; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图像到视频模型试错效率低的问题,提出Agentic Self-Improvement框架,通过两阶段优化提升视频与文本一致性,生成视频胜率达69%,为视频生成模型提供实用可控的优化方法。

AI 中文摘要

现代黑盒图像到视频(I2V)模型为自动化内容创作提供了强大能力,但它们缺乏细粒度控制和可靠性,给专业工作流程带来了重大挑战。其固有的随机性导致文本提示或超参数的微小变化会产生截然不同的输出,通常需要低效、蛮力的试错过程。为解决这些限制,我们提出了“智能体自我改进(Agentic Self-Improvement)”框架,该框架将视频合成重新定义为闭环、目标导向的优化。我们的框架采用新颖的两阶段方法系统地探索生成参数空间:第一阶段,迭代提示优化循环使用多模态大语言模型(mLLM)优化输入提示,该优化执行两项自动评估:戴维森场景图(DSG)查询确保语义一致性,以及常见错误问题(CMQ)用于检测伪影;第二阶段,我们使用贝叶斯优化高效地协同优化随机种子和分类器自由引导(CFG)尺度,该搜索由一系列质量指标引导,包括从DSG和CMQ评估衍生的新型视频文本一致性(VTA)分数。我们的框架显著优于无引导搜索方法:在人类偏好研究中,通过我们的智能体方法生成的视频比基线输出更受青睐,胜率高达69%。这项工作为增强最先进视频生成模型的可预测性和控制提供了实用且可扩展的方法,推动该领域从推测性好奇转向可靠、可生产的工具。

英文摘要

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement" framework, which reframes video synthesis into a closed-loop, goal-directed optimization. Our framework systematically navigates the generation parameter space using a novel two-stage approach. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement implements two automated evaluations: Davidsonian Scene Graph (DSG) queries ensure semantic adherence, and Common Mistake Questions (CMQ) for artifact detection. At the second stage, we use Bayesian optimization to efficiently co-optimize stochastic seeds and CFG scales. This search is guided by a suite of quality metrics, including the novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. Our framework significantly outperforms unguided search methods: in human preference studies, videos generated via our agentic approach were strongly preferred over baseline outputs, achieving win rates up to 69\%. This work provides a practical and extensible methodology for enhancing the predictability and control of state-of-the-art video generation models, moving the field beyond speculative curiosities toward reliable, production-ready tools.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑