开放评估智能体:视觉生成模型的高效且可提示的评估
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
- S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
- Agency for Science, Technology and Research (A*STAR)(科学、技术与研究局(A*STAR))
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出Evaluation Agent框架及基于其构建的Open-EA,可高效评估视觉生成模型,将评估时间减至传统方法的10%,还通过EA-CoT-10K和EA-3B减少对专有骨干的依赖,验证了其在多基准及跨家族的有效性。
AI中文摘要:
视觉生成模型的最新进展已实现高质量的图像和视频生成,但评估这些模型通常需要采样数百或数千张图像或视频,计算成本高昂。现有评估方法还依赖僵化的流程,忽略特定用户需求,仅提供无清晰解释的数值结果。借鉴人类仅从少量样本中快速形成对模型能力印象的方式,我们提出Evaluation Agent(评估智能体)框架,该框架采用类人策略进行高效、动态的多轮评估,提供详细、适配用户需求的分析。给定自然语言评估请求,智能体将其分解为多个子方面,生成针对性提示,从被评估模型中采样图像或视频,调用合适的评估工具,并根据观察到的证据迭代更新计划,涵盖预定义的基准维度和开放式用户关注点。该框架因此具备高效性、可提示性、可解释性,且可跨模型和工具扩展。实验表明,Evaluation Agent将评估时间减少至传统方法的10%,同时提供可比结果。我们进一步通过构建EA-CoT-10K引入Open Evaluation Agent(Open-EA),这是一个源自多轮评估滚动的历史条件步骤级指令微调记录语料库,并基于Qwen2.5-3B-Instruct训练EA-3B作为本地规划骨干,该骨干保留了基于API的智能体的结构化推理、工具调用和总结协议,同时减少了对专有骨干的依赖。实验验证了基于API的智能体在既定T2I/T2V基准和开放式查询上的性能,并在四个域内和三个域外T2V生成器家族上评估Open-EA,显示学习到的策略存在部分跨家族迁移。
英文摘要:
Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses. Given a natural-language evaluation request, the agent decomposes it into sub-aspects, generates targeted prompts, samples images or videos from the evaluated model, invokes suitable evaluation tools, and iteratively updates its plan from the observed evidence, covering both predefined benchmark dimensions and open-ended user concerns. The framework is thus efficient, promptable, explainable, and scalable across models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. We further introduce Open Evaluation Agent (Open-EA) by constructing EA-CoT-10K, a corpus of history-conditioned step-level instruction-tuning records derived from multi-round evaluation rollouts, and training EA-3B from Qwen2.5-3B-Instruct as a local planning backbone that preserves the structured reasoning, tool invocation, and summary protocol of the API-based agent while reducing dependence on proprietary backbones. Experiments validate the API-based agent on established T2I/T2V benchmarks and open-ended queries, and evaluate Open-EA on four in-domain and three out-of-domain T2V generator families, showing partial cross-family transfer of the learned policy.