arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13560cs.CVcs.AIcs.CL

AutoDesign:面向长视距智能体设计的元工具优化

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

  • Meituan(美团)
  • MBZUAI(Mohamed bin Zayed University of Artificial Intelligence)
  • Huazhong University of Science and Technology(华中科技大学)
  • Peking University(北京大学)
  • Tsinghua University(清华大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li

AI总结:

AutoDesign是符合人类设计先验的元工具优化框架,以论文转海报生成任务为实例,在PosterBench上性能优于Claude Design,集成其学习的DesignHarness可提升代码智能体性能,且获人类最高偏好。

AI中文摘要:

将多模态源转化为浓缩且结构化的媒体输出,可从根本上被概念化为以模型-工具系统为核心的长视距智能体过程。理想的工具系统应符合人类设计先验,并通过经验探索积累可复用的经验以驱动递归自我改进,而现有范式仍为静态,缺乏此能力。本文提出AutoDesign框架,其符合人类设计先验,其中元工具优化器指导代码智能体基于rollout反馈递归改进工具。为实例化和评估该框架,我们聚焦学术论文到海报的生成任务,引入PosterBench,包含覆盖五个学科的100篇论文主赛道,以及用于受控评估的共享10篇论文子集PosterBench-mini。在PosterBench主赛道上,AutoDesign取得78.32的最高分,超过闭源商业系统Claude Design 7.45分。在七个受控代码智能体-模型配置中,集成学习得到的DesignHarness始终提升性能,将平均PosterBench分数从54.99提高至67.39(+12.4%)。在完全自主的长视距循环中,它在40分钟内执行253次工具调用和11次编辑轮次,成本低于3美元,在人类评估中达到会议海报的平均质量;系统盲测人类研究进一步表明,AutoDesign在被评估系统中获得最高的人类偏好。

英文摘要:

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback. To instantiate and evaluate this framework, we focus on the academic paper-to-poster generation task and introduce PosterBench, comprising a 100-paper Main Track spanning five disciplines and PosterBench-mini, a shared 10-paper subset for controlled evaluation. On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points. Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under $3, reaching average conference-poster quality in human evaluation. A system-blind human study further demonstrates that AutoDesign achieves the highest human preference among evaluated systems.

补充信息

相关深度报道

↑